# Snapshots & recovery

> Whole-deployment snapshots, offsite copies, retention floors, and rebuilding a host that is gone.

Source: https://oddk.dev/docs/snapshots/
Project: https://github.com/AndrianBdn/oddk


A **snapshot** is one archive containing every instance's data *and* oddk's own
configuration. It is what a host migration or a disaster recovery restores from.

```bash
oddk snapshot make                      # capture everything, now
oddk snapshot setup-cron --utc-hour 3   # every night at 03:00 UTC
oddk snapshot list                      # what exists, and where the copies are
```

## Physical by default

Running instances are captured with `pg_basebackup` over the replication
protocol — fast, no locks, no long transaction. Stopped instances are captured as
a cold file copy of the data directory, which is a real, restorable capture.

Physical restores are byte-for-byte, so per-database GUCs, database-level ACLs
and ICU collations all survive.

Use `--logical` for a portable `pg_dump`-based archive when you need to cross
CPU architectures, or when unlogged tables must survive the restore.

## Configure offsite copies

Create a bucket in S3 or your S3-compatible storage, with credentials that can
list the bucket and upload, download and delete objects under your chosen
prefix. Give each deployment its own bucket or prefix so their retention
jobs do not share a destination.

Allow `s3:AbortMultipartUpload` on the archive prefix so failed or canceled
uploads can clean up their uploaded parts. Configure a bucket lifecycle rule
to abort incomplete multipart uploads left by a daemon crash or loss of
connectivity. For SSE-KMS buckets, the upload identity also needs `kms:Decrypt`
and `kms:GenerateDataKey` on the encryption key.

On the oddk server, export the configuration template to a private file:

```bash
(umask 077; oddk offsite get > offsite.json)
```

Edit `offsite.json`. The configuration is **one JSON object**:

```json
{
  "type": "s3",
  "bucket": "my-backup-bucket",
  "endpoint": "",
  "region": "us-east-1",
  "accessKeyId": "YOUR_ACCESS_KEY_ID",
  "secretAccessKey": "YOUR_SECRET_ACCESS_KEY",
  "bucketPath": "db01/",
  "ec2IamRole": false
}
```

Replace the bucket, region and credentials. Leave `endpoint` empty for AWS S3;
for S3-compatible storage, use its endpoint URL and region. A non-empty
`bucketPath` such as `db01/` must end with `/` and must not begin with `/`.
An empty path means the bucket root.

On an EC2 host using an instance role, set `ec2IamRole` to `true` and leave both
access-key fields empty. When editing an existing configuration,
`%SAME-AS-BEFORE%` keeps its stored secret.

Apply the file, then test the connection:

```bash
oddk offsite apply --file offsite.json
oddk offsite test
```

The test uploads, downloads and deletes a test file. After it passes, remove the local
JSON file if it contains credentials you no longer need on disk.

Scheduled snapshots use this destination automatically. To prove an actual
snapshot reaches the bucket, make one and find its ID:

```bash
oddk snapshot make --comment "offsite check"
oddk snapshot list
```

Replace `7` below with that snapshot's ID:

```bash
oddk snapshot upload 7
oddk snapshot list-remote
oddk offsite logs --limit 50
```

With offsite configured, failed uploads keep their local archive past retention
for retry.

Snapshot archives contain database contents in plaintext; the master key
encrypts stored credentials, not the archive. Control access to the bucket
and any downloaded copies.

## Back up the master key

The installer stores the key at `/var/lib/oddk/data/master.key`. If you run
the daemon with a custom `--data-dir`, the key is inside that directory instead.

**The key is not inside a snapshot.** Copy it to protected storage separate
from this server, such as an encrypted recovery drive. Confirm that drive is
mounted at `/mnt/oddk-recovery` before running these commands on the server:

```bash
sudo install -m 600 /var/lib/oddk/data/master.key \
  /mnt/oddk-recovery/db01-master.key
sudo cmp /var/lib/oddk/data/master.key \
  /mnt/oddk-recovery/db01-master.key
```

`cmp` exits successfully with no output when the copy matches. Detach the
drive and store it securely. A copy
left on the same server is not protection against losing that server.

Keep the source key with the recovery records for this deployment, along with
its bucket, endpoint, region and a way to obtain storage credentials without
the original host. A replacement host needs both an archive and the matching
source key.

## Retention

Choose local and offsite retention windows when setting the schedule. For example,
this captures nightly at 03:00 UTC, keeping local copies for 7 days and offsite
copies for 30 days:

```bash
oddk snapshot setup-cron --utc-hour 3 \
  --cleanup-local-days 7 --cleanup-remote-days 30
```

Retention preserves the newest **two surviving copies in each location**, even
past their age limit. The newest complete archive is also protected for that
location's retention window plus **30 days**. That extra protection expires;
address incomplete captures reported by `oddk checklist` rather than relying
on an old complete archive indefinitely.

## Restore one instance

The rest of the deployment stays up.

```bash
oddk snapshot restore-instance --instance app --id 7
```

## Rebuild a whole host

On an empty replacement host, [install oddk](../install/) with Docker running.
Use the same CPU architecture for a physical snapshot. The host must have
enough resources for the recorded instances.

Retrieve your archive from the bucket and your separately saved source key.
Place them at the paths below, readable by the `oddk` user; that user also
needs permission to traverse their parent directories. Keep the key private.

Stop the newly installed service, then apply. This runs locally against the
data directory, so recovery works even when the daemon cannot start:

```bash
sudo systemctl stop oddk
sudo -u oddk oddk snapshot apply \
      --file /mnt/restore/snapshot-db01-20260827.tar.zst \
      --master-key /mnt/restore/master.key
sudo systemctl start oddk
```

{{< callout type="warning" >}}
Use the **source host's key**, not the key created by the fresh installation.
Apply checks it before installing configuration. Do not delete the new host's
key by hand; apply replaces it and preserves the previous one.
{{< /callout >}}

Apply pauses every restored schedule, because the archive carries the *source*
host's offsite settings — an unpaused restore would upload into that bucket and
run retention against it. Resume deliberately, once this host owns its bucket:

```bash
oddk snapshot setup-cron --resume
```

On a rehearsal host, keep schedules paused or configure a separate offsite
destination before resuming. If the snapshot also contains per-instance backup
schedules, resume each deliberately with
`oddk backup setup-cron --instance NAME --resume`.

Check the result and test notification delivery:

```bash
oddk list
oddk checklist
oddk notify test
```

## Every archive is verified

Archives are fsynced, read back, and structurally asserted before being
catalogued — and verified again on download, before anything is allowed to depend
on them. A corrupt archive is refused rather than silently restored.

