Skip to content
SupaCovedocs

11 / 13

Restore & disaster recovery

Recovery kits, standalone restores without the metadata database, and the three key incidents

Routine restore: the recovery kit

AGE_IDENTITY_FILE=/path/identity.txt \
PGPASSWORD='<target db password>' \
sh restore-job7.sh 'postgresql://user@host:5432/restored?sslmode=require' backup-job7.dump.age

Order of operations: verify the ciphertext SHA-256 (mismatch exits 1 before any decryption) → refuse a non-empty target database → age-decrypt to a temporary plaintext → pg_restore → when the manifest carries a table count, assert EQUALITY and print it. A failed restore may leave partial writes: clean the target before retrying. Create the target database first (createdb). The manifest's backupId is the instance-local id (job-N), not the bucket object key (objects are keyed by the backup UUID).

Supabase backups

Their kit restores with a supabase profile: it needs a server that has Supabase's extensions, a new empty database and a superuser connection, and it checks all three before writing anything. SUPABACKUP_PROFILE=generic or supabase overrides a kit's default. Walkthrough: Back up a Supabase database.

Standalone restore: without this instance

If the instance is gone entirely, recovery needs exactly:

  1. the *.dump.age and its *.manifest.json from the bucket (or staging);
  2. your offline age identity;
  3. the recovery kit script, or manual age-decrypt + pg_restore guided by the manifest hash and metadata (the full commands live in the repository playbook, scenario 3).

The restore host needs its own tools: the kit checks for age, pg_restore, psql and sha256sum (or shasum) and exits 3 when missing — the service image embeds the age Go library, not a CLI for your restore box.

The manifest is self-describing: backup id (job-N, the instance-local id), server version, table count, SHA-256, recipient fingerprint; the bucket object key's backup UUID lives in the instance metadata database. The SQLite metadata database is NOT a prerequisite for recovery.

Key incidents (four cases)

Master secret file missing, age identity intact

Ciphertext stays decryptable. The instance generates a fresh master secret and keeps running — stored connection strings and destination credentials become unreadable and must be re-registered. Try to recover the original key from backups before rebuilding anything.

Master secret corrupted / bad permissions / replaced

Corruption and permission problems refuse startup (fail closed): repair the file first. A key that was merely replaced (valid format, different value) lets the instance start, but every old credential fails to decrypt — equivalent to losing it; rebuilding then requires re-entering database connection strings and destination credentials and re-binding them.

Age identity lost, master secret intact

Existing ciphertext is unrecoverable. The correct path for future protection: preserve the old data directory and recipient record as audit evidence (do not reuse them), then on a NEW instance (empty recipient config) run age init, store and age verify the new identity offline, re-register databases and destinations, restore schedules, and prove the chain with a fresh backup + restore. age init is not a rotation interface — an instance with a configured recipient refuses it. Writing the OLD recipient into the new instance (the DR playbook's scenario-1 recipe, valid only while the old identity still works) would keep producing backups nobody can decrypt.

Both lost

Existing backups are unrecoverable. Immediately: pause the old instance, preserve the data directory and bucket objects (for audit), then complete age init + destination credentials + database re-registration on a new instance so the protection chain restarts.

Mapping to the repository playbook

The command-level playbook lives in docs/disaster-recovery.md (developer documentation; until the repo has a remote, open the local path docs/disaster-recovery.md). Scenario index:

  • Scenario 1 (SQLite corruption / data-volume loss) → start a fresh instance with an empty data dir, bootstrap, write the original recipient back (only while the old age identity still works), re-register databases/destinations and schedules; historical data recovery follows scenario 3.
  • Scenario 2 (master secret unavailable: missing / corrupted / replaced) → the first two key-incident steps above.
  • Scenario 3 (whole instance gone, only bucket ciphertext + offline age identity left) → "Standalone restore"; object keys are <prefix>backups/<backup-uuid>.dump.age.
  • Scenario 4 (failed upgrade / interrupted migration) → pre-migrate snapshot rollback, isolate the old WAL/SHM, boot the previous binary.

Last updated

On this page