Skip to content

Restoring with Keeper

Every restore is a Restore resource. The console, the CLI and the API all create one. Restores never write into an existing database unless you ask for in place and type the target's name. Pick the mode:

Need Mode Role
Look at data, check a fix, run a report sandbox (default) operator
Get one table or a few rows back sandbox, then copy what you need operator
A copy next to production (new database name) new restorer
A dump file for someone outside the cluster download (presigned URL, 1 h) restorer
Roll the whole target back in place (after a safety backup) restorer-admin

1. Find the point

keeper targets                                   # PITR windows and last backups
keeper backups team/app                          # versions (ids, times, sizes, verified, pinned)
keeper changes team/app --since 6h               # what changed when: per-table counts and DDL per segment
keeper diff <backup-a> <backup-b>                # schema and row changes between two versions

In the console, open the target, then Changes. Each segment row has "sandbox before" and "sandbox after" links that fill in the time.

With PITR, any second inside a window works (UTC, RFC3339): --at 2026-01-02T03:04:05Z. Without PITR, pick a backup id: --at 20260102T020000Z-1a2b3c. If the time is not restorable, Keeper says why and prints the windows.

2. Restore into a sandbox (safe, the default)

keeper restore team/app --at 2026-01-02T03:04:05Z --ttl 6h --wait
keeper sandbox show <sandbox>                    # connection info, expiry
keeper query <sandbox> -e "select count(*) from orders where created_at > now() - interval '1 day'"
keeper query <sandbox> -e "select * from orders where id = 42" -o csv > order42.csv
  • The sandbox gets its own namespace keeper-sbx-<name>, random credentials and a TTL (default 24 h, extend with keeper sandbox extend <sandbox> --by 24h, up to the policy maximum).
  • The query console is read-only by default, with a 30 s timeout and a 1,000-row limit. Columns matching the target's mask patterns are masked. Every query is audited.
  • From another pod: give it the label keeper.republic.global/sandbox-client: <sandbox> and connect to the in-cluster address that sandbox show prints.
  • From a laptop, when the sandbox was created with --expose: keeper sandbox connect <sandbox> --port 15432. This opens the tunnel through Cloudflare Access and prints the connection string.
  • keeper sandbox delete <sandbox> when done. Expired sandboxes are deleted automatically.

3. Restore into a new database

keeper restore team/app --at 2026-01-02T03:04:05Z --to new \
  --host postgres.team.svc --database app_restored --credentials-secret keeper-admin --wait

Keeper first restores into a staging sandbox, runs the checks, then copies the result into the new database with a logical dump and load (ADR 0007). It never overwrites an existing database. The admin secret (keeper- prefix) lives in the target's namespace and needs CREATE DATABASE.

4. Download

keeper restore team/app --at <backup-id> --to download --wait

--wait prints a presigned URL to a decrypted, recompressed dump. It is valid for 1 hour. The URL is a secret: share it like one.

5. In place (last resort)

keeper restore team/app --at 2026-01-02T03:04:05Z --to in-place --confirm app --credentials-secret keeper-admin --wait
  1. Stop the writers first. Keeper does not stop applications. Scale them down, or put them in maintenance mode.
  2. Keeper takes a safety backup of the current state (reason safety). Skipping it needs --skip-safety-backup.
  3. It restores into a staging sandbox, checks the result, then replaces the target's databases.
  4. Start the applications again and check them.

If anything goes wrong, the safety backup is a normal version of the target: restore it the same way.

Without the CLI

The Restore resource is the API:

apiVersion: keeper.republic.global/v1alpha1
kind: Restore
metadata: { name: app-incident-42, namespace: team }
spec:
  target: app
  source: { time: "2026-01-02T03:04:05Z" }     # or backupID: ..., or latest: true
  destination:
    mode: sandbox
    sandbox: { ttl: 6h }

kubectl -n team get restore app-incident-42 -w shows the progress: plan, fetching base, replaying, checking, Succeeded.

When Keeper itself is down

The store is the source of truth, and the backup format is plain: - manifest.json lists every object with its size and SHA-256. - Objects are zstd-compressed and age-encrypted: age -d -i key.txt obj.zst.age | zstd -d. - Postgres physical backups are pg_basebackup tar streams plus WAL segments. - Postgres logical backups are pg_dump -Fc files, one per database. - MySQL backups are mysqldump output plus binlog files.

keeper catalog import rebuilds the Backup objects from S3 after a cluster loss.