Skip to content

0007: Restore paths: receiver for physical sandboxes, staging sandbox for other destinations

Date: 2026-10-07 · Status: accepted

Context

Keeper never mounts anything and touches databases only over the network. A physical Postgres restore must place files into a data directory before the server starts. The age identity must stay in Keeper's namespace (restorer Jobs only). The design also asks for new-database and in-place restores, and for downloads.

Decision

  • Physical sandbox restores. The sandbox pod has an init container running keeper receiver. It listens on TCP 9000 and requires a random token from the sandbox secret. The restorer Job (in Keeper's namespace, with the identity) decrypts the base backup and the WAL needed. It streams them as a tar with recovery.signal, recovery settings and a sandbox pg_hba.conf. The receiver writes the pod's own emptyDir. The database container then runs a small wrapper: the official entrypoint starts Postgres, recovery replays to the target time and promotes, and the wrapper creates or resets the keeper_sandbox superuser through the local socket. The pod becomes ready.
  • Logical sandbox restores. An empty official image with keeper_sandbox (Postgres) or root (MySQL) and a random password; the restorer loads over the network.
  • New database and in-place. The restore goes into a short-lived staging sandbox first (any mode, any point in time), then each database is copied with the engine's dump/load (pg_dump -Fc | pg_restore) into the destination. In-place drops and recreates the databases after an automatic safety backup.
  • Downloads. Produced from the staging sandbox as zstd-compressed logical dumps under downloads/, with a 1-hour presigned URL. GC removes them after 2 hours.
  • Restorer Jobs have backoffLimit: 0. The controller retries up to 3 times and resets the sandbox first (a fresh pod with an empty data directory), so a killed restore restarts cleanly.

Consequences

Every restore path is the same well-tested one (restore into a sandbox), and point-in-time recovery works for all destinations. Copying from staging doubles the work for new/in-place restores, which is acceptable for the sizes involved.