Skip to content

How Keeper works

Keeper is a Kubernetes operator. You declare where backups go, how often and for how long, and which databases; Keeper takes it from there.

flowchart LR
  subgraph Declared["Declared in git"]
    S[BackupStore<br/>bucket + keys]
    P[BackupPolicy<br/>schedule, retention, window]
    T[BackupTarget<br/>one database]
  end
  subgraph Created["Created by Keeper or by you"]
    B[Backup<br/>one full backup]
    R[Restore<br/>to a sandbox, new DB, download, in place]
    X[Sandbox<br/>a throwaway database]
  end
  T --> P --> S
  T -. schedule .-> B
  R -. sandbox .-> X
Resource Scope Who creates it What it is
BackupStore cluster you An S3 bucket, its credentials and the age recipients every object is encrypted to.
BackupPolicy namespace you (or the chart's presets) Schedule, point-in-time window, retention tiers, limits, load guard, verification, compression.
BackupTarget namespace, next to the database you One database server: engine, endpoint, credentials, policy, checks, masks.
Backup target namespace Keeper (schedule, backup-now, chain restart) One full backup run. Only the latest few are kept as objects; S3 holds them all.
Restore target namespace you, the console or the CLI A plan from a time or a backup ID to a destination.
Sandbox cluster a Restore to a sandbox A private database with a TTL, credentials and a query console.

The data path

flowchart LR
  db[(Postgres / MySQL)] -->|pg_basebackup, pg_dump, mysqldump| mover[mover Job]
  db -->|WAL receiver, mysqlbinlog| streamer[streamer]
  mover --> pipe[zstd → age → SHA-256 → parallel multipart]
  streamer --> pipe
  pipe --> s3[(S3 bucket<br/>manifest.json written last)]
  s3 --> restorer[restorer Job<br/>verify → decrypt → replay → checks]
  restorer --> dest[sandbox · new database · in place · download]
  • The controller only schedules. Movers and restorers are Jobs; each target with point-in-time recovery has one streamer Deployment. Bytes never pass through the controller or the API.
  • Nothing hangs. No mounts, iSCSI, FUSE or block devices; databases are reached only over the network. Every I/O call has a deadline, every stream a progress watchdog, and child tools run in their own process group so they are killed with it.
  • Crash-only. Any pod can be killed at any moment. State lives in CRDs and S3, and a backup exists only once its manifest is written last, so a half-written backup is never mistaken for a complete one.

The catalog and the console

keeper-api lists the manifests in each store at start-up and keeps an in-memory index: targets, backups, change segments and point-in-time windows. Inventories and masked previews are encrypted to an extra catalog key that can read them but cannot decrypt dumps or WAL (ADR 0009). The cache is disposable; S3 is the source of truth.