How Keeper works¶
Keeper is a Kubernetes operator. You declare where backups go, how often and for how long, and which databases; Keeper takes it from there.
flowchart LR
subgraph Declared["Declared in git"]
S[BackupStore<br/>bucket + keys]
P[BackupPolicy<br/>schedule, retention, window]
T[BackupTarget<br/>one database]
end
subgraph Created["Created by Keeper or by you"]
B[Backup<br/>one full backup]
R[Restore<br/>to a sandbox, new DB, download, in place]
X[Sandbox<br/>a throwaway database]
end
T --> P --> S
T -. schedule .-> B
R -. sandbox .-> X
| Resource | Scope | Who creates it | What it is |
|---|---|---|---|
BackupStore |
cluster | you | An S3 bucket, its credentials and the age recipients every object is encrypted to. |
BackupPolicy |
namespace | you (or the chart's presets) | Schedule, point-in-time window, retention tiers, limits, load guard, verification, compression. |
BackupTarget |
namespace, next to the database | you | One database server: engine, endpoint, credentials, policy, checks, masks. |
Backup |
target namespace | Keeper (schedule, backup-now, chain restart) | One full backup run. Only the latest few are kept as objects; S3 holds them all. |
Restore |
target namespace | you, the console or the CLI | A plan from a time or a backup ID to a destination. |
Sandbox |
cluster | a Restore to a sandbox |
A private database with a TTL, credentials and a query console. |
The data path¶
flowchart LR
db[(Postgres / MySQL)] -->|pg_basebackup, pg_dump, mysqldump| mover[mover Job]
db -->|WAL receiver, mysqlbinlog| streamer[streamer]
mover --> pipe[zstd → age → SHA-256 → parallel multipart]
streamer --> pipe
pipe --> s3[(S3 bucket<br/>manifest.json written last)]
s3 --> restorer[restorer Job<br/>verify → decrypt → replay → checks]
restorer --> dest[sandbox · new database · in place · download]
- The controller only schedules. Movers and restorers are Jobs; each target with point-in-time recovery has one streamer Deployment. Bytes never pass through the controller or the API.
- Nothing hangs. No mounts, iSCSI, FUSE or block devices; databases are reached only over the network. Every I/O call has a deadline, every stream a progress watchdog, and child tools run in their own process group so they are killed with it.
- Crash-only. Any pod can be killed at any moment. State lives in CRDs and S3, and a backup exists only once its manifest is written last, so a half-written backup is never mistaken for a complete one.
The catalog and the console¶
keeper-api lists the manifests in each store at start-up and keeps an in-memory index: targets, backups, change
segments and point-in-time windows. Inventories and masked previews are encrypted to an extra catalog key that
can read them but cannot decrypt dumps or WAL (ADR 0009). The cache is
disposable; S3 is the source of truth.