keeper
EN
GitHub

Backup & restore operator for Kubernetes · Postgres 15–17 · PostGIS · MySQL 8.4

Every second of your database, restorable.

Keeper backs up the databases you run in Kubernetes to any S3 bucket and streams every change as it happens. Pick a moment from the last few weeks and you get that exact database back in a throwaway sandbox, with checks already run, in less time than it takes to open a ticket.

Install with Helm Read the docs One binary, one chart. Your bucket, your keys.
Point-in-time window · team/app · policy prod window 7 days · chain unbroken
Restore to
2026-10-07 09:41:27 UTC
Base backup
20261007T020712Z-4c1e
Then replay
114
$ keeper restore team/app --at 2026-10-07T09:41:27Z --wait

Drag across the window. Amber ticks are the nightly full backups; the blue band is the change stream between them. An illustration of a prod policy: nightly fulls, 7-day window.

Measured with make test-perf and make test-ui on the k3d test harness (4 vCPU, in-cluster MinIO). The numbers are in the repository's docs/PERF.md.

See itone minute

One minute, start to restore.

Real commands and the real console against a demo catalog: list every database, see what changed overnight, restore it into a sandbox.

0:58 · with sound · 2.7 MB

The console

Keeper console overview: nine databases with their last full backup, point-in-time window, lag, stored size and health
Every database at a glance: last full backup, restore window, stream lag, size and health.
Keeper console target page: the restorable timeline, row and size trends, and the list of versions with verified badges
One database: its restorable timeline, its trends and every version, with verified restores marked.
Keeper console diff between two backups: a column added, a table emptied, rows and sizes per table
Two versions side by side: what was applied, which columns changed, rows and size per table.

The CLI

$ keeper targets
TARGET              ENGINE    ORG/PROJECT/ENV      LAST FULL             PITR WINDOW                                  STORED    HEALTH
analytics/events    postgres  acme/data/staging    2026-10-07 02:11:47Z  -                                            19.1GiB   ok
auth/users          postgres  acme/platform/prod   2026-10-07 02:11:06Z  2026-09-23 02:11:06Z → 2026-10-07 18:50:52Z  1.5GiB    ok
billing/invoices    mysql     acme/billing/prod    2026-10-07 02:09:03Z  2026-09-23 02:09:03Z → 2026-10-07 18:50:52Z  6.7GiB    ok
billing/ledger      postgres  acme/billing/prod    2026-10-07 02:08:22Z  2026-09-23 02:08:22Z → 2026-10-07 18:50:52Z  14.3GiB   ok
maps/geo            postgres  acme/maps/prod       2026-10-07 02:10:25Z  2026-09-23 02:10:25Z → 2026-10-07 18:50:52Z  23.3GiB   ok
shop-dev/orders-db  postgres  acme/shop/dev        2026-10-07 02:12:28Z  -                                            17.4MiB   ok
shop/catalog-db     postgres  acme/shop/prod       2026-10-07 02:07:41Z  2026-09-23 02:07:41Z → 2026-10-07 18:50:52Z  731.0MiB  ok
shop/orders-db      postgres  acme/shop/prod       2026-10-07 02:07:00Z  2026-09-23 02:07:00Z → 2026-10-07 18:50:52Z  9.6GiB    ok
web/wordpress       mysql     acme/marketing/prod  2026-10-07 02:09:44Z  2026-09-23 02:09:44Z → 2026-10-07 18:50:52Z  607.4MiB  ok
$ keeper diff 20261006T020700Z-bd79 20261007T020700Z-6da5
from  20261006T020700Z-bd79 (2026-10-06 02:07:00Z)
to    20261007T020700Z-6da5 (2026-10-07 02:07:00Z)

20261007 applied (was 20261006)
column orders.gift_note added
table audit_log +158,175 rows (+2%)
table order_items +62,923 rows (+1%)
table orders +15,655 rows (+1%)
table coupons −3,120 rows (−100%)
table customers +531 rows (+1%)

14 segments in between:
  00000001000000A100000000  18:51:32–19:20:32 · 800 commits · audit_log +3348 / ~1116 / −83 · order_items +1327 / ~442 / −33 · orders +330 / ~110 / −8 · customers +11 / ~3 / −0
  00000001000000A100000001  19:21:32–19:50:32 · 813 commits · audit_log +3352 / ~1116 / −83 · order_items +1328 / ~442 / −33 · orders +330 / ~110 / −8 · customers +13 / ~3 / −0
  00000001000000A100000002  19:51:32–20:20:32 · 826 commits · audit_log +3356 / ~1116 / −83 · order_items +1329 / ~442 / −33 · orders +330 / ~110 / −8 · customers +15 / ~3 / −0
$ keeper restore shop/orders-db --at 20261006T020700Z-bd79 --ttl 6h --reason "coupons truncated"
restore orders-db-sandbox-60414a15fc created (sandbox from 20261006T020700Z-bd79, 0 segments)
restore  shop/orders-db-sandbox-60414a15fc

$ keeper sandbox list
NAME      ENGINE    SOURCE            OWNER               PHASE         EXPIRES               IN-CLUSTER
sbx-7f3a  postgres  shop/orders-db    maria@acme.example  Ready         2026-10-08 00:37:32Z  keeper-sbx-sbx-7f3a.keeper-sandboxes.svc:5432
sbx-c210  mysql     billing/invoices  dev@acme.example    Provisioning  -

Output from the demo server in the repository: go run ./hack/demo, then point the CLI or a browser at it.

Restorefour destinations

Restore anywhere, without touching production.

A restore is a Kubernetes resource like any other. Ask for a time or a backup ID and pick where it goes. Keeper plans the base backup and the changes to replay, then runs your checks before it calls the restore a success.

SANDBOX

A throwaway copy

A private database with a TTL, credentials and a read-only query console. It cleans itself up.

--to sandbox --ttl 6h
NEW

A new database

Restored and checked in a staging sandbox, then copied to a new name. Existing objects are never overwritten.

--to new --database app_copy
IN PLACE

Roll back for real

Takes a safety backup first and asks you to type the target's name. Needs the restorer-admin role.

--to in-place --confirm app
DOWNLOAD

A dump file

A decrypted, compressed logical dump behind a link that expires in an hour. Audited.

--to download
Footprintsmall and steady

Small enough to forget it is there.

Keeper is one Go binary. The controller and the API only schedule and read; bytes move in short-lived Jobs that exit when they are done, so nothing heavy sits in your cluster between backups.

54MBone static binary: controller, API and console, movers, streamers, restorer and CLI
18MBcompressed download, no runtime to install, no sidecars
21MiBcontroller memory at rest, watching a cluster of databases
19MiBAPI and console memory at rest, with the catalog in memory

Working set reported by the kubelet in the k3d test cluster (12 databases) after the end-to-end and chaos suites; CPU at rest is a few millicores. Each database with point-in-time recovery also runs a streamer, about 70 MiB while it streams.

Fast by construction

Streaming pipelines with no temp files, parallel multipart uploads, multi-threaded zstd, and an in-memory catalog behind a console that answers in milliseconds.

Reliable by construction

Crash-only design: any pod can die at any moment. State lives in CRDs and S3, a backup exists only once its manifest is written last, and nothing waits forever.

Inspectbefore you restore

Know what changed between two versions.

Every backup carries an inventory of its tables, rows and sizes. Every change segment carries a summary of what it did. So before restoring anything, you can see which version still had the rows you lost.

  • Diffs between any two backups. Schema changes, row counts and size per table.
  • Change summaries per WAL or binlog segment. Inserts, updates, deletes and DDL per table, so a TRUNCATE at 22:48 is easy to find.
  • Masked previews. Sample rows with the columns you list masked, readable with a catalog key that cannot decrypt the data itself.
Featuresin the box

Built for the night something goes wrong.

Change streaming

Keeper's own WAL receiver on a replication slot for Postgres and mysqlbinlog for MySQL. The window reaches the present, even on a quiet database.

Verified, not hoped for

Scheduled verification restores run your SQL checks against a real restore and mark the backup verified. Their restore time is exported as a metric.

Encrypted before it leaves

Every object is encrypted client-side with age to your keys. The bucket never sees plaintext, and the console's catalog key cannot read data.

Compressed by default

zstd at a level you choose per policy, applied once in the pipeline. Dump tools run uncompressed, and the database server never spends CPU on it.

Never hangs

No mounts or block devices; databases only over the network. Every stream has a progress watchdog, and every tool runs in its own process group.

Crash-only

Kill any pod at any moment. A backup exists only once its manifest is written last, and the chaos suite kills Keeper mid-backup to prove it.

Gentle on production

Least-privilege, read-only users. Per-host slots, bandwidth limits, and a load guard that postpones a backup when the server is busy.

Console, CLI and API

One JSON API described by OpenAPI, with roles from Cloudflare Access. Every write and every query is audited.

GitOps native

Stores, policies, targets and restores are CRDs. Retention tiers, alerts and a Grafana dashboard ship with the chart.

Architecturedata path

Data moves in Jobs, never in the controller.

The controller only schedules. Movers and streamers carry the bytes through one streaming pipeline straight into your bucket, and the restorer brings them back.

Postgres · MySQL in your cluster mover Job full backups streamer WAL · binlog one streaming pipeline zstd → age → SHA-256 → 8 × multipart no temp files, back-pressure from S3 your S3 bucket manifest.json last *.zst.age restorer Job verify → decrypt → replay → checks sandbox · new database · in place · download with TTL, query console and audit
Comparedto a cron job

More than a nightly dump in a bucket.

When you need to…cron + pg_dump scriptKeeper
Get back to 09:41:27Last night's dump, up to 24 h lostAny second in the window
Prove a backup restoresFind out during the incidentScheduled verification restores with your checks
Find the version with the missing rowsRestore several and lookDiffs and change summaries in the console
Look without riskRestore over a shared staging databaseA private sandbox with a TTL
Keep the bucket blindServer-side encryption, if configuredClient-side age encryption to your keys
Survive a node dying mid-backupA half-written file that looks completeNothing counts until the manifest is written last
Know it is still workingSilenceAlerts for overdue backups, broken chains and lag
Installabout ten minutes

From zero to restorable.

  1. Install the chart

    CRDs, the controller, the API and console, RBAC, alerts and the dashboard.

    helm install keeper charts/keeper -n keeper-system --create-namespace \
      --set image.repository=registry.example.com/keeper/keeper --set image.tag=v0.1.0 \
      --set 'targetNamespaces={team-a}'
  2. Point it at a bucket and your keys

    Any S3-compatible store. Generate an age key with age-keygen and keep an offline copy.

    apiVersion: keeper.republic.global/v1alpha1
    kind: BackupStore
    metadata: { name: main }
    spec:
      s3:
        endpoint: https://s3.example.com
        region: us-east-1
        bucket: db-backups
        credentialsSecret: { namespace: keeper-system, name: keeper-s3 }
      encryption:
        ageRecipients: [age1…]   # public key only; the private key stays offline
  3. Declare a target next to the database

    Create the read-only users with the SQL in docs/sql/, then commit the target to git.

    apiVersion: keeper.republic.global/v1alpha1
    kind: BackupTarget
    metadata: { name: app, namespace: team-a }
    spec:
      engine: postgres
      endpoint: { host: postgres, port: 5432 }
      databases: [app]
      policy: prod             # nightly fulls, 7-day point-in-time window
      credentials:
        secretRef: { name: keeper-backup-postgres }
        replicationSecretRef: { name: keeper-repl-postgres }
  4. Restore something, on purpose

    The first full backup starts on its own. Then try the thing you hope you never need.

    keeper targets
    keeper restore team-a/app --at 2026-10-07T09:41:27Z --wait
    keeper query <sandbox> -e "select count(*) from orders"

The full guide, every CRD field and the REST API are in the documentation.

Test your restore before you need it.