Skip to content

Backups and the restore window

A target's history is a chain: full backups on a schedule, and between them a continuous change stream (Postgres WAL, MySQL binlog). Any second covered by a full backup and the changes after it can be restored. That span is the point-in-time window.

gantt
  dateFormat  YYYY-MM-DD HH:mm
  axisFormat  %a %d
  section Full backups
  daily full      :milestone, 2026-10-01 02:07, 0d
  daily full      :milestone, 2026-10-02 02:07, 0d
  daily full      :milestone, 2026-10-03 02:07, 0d
  daily full      :milestone, 2026-10-04 02:07, 0d
  section Change stream
  WAL / binlog segments :active, 2026-10-01 02:07, 2026-10-04 18:00

Full backups

Engine Mode Tool Restore
Postgres 15–17 physical (default) pg_basebackup -Ft -X fetch over the replication protocol base backup + WAL replay to the second
Postgres 15–17 logical pg_dump -Fc, one object per database; table exclusions and selective exports pg_restore, plus WAL replay through a physical chain when one exists
MySQL 8.4 logical mysqldump --single-transaction --source-data=2 (routines, events, triggers) load + binlog replay to the second

Every backup records an inventory: databases, tables, columns, indexes, row estimates and sizes, and the migration version when a migration table exists. That is what the console and keeper diff compare.

The change stream

Each target with point-in-time recovery has one streamer:

  • Postgres: Keeper's own WAL receiver on a physical replication slot (ADR 0005). The slot guarantees no WAL is lost while the streamer restarts.
  • MySQL: mysqlbinlog --read-from-remote-server --raw --stop-never from the binlog position of the last full.

Segments are uploaded as they complete, and the segment still being written is uploaded every partialInterval (default 60 s), so the last minute is restorable too. Each segment carries a summary: commits, inserts, updates and deletes per table, DDL, and optionally masked row previews.

If the stream breaks (a slot is lost, a binlog is purged), Keeper marks the chain broken, alerts (KeeperStreamChainBroken), starts a new chain and takes a full backup at once (reason chain-restart). The old window stays restorable up to the break.

Retention

Full backups are kept by tiers; segments are kept only as far back as the window needs; garbage collection removes everything else, never anything pinned.

Preset Window Daily Weekly Monthly Yearly
prod 7 days 14 8 12 5
dev 3 days 7 4 3 0
critical 14 days 30 12 24 7

The first committed full of a day is that day's daily; of an ISO week its weekly; and so on. One backup can carry several tiers; nothing is copied. Pins (keeper pin <id> --reason …) hold a backup regardless of tiers. GC writes its plan to the bucket before it deletes, deletes manifests before data, and cleans up orphaned and incomplete uploads after a day.

Storage format

<store prefix>/<org>/<project>/<namespace>/<target>/
  full/<backup-id>/manifest.json            # written last: the commit marker
  full/<backup-id>/data/<database>.zst.age  # dumps or the base backup tar
  full/<backup-id>/inventory.json.zst.age   # also readable with the catalog key
  full/<backup-id>/tags.json                # tiers, pins, verification results
  stream/…/<segment>.zst.age                # WAL / binlog segments and their summaries
  stream/position.json                      # the streamer's confirmed position
  • Compression: zstd at the policy's level, once, before encryption; on by default (ADR 0013). compression.algorithm: none stores data as is.
  • Encryption: every object is age-encrypted to the store's recipients on the client. The bucket never sees plaintext.
  • Integrity: every object's size and SHA-256 are in the manifest and verified on every read.
  • Commit: a backup exists only once manifest.json is written. A crash before that leaves orphans that GC removes, never a backup that looks complete.

Load control

Backups never compete blindly with production: per-host, per-target and global slots, a bandwidth limit, a priority, a jitter on schedules, and a load guard that postpones a full backup while the server has too many active connections, too much replication lag or long-running transactions. All of it is set in the BackupPolicy.