Keeper — Database Backup & Restore Platform: Design¶
Oct 6, 2026 · @Eduardo
Purpose and context¶
Keeper is a small, in-house Kubernetes operator plus web console that backs up every database we run to S3, lets admins browse what each backup contains, and restores any backup (to any point in time) into a throwaway sandbox or a real database. We build only what we need, in Go, and it must never hang.
Why now. Today there are no database backups at all. All 9 databases live on local disks (local-path volumes) of a 3-node RKE2 cluster whose nodes are LXD system containers on one physical Dell PowerEdge R630. One hardware or disk failure loses everything. Backups must leave that machine.
Why build instead of adopt. We evaluated CloudNativePG, Percona operators / OpenEverest, K8up, Velero, WAL-G, pgBackRest and restic. They are good engines, but none gives one console across engines with data inspection, version diffs and sandbox restores, and Longhorn's history of freezing made us want full control over failure behaviour. Keeper reuses the engines' proven primitives (pg_basebackup, pg_receivewal, pg_dump, mysqldump, mysqlbinlog, later mongodump) and owns only orchestration, catalog, UI and safety.
Goals
- Back up Postgres and MySQL now (MongoDB and others later through the same engine interface).
- Full backups on schedules plus continuous change capture, so restores can target any second inside a window.
- Rotation by daily/weekly/monthly/yearly tiers, configurable per organization, project and environment.
- Never overload a server: staggered starts, concurrency slots, load checks, bandwidth limits.
- Inspect every backup (tables, row estimates, sizes, schema) and summarize what changed between versions.
- Restore into a sandbox database with a TTL that admins can query from the console or reach through a Cloudflare Tunnel; restore into new or existing databases.
- Prove backups work with scheduled automatic restore tests.
- Encrypt everything before it leaves the cluster.
- Never hang; any component can be killed at any time and self-heals.
Non-goals (for now)
- Block-level volume replication (Longhorn/DRBD style). Replication belongs at the database layer and needs a second physical failure domain first.
- Backing up Kubernetes resources (Flux and git already hold them).
- A public product; this is an internal admin tool.
Environment facts¶
Keeper's first home is one RKE2 cluster on one physical server; everything below was measured on 2026-10-05/06.
Cluster
- RKE2 v1.28.11, 3 nodes (
node110.212.66.84,node210.212.66.250,node310.212.66.190), Ubuntu 22.04, kernel 6.8. - Nodes are LXD system containers on one Dell PowerEdge R630 (same DMI, same kernel, identical boot time). No kernel modules can be loaded inside nodes; no iSCSI; Longhorn does not work here.
- 64 host threads visible per node; RAM limits 56 GiB (node1), 32 GiB (node2, node3).
- Root disks 137 GiB each, 77–83% used, mostly container image cache (kubelet image GC at 85%/80%). node1 also sees an 876 GiB host disk (45% used).
- Storage class
local-path(/opt/local-path-provisioner), no snapshot support. - Ingress class
republic-nginx(nginx.org controller, NodePort). Public traffic arrives via Cloudflare (proxied DNS to origin 107.155.66.195). - Flux v2.2.3 GitOps from
republic-global/republic-gitops(main, SSH deploy key). kube-prometheus-stack inmonitoring(Prometheus, Grafana, kube-state-metrics, node-exporter). cert-manager installed. No External Secrets Operator, no Kyverno, no Cloudflare tunnel yet.
Databases to protect (first wave)
| Namespace | Workload | Engine / image | Credentials secret | Real size |
|---|---|---|---|---|
| republic-prod | postgres | postgres:16.2 | republic-db-prod/POSTGRES_PASSWORD | 70 MiB |
| republic-billing | mysql-billing | mysql:8.4 | db-password, db-root-password | 264 MiB |
| solar-search-dev | postgres-dev-solar | postgis/postgis:16-3.4 | db-admin-solar-dev/POSTGRES_PASSWORD | 73 MiB |
| solar-search-dev | mysql-wp | mysql:8.4 | solar-search-db-secrets | 680 MiB |
| republic-dev | postgres-dev-public | postgres:16.2 | republic-db-dev-public | 66 MiB |
| republic-dev | postgres-doc-dev | postgres:16.2 | postgres-doc-dev | 182 MiB |
| republic-dev | postgres-infra | postgres:16.2 | republic-db-dev | 188 MiB |
| republic-dev | postgres-sonarqube | postgres:15.7 | republic-postgres-sonarqube/PASSWORD | 153 MiB |
| crowdgig-dev | postgres-dev-cg1 | postgres:16.2 | db-admin-dev/POSTGRES_PASSWORD | 68 MiB |
MongoDB is not deployed yet but must be supported later.
Object storage
- Telnyx Cloud Storage, S3-compatible, path-style addressing, signature v4. Endpoints per region:
https://us-west-1.telnyxcloudstorage.com,https://us-central-1.telnyxcloudstorage.com. - One Telnyx API key is both access key id and secret.
- Backups go to a dedicated bucket (proposed
republic-db-backups, us-central-1, see ADR 0001) with prefixes per cluster/org/project/target.
GitOps rules (republic-gitops)
- Never push to
main; work on a branch and open a PR; a person merges. Merges are squashed. - Run
git config core.hooksPath .githooksbefore the first push (pre-push hook blocksmain). - Layout:
apps/,apps-infra/(bases inapps-infra/base/*, overlays per org/env),infrastructure/(SOPS secrets, PGP key25F4CBA7…BD44, public key in.sops.pub.asc),charts/republic-server. - Secrets are SOPS-encrypted (
encrypted_regex: ^(data|stringData)$). New secrets are built withkubectl create secret --dry-run=client -o yamlthensops -e.
Access
- Cloudflare zone
republic.global(zone id2622eceb21f616dceba14bbeb61c0f77). The console will be exposed only through a Cloudflare Tunnel protected by Cloudflare Access. - Container registry
therepublic.azurecr.io(pull secrettherepublic-acr).
Key decisions and principles¶
Keeper is written in Go, keeps all state in Kubernetes and S3, moves data only in short-lived Jobs, and touches databases only over the network.
| Decision | Choice | Why |
|---|---|---|
| Build or adopt | Build a minimal operator + console | One console for all engines, inspection, diffs and sandboxes; full control of failure behaviour |
| Language | Go 1.23+ | Kubernetes-native (controller-runtime, client-go), static binary, context cancellation everywhere, fast streaming |
| UI | Server-rendered Go templates (templ or html/template) + HTMX + SSE, embedded in the binary | No Node runtime, millisecond pages, live progress |
| State | CRDs (desired state + status) and S3 manifests (finished backups) | Crash-only: nothing important in memory |
| Data path | Kubernetes Jobs and per-target streamer pods | The controller can never hang on a database or S3 |
| Database access | Network protocols only | No mounts, no block devices, no kernel paths that can freeze |
| Incremental | Full backup + continuous change stream (WAL, binlog, oplog) | Point-in-time restore to any second, no extra load spikes |
| Compression | zstd (klauspost/compress) | Fast, multi-core |
| Encryption | age (X25519) per object, client-side | Bucket leak exposes nothing; decryption key only fetched for restores |
| Object storage | Telnyx S3 via minio-go or aws-sdk-go-v2 (path-style) | Off the physical host, cheap |
| Config | Declarative CRDs in git (Flux), labels for org/project | Reviewed, versioned, next to the database |
| Admin access | Cloudflare Tunnel + Cloudflare Access (SSO), JWT verified by Keeper | No public IP, no VPN |
Never-hang rules (non-negotiable)
- No kernel-level I/O: Keeper never mounts volumes, never uses iSCSI, FUSE or block devices. Everything is a network stream that can be cancelled.
- The controller never moves data. All heavy work happens in Job pods with
activeDeadlineSeconds; streamers are separate Deployments. - Every network call has a
contextdeadline. Every data stream has a progress watchdog (default: abort if no bytes for 60 s). - Crash-only: any pod can be killed at any time. On start, components rebuild state from CRDs and S3 and continue.
- Commit by manifest: data objects first, manifest last. No manifest means the backup does not exist; GC removes leftovers.
- No blocking finalizers. Where a finalizer is unavoidable it has a timeout and an escape hatch.
- Idempotent steps: re-running any step with the same inputs produces the same result.
- Bounded everything: queues, retries (exponential backoff with cap), memory buffers, concurrency.
- Two controller replicas with leader election; liveness checks reconcile progress, not just process liveness.
- Every feature lands with tests in the harness, including a kill test.
Architecture¶
One Go binary (keeper) runs in different modes; the controller decides, Jobs and streamers move data, and the API serves the console from an in-memory catalog rebuilt from S3.
flowchart TB
admins["Admins (browser)"] --> tunnel["Cloudflare Tunnel + Access"]
subgraph ks["keeper-system namespace"]
controller["keeper-controller<br/>scheduler, slots, GC<br/>leader election"]
k8s["Kubernetes API<br/>CRDs: desired + status<br/>Leases for slots"]
api["keeper-api<br/>console, REST, SSE<br/>catalog rebuilt from S3"]
controller <--> k8s
api <--> k8s
end
tunnel --> api
controller -- "spawns Jobs and Deployments" --> mover["mover<br/>Job: full backup<br/>inspect + manifest"]
controller --> streamer["streamer<br/>WAL / binlog stream<br/>resumes from position"]
controller --> restorer["restorer<br/>Job: PITR restore<br/>checksums verified"]
controller --> sandbox["sandbox<br/>own namespace + TTL<br/>query, tunnel"]
restorer --> sandbox
mover & streamer & restorer -. "network streams only" .- dbs["Databases (app namespaces)<br/>Postgres, MySQL; MongoDB later"]
mover & streamer & restorer -. "network streams only" .- s3["Telnyx S3 (off this host)<br/>zstd + age, manifest written last"]
Keeper architecture · 1 binary, 5 runtime roles
The API and controller only read and write Kubernetes resources; data flows over the network from databases through mover, streamer and restorer pods to S3, and back into sandboxes.
| Component | Mode / kind | Responsibility |
|---|---|---|
| keeper-controller | keeper controller, Deployment, 2 replicas, leader election |
Reconciles CRDs; scheduler (cron + jitter); concurrency slots; creates mover Jobs, streamer Deployments, restore Jobs and sandboxes; retention and GC; verification schedules |
| keeper-api | keeper api, Deployment, 2 replicas (stateless) |
Web console + JSON API + SSE progress; reads the catalog; writes Restore/Sandbox/Backup CRs on admin actions; query console proxy; audit log |
| mover | keeper mover, Job per run |
Full backup: runs the engine dump tool, streams through zstd + age + parallel multipart upload, inspects the database, writes the manifest last |
| streamer | keeper streamer, Deployment per target (1 replica) |
Change capture: WAL receiver on a slot (ADR 0005) / mysqlbinlog --stop-never; ships segments, records positions in S3, produces change summaries |
| restorer | keeper restorer, Job per restore |
Fetches parts in parallel, verifies checksums, decrypts, decompresses, loads or replays to the target time, runs checks |
| sandbox | Namespace keeper-sbx-<id> with a DB pod, Service, random credentials, TTL |
Isolated database for inspection, testing and external apps |
| cloudflared | Deployment in keeper-system |
Tunnel for the console (HTTP) and for sandbox TCP routes |
| CLI | keeper locally |
Same API: backup now, list, restore, sandbox, pin |
Catalog. Each committed backup has a manifest in S3 (JSON) with an inventory summary (tables, rows, sizes, schema hash) and a pointer to its full inventory object; change segments have a summary. Inventories and masked previews are encrypted to the store recipients plus optional catalog recipients, so keeper-api can read them with a catalog identity that cannot decrypt data (ADR 0009). keeper-api lists manifests at start-up and keeps an index in memory; Keeper resources come from an informer-fed index, and target views are pre-computed per change (ADR 0010). The cache is disposable: S3 is the source of truth. CRDs keep only recent Backup objects (default: last 20 per target) to protect etcd.
Images. One image therepublic.azurecr.io/keeper/keeper:<tag> (distroless or slim Debian) containing the keeper binary plus engine tools: pg_dump, pg_basebackup, pg_receivewal, pg_waldump, pg_restore (Postgres client tools for majors 15, 16 and 17 side by side under /usr/lib/postgresql/\mysqldump, mysqlbinlog, mysql (8.4), and later mongodump. Sandbox and restore targets use the official engine images matching the source major version (for example postgis/postgis:16-3.4).
Data model (CRDs)¶
Six custom resources in API group keeper.republic.global/v1alpha1; adding a database means adding one BackupTarget.
| Kind | Scope | Purpose | Written by |
|---|---|---|---|
| BackupStore | Cluster | Bucket, endpoint, region, prefix, credentials secret ref, age recipients | Git |
| BackupPolicy | Namespace (the target's, then Keeper's for shared presets; ADR 0008) | Schedules, PITR window, retention tiers, limits, priority, verification, compression | Git |
| BackupTarget | Namespace (next to the DB) | Engine, endpoint, credentials, policy ref, scope (selective), labels org/project/env | Git |
| Backup | Namespace | One run (full or manual); status, sizes, manifest key, inventory summary | Controller / API |
| Restore | Namespace | Restore request: source backup or point in time, destination (sandbox, new, in-place), checks | API / CLI / controller (verification) |
| Sandbox | Cluster | Ephemeral database: engine, image, size, TTL, exposure, owner | Restore / API |
apiVersion: keeper.republic.global/v1alpha1
kind: BackupStore
metadata: { name: telnyx-us-central-1 }
spec:
s3:
endpoint: https://us-central-1.telnyxcloudstorage.com
region: us-central-1
bucket: republic-db-backups
pathStyle: true
credentialsSecret: { namespace: keeper-system, name: keeper-s3 } # keys: accessKeyId, secretAccessKey
prefix: rke2-main/
encryption:
ageRecipients: ["age1..."] # public keys; private key only in Key Vault / offline
identitySecret: { namespace: keeper-system, name: keeper-age-identity, optional: true }
---
apiVersion: keeper.republic.global/v1alpha1
kind: BackupPolicy
metadata: { name: prod }
spec:
store: telnyx-us-central-1
full: { schedule: "0 2 * * *", timeZone: America/Edmonton, jitter: 20m }
pitr: { enabled: true, window: 7d }
retain: { daily: 14, weekly: 8, monthly: 12, yearly: 5 }
limits: { perHost: 2, perTarget: 1, global: 4, bandwidth: 200Mi, priority: 100 }
loadGuard: { maxActiveConnections: 200, maxReplicationLagSeconds: 60, postponeFor: 15m, maxPostpones: 8 }
verify: { schedule: "0 4 * * 0", checks: [inventory-matches, row-estimates-within-10pct] }
---
apiVersion: keeper.republic.global/v1alpha1
kind: BackupTarget
metadata:
name: postgres
namespace: republic-prod
labels: { keeper.republic.global/org: republic, keeper.republic.global/project: platform, keeper.republic.global/env: prod }
spec:
engine: postgres # postgres | mysql | mongodb (later)
endpoint: { host: postgres.republic-prod.svc, port: 5432 }
databases: ["*"] # or explicit names
credentials:
secretRef: { name: keeper-backup-postgres } # keys: username, password (read-only backup user)
replicationSecretRef: { name: keeper-repl-postgres } # for WAL streaming
policy: prod
scope: {} # selective options, see below
checks:
- name: users-not-empty
sql: "select count(*) > 0 from public.users"
---
# SonarQube style: keep only what cannot be rebuilt
spec:
engine: postgres
policy: dev
pitr: { enabled: false }
scope:
mode: logical # pg_dump instead of pg_basebackup
excludeData: [file_sources, live_measures, project_measures, ce_activity, ce_scanner_context, ce_task_input, snapshots, duplications_index] # verify names against the live SonarQube schema; keep issues + issue_changes if triage must survive
queries: # optional exports, stored as CSV next to the dump
- name: user-settings
sql: "select * from properties where user_uuid is not null"
Status fields (all CRDs) carry phase, conditions[] (Kubernetes style), observedGeneration, timestamps, and for Backup: manifestKey, bytes, compressedBytes, durationSeconds, startLSN/endLSN or binlog file/position, inventorySummary (table count, total rows estimate), tier (daily/weekly/monthly/yearly), pinned, verified.
Org / project / environment. Labels on targets drive console filters, S3 prefixes (<prefix>/<org>/<project>/<namespace>/<target>/) and RBAC in the console. Policies can be cluster-wide (shared presets: prod, dev, critical) or namespaced (project-specific).
Backup engines and storage format¶
Every engine provides a full backup plus a continuous change stream; together they restore to any second inside the PITR window, and the 2–4 hour incremental idea is replaced by continuous capture.
sequenceDiagram
participant C as controller
participant L as Lease slots
participant M as mover Job
participant D as database
participant S as Telnyx S3
C->>L: acquire slots
L-->>C: granted or queued
C->>M: create Job, deadline 2 h
M->>D: probe load
M->>D: start dump stream
D-->>M: bytes, 60 s watchdog
M->>S: zstd + age, parallel parts
M->>D: inventory queries
M->>S: manifest last = commit
M-->>C: Job succeeded
C->>L: release slots
Scheduled full backup · sequence
If the mover dies or stalls at any step before the manifest, nothing is committed: the Job is retried and GC later removes the partial objects.
| Engine | Full (physical) | Full (logical, selective) | Change stream | Point-in-time restore |
|---|---|---|---|---|
| Postgres 15/16/17 | pg_basebackup -D - -Ft -X fetch over the replication protocol (ADR 0006) |
pg_dump -Fc per database, scope options (ADR 0006) |
Keeper's WAL receiver on a physical replication slot (ADR 0005) | Base backup + restore_command fetching WAL, recovery_target_time |
| MySQL 8.4 | mysqldump --single-transaction --source-data=2 --routines --events --triggers (logical; physical later if needed) |
same dump with table filters | mysqlbinlog --read-from-remote-server --raw --stop-never from the recorded position (requires binlog_format=ROW, GTID recommended) |
Load dump, replay binlog with --stop-datetime |
| MongoDB (later) | mongodump --oplog --archive |
collection filters | Tail local.oplog.rs |
mongorestore --oplogReplay --oplogLimit |
Engine interface (Go). Each engine implements one interface so adding an engine never touches the controller:
type Engine interface {
Name() string
Probe(ctx context.Context, t Target) (ServerInfo, error) // version, size, settings, load metrics
Full(ctx context.Context, t Target, w ObjectWriter) (FullResult, error)
Stream(ctx context.Context, t Target, from Position, w SegmentWriter) error // blocks until ctx done
Inspect(ctx context.Context, t Target) (Inventory, error) // tables, rows, sizes, schema
SummarizeSegment(ctx context.Context, seg Segment) (ChangeSummary, error)
Restore(ctx context.Context, plan RestorePlan, dst Destination, r ObjectReader) error
Check(ctx context.Context, dst Destination, checks []Check) ([]CheckResult, error)
}
Streaming pipeline (no temp files). Tool stdout → bounded buffer → zstd (multi-threaded) → age encrypt → SHA-256 → S3 multipart upload with N parallel parts (default 8 × 16 MiB). Back-pressure is natural: if S3 slows, the dump tool blocks on its pipe. Each database dump (pg_dump -Fc) is its own object, so databases restore independently (ADR 0006).
Compression. On by default: zstd at level default, applied once in the pipeline before encryption. Dump tools run uncompressed, and the database server never compresses. BackupPolicy.spec.compression sets the algorithm (zstd or none), the level (fastest, default, better or best) and the threads. Readers detect zstd from the data, so changing the setting never affects restores of older backups (ADR 0013).
S3 layout
<prefix>/<org>/<project>/<namespace>/<target>/
full/<backupID>/data/<part-or-file>.zst.age # data objects
full/<backupID>/inventory.json.zst.age # tables, rows, sizes, schema (also summarized in manifest)
full/<backupID>/manifest.json # written LAST = commit marker (not encrypted, no secrets)
stream/<timeline-or-server>/<segment>.zst.age # WAL / binlog segments
stream/<timeline-or-server>/<segment>.summary.json # change summary per segment
stream/position.json # last confirmed position (atomic overwrite)
exports/<backupID>/<query-name>.csv.zst.age # selective query exports
Manifest (commit protocol). manifest.json lists every object with size and SHA-256, the engine, server version, start/end positions (LSN or binlog file:pos and GTID set), timestamps, tool versions, encryption recipients, inventory summary and tier tags. It is written with a conditional PUT when the store supports it, otherwise PUT then read-back verify. A backup is visible to restore and catalog only if its manifest exists and every listed object exists with the right size. Tier tags (weekly, monthly, yearly) are metadata in the manifest plus a tiny tags.json sidecar, not copies.
Change-stream safety.
- Postgres: a dedicated physical slot
keeper_<target>; Keeper requiresmax_slot_wal_keep_sizeto be set (for example 2 GB) so a stopped streamer can never fill the database disk; lag is exported as a metric and alerted. - MySQL: the streamer resumes from the last confirmed file:position; if the server purged the binlog it needs, Keeper marks the PITR chain broken from that time, alerts, and triggers an immediate full backup to start a new chain.
- Each committed full backup records the stream position it is consistent with, so restore plans are computed, never guessed.
Snapshot inspection and change previews¶
Every backup records what the database looked like, and every change segment records what changed, so the console can show the database's history like versions in git.
Inventory per full backup (fast queries only, no table scans). Collected by the mover over the same connection, inside the same snapshot when possible:
| Item | Postgres source | MySQL source |
|---|---|---|
| Databases, schemas, tables, views, extensions | pg_namespace, pg_class, pg_extension |
information_schema.SCHEMATA, TABLES |
| Row estimate per table | pg_class.reltuples (+ pg_stat_user_tables.n_live_tup) |
information_schema.TABLES.TABLE_ROWS |
| Exact counts (opt-in, small tables only, with timeout) | count(*) under statement_timeout |
count(*) under max_execution_time |
| Table / index / toast size | pg_total_relation_size, pg_indexes_size |
DATA_LENGTH, INDEX_LENGTH |
| Columns, types, PK, indexes, FKs | information_schema.columns, pg_index, pg_constraint |
information_schema.COLUMNS, STATISTICS, KEY_COLUMN_USAGE |
| Activity since stats reset | n_tup_ins/upd/del, last_autovacuum |
— |
| Schema fingerprint | SHA-256 of normalized DDL per table and overall | same |
| Migration version (if found) | flyway_schema_history, schema_migrations, alembic_version |
same tables |
| Server facts | version, settings that matter, DB size | version, binlog_format, gtid_mode |
Version diff between any two backups. Computed from inventories, instant in the console: tables added/dropped/renamed (by fingerprint similarity), columns added/removed/type changed, index changes, row-estimate deltas and growth rates, size deltas, migration version change. Displayed as a timeline with "what happened" lines, for example: V2.9.5 applied; table solar_design.jobs +1,240 rows (+18%); column public_quotes.limit_reached added.
Change summary per stream segment.
- MySQL (ROW binlog): parse events with
mysqlbinlog --base64-output=decode-rows -v(or a Go binlog parser such as go-mysql) → per table insert/update/delete counts, DDL statements, transaction count, first/last timestamp, and a preview of up to N rows per table (default 5) with column values truncated and sensitive columns masked. - Postgres (physical WAL):
pg_waldump --stats=recordand record parsing → per relation (resolved to table names via the inventory's relfilenode map) insert/update/delete/truncate counts, DDL markers, commit count, time range. No row values (physical WAL has none in usable form). - Postgres row previews (optional per target): a second, logical slot (
pgoutput, built in) streamed by the same streamer, sampled to N rows per table per segment. Off by default because it retains extra WAL on the database. - MongoDB (later): oplog entries already carry documents; same summary shape.
Privacy. Previews pass through a masking rule set per target (mask: [email, phone, password*, token*], regex by column name) before storage, and previews can be disabled entirely per target. Previews are encrypted like data.
Explanations. Each version and segment gets a one-line generated description from the summary (deterministic template, no AI dependency), for example 14:00–14:05 · 312 commits · orders +120 / ~45 / −2 · DDL: CREATE INDEX ix_orders_date.
Scheduling and load control¶
Backups never start all at once and never push a server over its limits: schedules are jittered, every run needs a slot, and a load check can postpone it.
- Schedule + deterministic jitter. Cron (with
timeZone) per policy, plus a per-target offset = hash(target UID) modjitter. Same target, same minute every day; different targets spread out. Avoid :00/:15/:30/:45 where existing CronJobs fire. - Slots. Before creating a mover Job, the controller acquires slots in three pools:
perTarget(default 1),perHost(keyed by the database server, default 2) andglobal(default 4). Slots are KubernetesLeaseobjects with holder identity and expiry, so they survive controller restarts and expire if a holder dies. Waiting runs queue by priority, then age. - Load guard. The mover probes the server first (active connections, long transactions, replication lag, disk free on the server if exposed). Over thresholds → postpone (default 15 min, max 8 times), then run anyway and record a warning.
- Bandwidth. A token bucket per mover limits upload rate (
limits.bandwidth). Restores have their own limit. - Resource limits. Mover and restorer pods have CPU/memory requests and limits; zstd threads = CPU limit;
nice/ionicefor the dump tool inside the pod. - Priority. Policies carry a numeric priority (prod 100, dev 10). Verification and sandbox restores run below scheduled backups.
- Streams are not scheduled. Streamers run continuously at low, steady cost; they are restarted with backoff by Kubernetes.
- Missed runs. If the controller was down during a schedule, it runs the missed backup once on start (not once per missed slot), within a
startingDeadline(default 6 h). - Multiple servers and clusters. Pools key on the database server identity (host:port), so many databases on one server share its slots. Each cluster runs its own Keeper; stores can be shared read-only across clusters.
Retention and garbage collection¶
Full backups are kept by daily/weekly/monthly/yearly tiers; change segments are kept only as far back as the PITR window needs; GC removes everything else, never anything pinned.
| Preset | PITR window | Daily | Weekly | Monthly | Yearly |
|---|---|---|---|---|---|
| prod | 7 days | 14 | 8 | 12 | 5 |
| dev | 3 days | 7 | 4 | 3 | 0 |
| critical (example) | 14 days | 30 | 12 | 24 | 7 |
- Tier assignment is a pure function of the backup's timestamp in the policy's time zone: the first committed full of a day is that day's daily; of an ISO week its weekly; of a month its monthly; of a year its yearly. One backup can carry several tiers. No copies are made.
- Keep set = union of: newest N per tier, every full needed as the base of the PITR window, all pinned backups, backups referenced by active Restores or Sandboxes.
- Segments are kept from the base backup of the oldest point in the PITR window onward.
- GC run (daily, in the controller): computes the keep set from manifests, deletes manifests first then data objects (so a half-deleted backup is never restorable), then deletes orphans: objects with no manifest older than 24 h (failed or killed uploads), and incomplete multipart uploads older than 24 h.
- Dry run and audit. GC writes a plan object (
gc/<timestamp>.json) before deleting; the console shows what will expire next. - Pins and legal hold.
pinned: truewith an optional reason and expiry; pinned backups are never deleted by GC.
Restore, sandboxes and query console¶
Any backup, or any second inside the PITR window, can be restored into a sandbox, a new database or (with safeguards) in place; sandboxes are queryable from the console and reachable through the tunnel.
sequenceDiagram
participant A as admin
participant API as keeper-api
participant C as controller
participant R as restorer Job
participant SB as sandbox DB
participant T as tunnel
A->>API: pick DB + time T
API->>C: Sandbox CR
C->>SB: namespace, DB, credentials
C->>R: start restore
R->>SB: full + WAL to T
SB-->>R: checks pass
R-->>C: Ready
C->>T: add route sbx-id
API-->>A: creds + tunnel cmd
A->>API: query console
API->>SB: read-only SQL, 30 s limit
C->>SB: TTL: delete namespace
Restore to a sandbox from the console · sequence
The restorer fetches the full backup and the change segments up to T from S3 in parallel; the admin sees live progress over SSE throughout.
Restore modes
| Mode | Destination | Safeguards |
|---|---|---|
| Sandbox (default) | New namespace keeper-sbx-<id>, fresh DB pod from the matching engine image, random credentials, TTL |
Isolated; NetworkPolicy allows only keeper-api, cloudflared and labelled client pods; auto-deleted at TTL |
| New database | A new database name on an existing server, or a new Deployment in a chosen namespace | Restored into a staging sandbox, then copied (ADR 0007); never overwrites existing objects |
| In place | The original target | Typed confirmation of the target name, role restorer-admin, automatic safety backup first, writes paused by the operator (Keeper does not stop apps) |
| Download | Time-limited presigned URL to a decrypted, recompressed dump produced by a Job | Role restorer, audited, expires in 1 h |
| Table-level (later) | Specific tables into sandbox or new database | Postgres: pg_restore -t; MySQL: filtered dump |
Restore plan. Given a target and a time T, the planner picks the newest committed full whose consistent position is ≤ T, then the ordered list of segments up to T, and verifies continuity (no gaps in LSN or binlog positions). If the chain is broken, the plan says so and offers the nearest restorable times.
Sandbox lifecycle. Pending → Provisioning → Restoring → Checking → Ready → Expiring → Deleted. Default TTL 24 h, extendable up to a policy maximum (7 days). Resources sized from the backup (for example 2× data size for the volume, emptyDir with sizeLimit by default so nothing lingers on node disks; optional local-path PVC for large restores). Credentials are shown once and stored only in the sandbox namespace secret.
Exposure.
- In-cluster apps: the sandbox Service DNS name; client pods need label
keeper.republic.global/sandbox-client: <id>. - External apps: a Cloudflare Tunnel TCP route
sbx-<id>.db.republic.globalprotected by Cloudflare Access. The client runscloudflared access tcp --hostname sbx-<id>.db.republic.global --url localhost:5432and connects to localhost. Routes are added and removed by the controller through the Cloudflare API.
Query console. In the web console, a SQL editor against a sandbox only (never against live databases): read-only transaction by default, statement_timeout / max_execution_time 30 s, result limit 1,000 rows, results streamed, CSV export, query history per user, every query audited. A table browser shows columns, row estimates and a 50-row preview per table, with masking rules applied.
Automatic verification. On the policy's verify schedule, the controller creates a Restore into a sandbox from the latest backup (and a random PITR point when streams exist), runs inventory comparison (tables and schema fingerprint must match, row estimates within tolerance) plus target-defined SQL checks, records duration (RTO) and marks the backup verified or alerts, then deletes the sandbox.
Web console, API and access¶
The console feels like a cloud provider's database backup page, served by keeper-api from memory, reachable only through Cloudflare Tunnel + Access.
Pages
| Page | Shows | Actions |
|---|---|---|
| Overview | Every target: engine, org/project/env, last full, PITR window (with gaps in red), last verification, stream lag, storage used, health | Filter, search, back up now |
| Target | Timeline of fulls (tiers, pins, verified badges) and stream coverage; growth charts (rows, size); schedule and next run | Back up now, restore, sandbox at time T, pin, pause/resume |
| Backup (version) | Manifest, sizes, duration, tool versions, inventory (tables, row estimates, sizes, schema), checks | Diff with another version, restore, sandbox, download, pin |
| Diff | Schema and row/size changes between two versions; segments in between with change summaries | Jump to a segment, sandbox at a time |
| Changes | Segment list with per-table insert/update/delete counts, DDL, previews (masked) | Sandbox just before or after a segment |
| Sandboxes | Active sandboxes, owner, TTL, size, connection info, tunnel route | Query console, table browser, extend, delete |
| Query console | SQL editor on a sandbox, results grid, history | Run, export CSV |
| Activity | Jobs and restores with live progress (SSE) and logs | Cancel, retry |
| Settings | Stores, policies, targets (read-only from git), alerts | Links to the git files |
| Audit | Who did what, when, from where | Export |
API. JSON over HTTP under /api/v1 mirrors the pages (/targets, /targets/{ns}/{name}/backups, /backups/{id}, /backups/{id}/diff/{other}, /segments, /restores, /sandboxes, /sandboxes/{id}/query, /events SSE). The CLI and the console use the same API. The OpenAPI document api/openapi/v1.yaml defines it, and the server serves it at /api/v1/openapi.yaml. The DTOs and the CLI's client are generated from that document, and every request is validated against it; each operation's role is set in the document (ADR 0012). All writes create or patch CRDs; the API never talks to databases directly except the query proxy to sandboxes.
CLI (keeper). The same binary used from an admin's laptop or CI, talking only to the API (never to Kubernetes or databases directly), so it gets the same auth, roles and audit as the console. Output is a table by default, -o json|yaml for scripts; exit codes are non-zero on failure; long operations stream progress and --wait blocks until done (with --timeout).
| Command | Does |
|---|---|
keeper login [--server https://keeper.republic.global] |
Gets a Cloudflare Access token (cloudflared access login flow or a service token from env KEEPER_ACCESS_CLIENT_ID/SECRET for CI), stores it in the OS keychain or ~/.config/keeper |
keeper targets [--org --project --env] |
List targets with last backup, PITR window, health |
keeper backups <ns>/<target> |
List versions with tiers, sizes, verified, pinned |
keeper show <backup-id> |
Manifest + inventory (tables, rows, sizes) |
keeper diff <backup-a> <backup-b> |
Schema and row/size changes between two versions |
keeper changes <ns>/<target> --since 2h |
Segment summaries (per-table counts, DDL) |
keeper backup now <ns>/<target> [--wait] |
Trigger a full backup |
keeper restore <ns>/<target> --at <time-or-backup-id> --to <sandbox/new/in-place> [--wait] |
Create a Restore; in-place also needs --confirm <target-name> |
keeper sandbox <list/show/extend/delete> [<id>] |
Manage sandboxes |
keeper sandbox connect <id> [--port 15432] |
Opens the tunnel locally (wraps cloudflared access tcp) and prints the connection string |
keeper query <sandbox-id> -e "select ..." |
Run a read-only query through the query proxy |
keeper pin <backup-id> [--reason], keeper unpin <backup-id> |
Retention hold on or off |
keeper verify <ns>/<target> [--wait] |
Run an on-demand verification restore |
keeper catalog import |
Admin/recovery: rebuild Backup CRs from S3 manifests (runs in-cluster with kubeconfig) |
Performance targets. p99 page render < 50 ms for 1,000 targets and 100,000 backups (in-memory index, pre-computed aggregates); first byte < 20 ms; SSE progress updates every 500 ms; no client-side framework; assets embedded and cached immutably.
Access.
- Cloudflare Tunnel (
cloudflaredDeployment inkeeper-system), hostname for examplekeeper.republic.global, no public IP or Ingress. - Cloudflare Access (SSO) in front; keeper-api verifies the
Cf-Access-Jwt-AssertionJWT (team domain + audience tag) on every request and rejects anything else, so bypassing the tunnel is useless. - Roles mapped from Access groups or emails in a ConfigMap:
viewer(read),operator(back up now, sandboxes, query console),restorer(restore to new, download),restorer-admin(in-place restore, pin/unpin, GC overrides). Optional scoping by org/project label. - Every write and every query is written to an audit log (stdout JSON + S3
audit/daily objects).
Security and encryption¶
Everything is encrypted before upload, databases are accessed with least-privilege users, and the decryption key is needed only for restores.
- Encryption. Every data object, inventory, segment, preview and export is encrypted with age (X25519) to the store's recipients. Manifests contain no data or secrets and stay plaintext for fast catalog loading. Recommended: two recipients, an operational key (in Key Vault, mounted only into restorer Jobs) and an offline break-glass key.
- Integrity. SHA-256 per object in the manifest; verified on every restore and by a weekly sampled re-read job.
- Credentials. One S3 key for the store (in
keeper-systemonly). Per target: a read-only backup user (Postgrespg_read_all_data+pg_monitor; MySQLSELECT, SHOW VIEW, TRIGGER, EVENT, LOCK TABLES, PROCESS, RELOAD) and, for streams, a replication user (PostgresREPLICATION; MySQLREPLICATION SLAVE, REPLICATION CLIENT). Target secrets live next to the database; Keeper's ServiceAccount reads only Secrets namedkeeper-*via a Role per target namespace (generated from git). - Secret sources. Phase 1: SOPS secrets in republic-gitops. Phase 2: External Secrets Operator with Azure Key Vault (ClusterSecretStore) so S3, age and DB backup credentials come from one place.
- RBAC (Kubernetes). The controller ServiceAccount can manage Keeper CRDs, Jobs and Deployments in
keeper-system, Leases, sandbox namespaces labelledkeeper.republic.global/sandbox, and readkeeper-*Secrets in target namespaces. No cluster-admin. - Sandboxes. Own namespace, NetworkPolicy default-deny, random credentials, TTL, masking optional at restore time (SQL masking scripts per target), never reachable without Cloudflare Access.
- Supply chain. Pinned base images and tool versions,
go mod verify,govulncheckin CI, SBOM, image signed (cosign) once CI exists. - Audit. All console writes, restores, downloads and queries are logged with user, time, target and outcome.
Reliability: never hang, self-heal¶
Any Keeper pod can be killed at any moment without losing or corrupting a backup; on restart everything resumes from CRDs and S3 within seconds.
flowchart TB
start["Pod starts"] --> read["Read last position from S3"]
read --> connect["Connect: slot / binlog position"]
connect --> receive["Receive next segment"]
receive --> upload["Compress, encrypt, upload"]
upload --> advance["Advance position in S3"]
advance -- "next segment" --> receive
connect -- "gap" --> gap["Gap: binlog purged or slot lost"]
gap --> broken["Mark PITR chain broken, alert"]
broken --> full["Full backup now, new chain"]
full -- "restart chain" --> read
advance -. "any step fails" .-> err["Error, timeout or pod killed<br/>Kubernetes restarts the pod"]
err -- "resume" --> read
Change-stream streamer · self-healing loop
The position is advanced only after a segment is safely in S3, so a kill at any point repeats at most one segment and never skips one.
| Failure | What happens | Recovery |
|---|---|---|
| Controller killed mid-reconcile | Standby takes the lease in < 15 s | Reconcile is idempotent; status phases resume |
| Mover killed mid-upload | No manifest written | Job retried (backoffLimit 2) with a new backup ID; GC removes orphans and aborts multipart uploads |
| Database stalls during dump | Watchdog sees no bytes for 60 s | Mover exits non-zero; Job retried after backoff; alert after final failure |
| S3 slow or down | Uploads retry with backoff; dump tool blocks on its pipe | Job deadline (default 2 h) kills it; next schedule runs |
| Streamer killed | Kubernetes restarts the pod | Resumes from position.json; Postgres slot guarantees no gap |
| Quiet database | No WAL/binlog for hours; server sends no keepalives | Streamer asks the server for a reply every 30 s (liveness), records caughtUpAt so the window reaches the present (ADR 0011) |
| PITR turned off or target deleted | Streamer removed | Cleanup Job drops the replication slot behind a non-blocking finalizer (ADR 0011) |
| Stream gap (binlog purged, slot lost) | Chain marked broken from time T | Alert + immediate full backup starts a new chain |
| Restore killed | Restore CR stays in its phase | Restorer Job is recreated; restore into a sandbox restarts cleanly (sandbox reset) |
| keeper-api killed | Other replica serves | Catalog rebuilt from S3 on start (seconds) |
| etcd/CRD loss | Backups still in S3 | keeper catalog import recreates Backup CRs from manifests |
| Sandbox namespace stuck terminating | Controller gives up after timeout | Removes its own finalizers only; reports in console |
Rules in code.
- No unbounded waits: every
selecthas actx.Done()branch; every HTTP/S3/DB client has timeouts; exec of tools usesexec.CommandContextwith process-group kill. - Kill the whole process group of child tools on cancel (
Setpgid+syscall.Kill(-pgid, SIGKILL)after a grace period). - Status updates use optimistic concurrency and retries; no in-memory queues that are the only copy of work.
- Leases for slots and leadership have short durations (15 s renew, 30 s expiry).
- Health:
/livezfails if the reconcile loop has not completed a cycle in 2 minutes;/readyzreflects informer sync and S3 reachability. - Panics are recovered per reconcile, logged, and turned into retries with backoff.
Observability¶
Keeper exports Prometheus metrics and ships PrometheusRules so a missing or failing backup always pages someone.
Metrics (prefix keeper_): last_success_timestamp_seconds{target,kind}, backup_duration_seconds, backup_bytes{stage=raw|compressed}, stream_lag_seconds{target}, stream_lag_bytes, pitr_window_seconds{target}, chain_broken{target}, verify_last_success_timestamp_seconds, verify_rto_seconds, restore_duration_seconds, slots_in_use{pool}, queue_depth, gc_deleted_objects_total, store_bytes{target}, sandboxes_active, plus Go runtime and controller-runtime metrics. ServiceMonitor included.
Alerts (PrometheusRule)
- BackupOverdue: no successful full for a target in 26 h (prod) / 50 h (dev).
- StreamLagHigh: lag > 5 min for 10 min; StreamChainBroken:
chain_broken == 1. - VerificationFailed or overdue (> 8 days).
- SlotWALRetentionHigh: Postgres slot retained WAL > 75% of
max_slot_wal_keep_size. - StoreUnreachable, GCFailing, SandboxExpiryFailing.
Logs and events. Structured JSON logs (slog) with target, backup ID and trace ID; Kubernetes Events on targets for every run outcome; Grafana dashboard JSON shipped with the Helm chart.
Testing harness¶
Every feature is proven by an automated test that restores data and compares it, including tests that kill components mid-flight; make test-all runs everything locally in a disposable cluster.
| Layer | Tool | What it proves | Command |
|---|---|---|---|
| Unit | go test, table tests, fakes |
Planner, tiering, retention, slot logic, manifest parsing, diff, masking, cron + jitter | make test-unit |
| Component | go test + testcontainers-go (Postgres 15/16/17, PostGIS 16-3.4, MySQL 8.4, MinIO) |
Each engine: full, stream, inspect, summarize, restore, PITR to exact time | make test-engines |
| Controller | controller-runtime envtest |
Reconcile loops, status phases, leases, finalizer timeouts | make test-controller |
| End to end | k3d (or kind) cluster + MinIO + DB workloads + Keeper Helm chart + fake cloudflared | Full user journeys through the API: schedule, backup, stream, sandbox restore, query, diff, GC, verification | make test-e2e |
| Chaos | e2e + a chaos driver (Go) | Kill controller, mover, streamer, restorer, API at random points; pause DB (SIGSTOP); throttle/blackhole S3 (toxiproxy) | make test-chaos |
| Performance | e2e + data generator | Throughput MB/s backup and restore, console p99 with 1,000 targets / 100k backups, memory ceilings | make test-perf |
| UI | Playwright (Go) or chromedp | Pages render, actions work, no console errors | make test-ui |
Data generator and golden checks. A Go tool creates schemas with all common types (including PostGIS geometry, JSONB, BLOB, unicode), loads a configurable volume, then runs a write workload while streams are active, recording (timestamp, checksum of every table) checkpoints. A PITR test restores to a random checkpoint time and must reproduce exactly that checksum set. Checksums use ordered md5(string_agg(...)) per table in Postgres and CHECKSUM TABLE + ordered hash in MySQL.
Invariants checked after every chaos run: no committed manifest references a missing or wrong-size object; every committed backup restores and matches its inventory; no orphan older than the GC horizon remains after GC; no slot or Lease leaked; no pod stuck Terminating > 60 s; controller and API recovered within 30 s.
Harness entrypoint. hack/harness.sh (or go run ./test/harness) creates the k3d cluster with a local registry, builds and loads images, installs MinIO and the fixtures, runs the selected suites, collects logs and a JUnit + HTML report into artifacts/, and tears down. It must run unattended in the Claude Code cloud workspace and in GitHub Actions. Default make test-all time budget: under 20 minutes; chaos and perf suites can run nightly.
Rule: a change is done only when make test-all is green; failing tests are fixed, never skipped.
Repository layout, conventions and stack¶
One Go module in github.com/republic-global/keeper, one binary with subcommands, one Helm chart, one harness.
keeper/
CLAUDE.md # rules for Claude sessions + link to this design doc
README.md # what it is, how to run, how to test
go.mod
cmd/keeper/ # main: controller | api | mover | streamer | restorer | cli subcommands
api/v1alpha1/ # CRD Go types + deepcopy (kubebuilder markers)
api/openapi/v1.yaml # the JSON API (OpenAPI 3.0); internal/apiv1 is generated from it
internal/controller/ # reconcilers: store, policy, target, backup, restore, sandbox, gc, verify
internal/scheduler/ # cron, jitter, slots (Leases), priority queue
internal/engine/ # Engine interface + registry
internal/engine/postgres/ # full, stream, inspect, summarize, restore, check
internal/engine/mysql/
internal/pipeline/ # zstd, age, sha256, multipart writer/reader, watchdog, rate limit
internal/store/ # S3 client, layout, manifests, commit, listing
internal/catalog/ # in-memory index, inventory, diffs, explanations
internal/planner/ # restore plans and continuity checks
internal/retention/ # tiers, keep set, GC
internal/sandbox/ # provisioning, TTL, network policy, tunnel routes
internal/web/ # handlers, templates (templ), HTMX, SSE, auth (Access JWT), RBAC, audit
internal/query/ # sandbox query proxy with limits
internal/cloudflare/ # tunnel route management
internal/proc/ # exec with process groups, deadlines, watchdog
internal/metrics/
config/crd/ # generated CRD YAML
charts/keeper/ # Helm chart (controller, api, RBAC, CRDs, ServiceMonitor, PrometheusRule, cloudflared)
images/keeper/Dockerfile # binary + engine tools (pg 15/16/17 clients, mysql 8.4 clients)
test/e2e/ test/chaos/ test/perf/ test/datagen/ test/fixtures/
hack/harness.sh hack/k3d.yaml
docs/ # user docs (MkDocs Material, mkdocs.yml), ADRs, runbooks; docs/reference is generated
site/ # website: i18n landing page, Worker (R2 media), demo video sources
hack/demo/ # console + API over a synthetic catalog (screenshots, trying Keeper without a cluster)
.github/workflows/ # lint, unit, engines, e2e (nightly chaos/perf), image build and push
Makefile
Stack. Go 1.23+, controller-runtime + kubebuilder markers (controller-gen), client-go, minio-go or aws-sdk-go-v2 for S3, klauspost/compress (zstd), filippo.io/age, robfig/cron v3, oapi-codegen + kin-openapi (API document, ADR 0012), a-h/templ + HTMX + SSE, log/slog, prometheus/client_golang, testcontainers-go, envtest, k3d, MinIO, toxiproxy, golangci-lint, govulncheck.
Conventions.
- Small packages with clear interfaces; engines never import controllers.
- Every exported function that does I/O takes
context.Contextfirst. - Errors wrapped with
%w; no panics in library code. - Config only through CRDs, flags and env; no hidden files.
- Conventional commits; PRs small and green;
mainprotected after the bootstrap commit; Claude sessions never push tomain(same pre-push hook as republic-gitops). - Image tags
main-<unix time>-<short sha>(same pattern as solar-search-quote, so Flux image automation works).
Roadmap, definition of done and rollout¶
Build in eight milestones, each ending green in the harness; then deploy to the real cluster for two databases through one republic-gitops PR, and only after that roll out to the rest.
- M0 — Skeleton. Repo, Makefile, CI, CRD types, controller and API skeletons, Helm chart, image with engine tools, harness that boots k3d + MinIO + Postgres + MySQL. Gate:
make test-allgreen on an empty feature set. - M1 — Postgres full backups. Pipeline (zstd, age, sha256, multipart), manifest commit, inventory, scheduler with jitter, slots, load guard, retention + GC. Gate: e2e backup → restore to sandbox → checksums match; kill tests pass.
- M2 — Postgres streaming + PITR. Streamer with slot, segment summaries via pg_waldump, restore planner, PITR restore. Gate: restore to 20 random checkpoint times all exact.
- M3 — MySQL. Full dump, binlog streaming, summaries with masked previews, PITR. Gate: same as M1 + M2 for MySQL 8.4.
- M4 — Console (read). Overview, target timeline, backup inventory, diff, changes, activity with SSE. Gate: UI tests + p99 < 50 ms with 1,000 targets.
- M5 — Console (act) + sandboxes. Restore modes, sandbox lifecycle, query console, table browser, audit, roles via Access JWT (mocked in tests), and the `keeper` CLI covering every API action. Gate: all journeys through the UI and through the CLI (e2e runs the CLI binary).
- M6 — Verification, alerts, selective scope. Scheduled verification, PrometheusRules, Grafana dashboard,
scope.excludeDataand query exports, Postgres logical previews (optional). Gate: chaos suite green, alert rules unit-tested (promtool). - M7 — Tunnel + hardening. Cloudflare route management (mocked in tests), docs (README, RESTORE.md, ONCALL.md), perf report, security review, v0.1.0 release image.
Definition of done (whole project). All milestones gated green; make test-all < 20 min; chaos and perf suites green; README and runbooks complete; image published to therepublic.azurecr.io/keeper/keeper; Helm chart versioned.
Rollout (after DoD).
- One PR in
republic-gitops(branch, nevermain) that adds: Keeper CRDs and chart (HelmRelease ininfrastructure/orapps-infra/keeper/),keeper-systemnamespace, SOPS secrets placeholders (S3 key, age recipient public key), aprodanddevBackupPolicy, a BackupStore for Telnyx, and BackupTargets for exactly two databases:solar-search-dev/postgres-dev-solar(Postgres + PostGIS) andsolar-search-dev/mysql-wp(MySQL). Read-only and replication users are created by a documented one-off SQL script, not by the PR. - The PR description lists the manual steps (create bucket and key, add secrets, create DB users, Cloudflare tunnel token) and the verification checklist. A person merges.
- After a week of green backups, verified restores and a restore drill, a second PR adds the remaining databases (prod first).
Open decisions and assumptions¶
The build can start now; these items only need answers before the real rollout.
- Bucket name and region (
republic-db-backupsin Telnyx us-central-1, confirmed) and a key scoped to it, if Telnyx supports scoped keys. - Console hostname (assumed
keeper.republic.global) and sandbox TCP hostnames (sbx-<id>.db.republic.global); Cloudflare Access policy (which emails/groups). - Secret source for phase 1: SOPS (the republic-gitops PR uses the repository's SOPS key; placeholders to fill).
- Retention presets confirmed (prod 7d PITR, 14/8/12/5; dev 3d PITR, 7/4/3/0).
- SonarQube: keep issue triage (
issues,issue_changes) or treat as rebuildable. - Postgres row-level previews (logical slot) default off: confirm.
- Telnyx S3 compatibility: multipart upload, conditional PUT, listing and presigned URLs verified by
make test-live-s3againstrepublic-db-backups(us-central-1). - Disk headroom on nodes before rollout (image cache at \~80%); sandboxes use
emptyDirwith size limits by default. - Assumption: databases keep running as today (plain Deployments); Keeper does not manage their lifecycle. Replication and a second failure domain are future work.
Appendix: bootstrap prompt for the Claude Code cloud session¶
Paste this into a Claude Code cloud session opened on republic-global/keeper (with access to republic-global/republic-gitops for the final PR).
You are building **Keeper**, our in-house Kubernetes database backup & restore operator with a web console, in Go.
## Read first (mandatory, fully)
1. `CLAUDE.md` in this repo (rules).
2. `docs/DESIGN.md` in this repo: the complete design (copy of the Claude Doc
https://claude.ai/code/artifact/9f366401-b635-42c5-8fab-81f75a96bf63 ; if they differ, the newer one wins
and you update the other).
The design is the spec. Do not re-decide what it already decides; if something is missing or wrong, write an ADR in
`docs/adr/` with your decision and reason, then continue.
## What to build
Everything in the design: CRDs (BackupStore, BackupPolicy, BackupTarget, Backup, Restore, Sandbox), controller with
scheduler/jitter/Lease slots/load guard/retention/GC/verification, mover (full backups), streamer (Postgres WAL via
slot, MySQL binlog), restorer with point-in-time restore, inventory + version diffs + change summaries with masked
previews, sandboxes with TTL and Cloudflare tunnel routes, web console (Go templates/templ + HTMX + SSE, embedded),
query console on sandboxes, audit, Cloudflare Access JWT auth with roles, Prometheus metrics + PrometheusRules +
Grafana dashboard, CLI, Helm chart, image with Postgres 15/16/17 and MySQL 8.4 client tools. Engines: Postgres
(incl. PostGIS) and MySQL now, behind the Engine interface so MongoDB can be added later.
## Non-negotiable
- Never hang: no mounts/iSCSI/FUSE; data moves only in Jobs/streamers; every call has a context deadline; watchdog on
every stream; kill child process groups; crash-only; commit by manifest written last; no blocking finalizers.
- Performance: streaming pipeline (no temp files), parallel multipart, zstd multi-thread, console p99 < 50 ms.
- Security: age encryption client-side for all data objects; least-privilege DB users; no secrets in logs.
## How to work
- Work milestone by milestone (M0..M7 in the design). One branch + one PR per milestone in this repo; never push to
`main` after the bootstrap commit (install the pre-push hook as CLAUDE.md says). Merge your own PR only when
`make test-all` is green; if branch protection blocks you, leave the PR open and continue on a branch stacked on it.
- Test all the time in a local k3d cluster with MinIO and real Postgres/MySQL (testcontainers + e2e). Build the harness
in M0 (`hack/harness.sh`, `make test-all`) and keep it green. Add chaos tests (kill controller/mover/streamer/
restorer/api mid-flight, SIGSTOP the DB, toxiproxy on S3) and the PITR golden-checksum tests. Never skip or delete
a failing test to get green; fix the cause.
- Do not touch any real cluster, Cloudflare account or Telnyx bucket. Mock Cloudflare and Access in tests; use MinIO
for S3 (also run the S3 suite against a path-style, sigv4 config identical to Telnyx).
- Keep `docs/DESIGN.md`, README, `docs/RESTORE.md`, `docs/ONCALL.md` and ADRs up to date as you go.
- Commit often with conventional commits; push the branch regularly so progress is visible.
## When everything is done (all milestones gated green)
1. Publish v0.1.0: tag, CI builds `therepublic.azurecr.io/keeper/keeper:<tag>` (if registry credentials are not
available to you, make CI ready and document the one secret to add; do not invent credentials).
2. Open ONE PR in `republic-global/republic-gitops` (read its CLAUDE.md; branch, never `main`; squash merges) that
installs Keeper (CRDs, chart HelmRelease, keeper-system namespace, prod/dev BackupPolicy, Telnyx BackupStore
us-central-1 bucket `republic-db-backups`) and adds BackupTargets for exactly two databases:
`solar-search-dev/postgres-dev-solar` (postgis/postgis:16-3.4, db `solar-dev`) and `solar-search-dev/mysql-wp`
(mysql:8.4). Secrets as SOPS placeholders with clear TODOs (you cannot encrypt real values); include the one-off SQL
for read-only + replication users and a verification checklist in the PR description. Do not add other databases.
3. Report: what was built, test results (counts, durations, perf numbers), open issues, and the manual steps left.