Keeper on-call runbook¶
Every Keeper alert links here. First steps for any alert:
kubectl -n keeper-system get pods # controller, api, streamers, jobs
kubectl get backuptargets -A # phase, last backup, PITR window
kubectl -n <ns> describe backuptarget <name> # conditions, events, messages
kubectl -n keeper-system logs deploy/keeper-controller --since=1h # JSON logs (never contain secrets)
keeper targets --health error # the console's view (CLI)
State lives in two places only: Keeper resources in the cluster and objects in the store. Any Keeper pod can be
deleted at any time. A backup exists only once its manifest.json is written, and a half-finished one is cleaned up
by garbage collection.
BackupOverdue¶
No successful full backup for 26 h (prod) or 50 h (other environments). KeeperBackupNeverSucceeded fires when a
target has never had one.
kubectl -n <ns> get backups -l keeper.republic.global/target=<name> --sort-by=.metadata.creationTimestamp. Look at the newest backup'sstatus.phaseandstatus.message:Postponed(load guard): the database was busy at schedule time. The run is retriedloadGuard.maxPostponestimes, then runs anyway.Queued: waiting for a concurrency slot (keeper_slots_in_use,keeper_queue_depth). Another long backup on the same host holds the per-host slot.Failed: read the message. The mover Job's logs are inkeeper-system(kubectl -n keeper-system logs job/<job>). Common causes:- Credentials: the
keeper-*secret is wrong or missing in the target namespace. - Network: a NetworkPolicy in the target namespace must allow the
keeper-systemmover pods. - Stalled: no progress for
stallTimeout; the watchdog killed the dump. - The store is unreachable: see StoreUnreachable.
- Credentials: the
- Fix the cause, then run
keeper backup now <ns>/<name> --wait. - A suspended target (
spec.suspend: true, or paused from the console) never backs up. Resume it withkubectl patchor from the console. Git (Flux) restoresspec.suspendon its next sync.
StreamLagHigh¶
The change stream is more than 5 minutes behind: recent seconds are not yet restorable.
kubectl -n keeper-system logs deploy/keeper-stream-<ns>-<name>-…: look for reconnect loops and upload errors.- The store is slow or unreachable (uploads retry with back-off): see StoreUnreachable.
- A burst of writes larger than the upload bandwidth (
limits.bandwidth): the lag drains by itself. - The streamer restarts by itself when it sees no stream activity for 3 minutes (liveness).
StreamChainBroken¶
Point-in-time restore has a gap. The streamer records why in stream/position.json, which shows as the target's
status.pitr message:
- Postgres: the slot was invalidated (
max_slot_wal_keep_size), dropped, or the database was recreated or promoted (new system identifier or timeline). - MySQL: the binlog needed next was purged (
binlog_expire_logs_secondstoo short for an outage), or the server's identity changed.
Keeper starts a new chain at once and takes a full backup for it (reason chain-restart). Restores are possible
up to the break and from the new base on; keeper restore refuses times in the gap and prints the restorable
windows. Act on the cause: raise max_slot_wal_keep_size or binlog retention, and look for long outages of the
streamer.
VerificationFailed¶
A scheduled or on-demand verification restore failed, or its checks failed. keeper-system events and the
restore show which:
kubectl -n <ns> get restores -l keeper.republic.global/verification=true --sort-by=.metadata.creationTimestamp
kubectl -n <ns> describe restore <name> # status.checks: inventory-matches, schema-matches, row estimates, SQL
inventory-matchesorschema-matchesfailed: the restored database differs from what the backup recorded. Treat this as a damaged backup. Pin the last good backup (keeper pin <id> --reason …), take a new full backup, and verify it (keeper verify <ns>/<name> --wait).- A row estimate is outside tolerance: usually a large table whose statistics were stale. Run
ANALYZEon the source, or exclude the table from the check. - A target SQL check failed: look at the check's message. It is the target owner's assertion about their data.
- The restore itself failed: the restorer logs are kept for 24 h (
kubectl -n keeper-system logs job/<job>).
KeeperVerificationOverdue means no successful verification in 8 days: check the policy's verify.schedule, and
whether verification restores fail to start (sandbox quota, images).
SlotWALRetentionHigh¶
A Postgres replication slot holds more than 75 % of max_slot_wal_keep_size. At 100 % Postgres invalidates the
slot (the chain breaks), or, without a limit, the database disk fills.
- The streamer is down or stuck: restart it (
kubectl -n keeper-system rollout restart deploy/keeper-stream-…). It resumes from its confirmed position. - The slot belongs to a target that no longer streams. Keeper drops slots when PITR is turned off or the target
is deleted (ADR 0011). If the cleanup failed, for example because the credentials secret went first, drop the
slot by hand:
select pg_drop_replication_slot('keeper_<ns>_<target>');(names:select slot_name, active from pg_replication_slots;).
StoreUnreachable¶
keeper_store_up == 0 for 10 min: Keeper cannot list or write the bucket.
- Credentials: the store's
credentialsSecret(S3 key id and secret; on Telnyx both are the API key). - Endpoint and region: path-style SigV4. A wrong region usually returns 301 or 403.
- Network: egress from
keeper-systemto the endpoint.
Backups fail (and retry on schedule) while the store is down. Change streams keep at most the current segment locally: they stop confirming, so the slot or binlog position holds the changes on the database until the store returns.
GCFailing¶
Garbage collection has not completed for 2 days. Expired backups stay, which costs storage but loses nothing.
Look at the controller logs (gc messages) and the store's status.message. The plan of each run is written to
gc/<time>.json in the store before anything is deleted.
SandboxExpiryFailing¶
Sandboxes are past their TTL and their namespaces still exist. Look at kubectl get sandboxes, the namespace
keeper-sbx-<name> (stuck finalizers on objects inside it), and the controller logs. A sandbox finalizer is
dropped after a timeout whatever happens, so a leftover namespace can be deleted by hand.
Restoring under pressure¶
See RESTORE.md.