Skip to content

Keeper on-call runbook

Every Keeper alert links here. First steps for any alert:

kubectl -n keeper-system get pods                                  # controller, api, streamers, jobs
kubectl get backuptargets -A                                       # phase, last backup, PITR window
kubectl -n <ns> describe backuptarget <name>                       # conditions, events, messages
kubectl -n keeper-system logs deploy/keeper-controller --since=1h  # JSON logs (never contain secrets)
keeper targets --health error                                      # the console's view (CLI)

State lives in two places only: Keeper resources in the cluster and objects in the store. Any Keeper pod can be deleted at any time. A backup exists only once its manifest.json is written, and a half-finished one is cleaned up by garbage collection.

BackupOverdue

No successful full backup for 26 h (prod) or 50 h (other environments). KeeperBackupNeverSucceeded fires when a target has never had one.

  1. kubectl -n <ns> get backups -l keeper.republic.global/target=<name> --sort-by=.metadata.creationTimestamp. Look at the newest backup's status.phase and status.message:
  2. Postponed (load guard): the database was busy at schedule time. The run is retried loadGuard.maxPostpones times, then runs anyway.
  3. Queued: waiting for a concurrency slot (keeper_slots_in_use, keeper_queue_depth). Another long backup on the same host holds the per-host slot.
  4. Failed: read the message. The mover Job's logs are in keeper-system (kubectl -n keeper-system logs job/<job>). Common causes:
    • Credentials: the keeper-* secret is wrong or missing in the target namespace.
    • Network: a NetworkPolicy in the target namespace must allow the keeper-system mover pods.
    • Stalled: no progress for stallTimeout; the watchdog killed the dump.
    • The store is unreachable: see StoreUnreachable.
  5. Fix the cause, then run keeper backup now <ns>/<name> --wait.
  6. A suspended target (spec.suspend: true, or paused from the console) never backs up. Resume it with kubectl patch or from the console. Git (Flux) restores spec.suspend on its next sync.

StreamLagHigh

The change stream is more than 5 minutes behind: recent seconds are not yet restorable.

  • kubectl -n keeper-system logs deploy/keeper-stream-<ns>-<name>-…: look for reconnect loops and upload errors.
  • The store is slow or unreachable (uploads retry with back-off): see StoreUnreachable.
  • A burst of writes larger than the upload bandwidth (limits.bandwidth): the lag drains by itself.
  • The streamer restarts by itself when it sees no stream activity for 3 minutes (liveness).

StreamChainBroken

Point-in-time restore has a gap. The streamer records why in stream/position.json, which shows as the target's status.pitr message:

  • Postgres: the slot was invalidated (max_slot_wal_keep_size), dropped, or the database was recreated or promoted (new system identifier or timeline).
  • MySQL: the binlog needed next was purged (binlog_expire_logs_seconds too short for an outage), or the server's identity changed.

Keeper starts a new chain at once and takes a full backup for it (reason chain-restart). Restores are possible up to the break and from the new base on; keeper restore refuses times in the gap and prints the restorable windows. Act on the cause: raise max_slot_wal_keep_size or binlog retention, and look for long outages of the streamer.

VerificationFailed

A scheduled or on-demand verification restore failed, or its checks failed. keeper-system events and the restore show which:

kubectl -n <ns> get restores -l keeper.republic.global/verification=true --sort-by=.metadata.creationTimestamp
kubectl -n <ns> describe restore <name>      # status.checks: inventory-matches, schema-matches, row estimates, SQL
  • inventory-matches or schema-matches failed: the restored database differs from what the backup recorded. Treat this as a damaged backup. Pin the last good backup (keeper pin <id> --reason …), take a new full backup, and verify it (keeper verify <ns>/<name> --wait).
  • A row estimate is outside tolerance: usually a large table whose statistics were stale. Run ANALYZE on the source, or exclude the table from the check.
  • A target SQL check failed: look at the check's message. It is the target owner's assertion about their data.
  • The restore itself failed: the restorer logs are kept for 24 h (kubectl -n keeper-system logs job/<job>).

KeeperVerificationOverdue means no successful verification in 8 days: check the policy's verify.schedule, and whether verification restores fail to start (sandbox quota, images).

SlotWALRetentionHigh

A Postgres replication slot holds more than 75 % of max_slot_wal_keep_size. At 100 % Postgres invalidates the slot (the chain breaks), or, without a limit, the database disk fills.

  • The streamer is down or stuck: restart it (kubectl -n keeper-system rollout restart deploy/keeper-stream-…). It resumes from its confirmed position.
  • The slot belongs to a target that no longer streams. Keeper drops slots when PITR is turned off or the target is deleted (ADR 0011). If the cleanup failed, for example because the credentials secret went first, drop the slot by hand: select pg_drop_replication_slot('keeper_<ns>_<target>'); (names: select slot_name, active from pg_replication_slots;).

StoreUnreachable

keeper_store_up == 0 for 10 min: Keeper cannot list or write the bucket.

  • Credentials: the store's credentialsSecret (S3 key id and secret; on Telnyx both are the API key).
  • Endpoint and region: path-style SigV4. A wrong region usually returns 301 or 403.
  • Network: egress from keeper-system to the endpoint.

Backups fail (and retry on schedule) while the store is down. Change streams keep at most the current segment locally: they stop confirming, so the slot or binlog position holds the changes on the database until the store returns.

GCFailing

Garbage collection has not completed for 2 days. Expired backups stay, which costs storage but loses nothing. Look at the controller logs (gc messages) and the store's status.message. The plan of each run is written to gc/<time>.json in the store before anything is deleted.

SandboxExpiryFailing

Sandboxes are past their TTL and their namespaces still exist. Look at kubectl get sandboxes, the namespace keeper-sbx-<name> (stuck finalizers on objects inside it), and the controller logs. A sandbox finalizer is dropped after a timeout whatever happens, so a leftover namespace can be deleted by hand.

Restoring under pressure

See RESTORE.md.