Skip to content

Monitoring

Turn on the chart's monitoring objects (they need the Prometheus Operator and a Grafana sidecar):

metrics:
  serviceMonitor: { enabled: true }
  prometheusRule: { enabled: true, overdueHours: { prod: 26, default: 50 } }
  grafanaDashboard: { enabled: true }

Alerts

Alert Fires when Runbook
KeeperBackupOverdue No successful full backup for overdueHours (per env label) BackupOverdue
KeeperBackupNeverSucceeded A target has never had a successful backup BackupOverdue
KeeperStreamLagHigh The change stream is behind StreamLagHigh
KeeperStreamChainBroken The point-in-time chain broke (slot lost, binlog purged) StreamChainBroken
KeeperVerificationFailed The last verification restore failed VerificationFailed
KeeperVerificationOverdue No successful verification for 8 days VerificationFailed
KeeperSlotWALRetentionHigh A replication slot retains a lot of WAL on the database SlotWALRetentionHigh
KeeperStoreUnreachable A BackupStore cannot be reached StoreUnreachable
KeeperGCFailing Garbage collection has not succeeded recently GCFailing
KeeperSandboxExpiryFailing Sandboxes past their TTL are not being deleted SandboxExpiryFailing

The rules ship in charts/keeper/files/prometheus-rules.yaml and are unit-tested with promtool. Every metric is listed in Metrics.

Dashboard

The Keeper Grafana dashboard shows, per target: last success, backup duration and size (raw and stored), stream lag, point-in-time window, verification results and restore time, slot WAL retention, store size, queue depth and active sandboxes.