Monitoring¶
Turn on the chart's monitoring objects (they need the Prometheus Operator and a Grafana sidecar):
metrics:
serviceMonitor: { enabled: true }
prometheusRule: { enabled: true, overdueHours: { prod: 26, default: 50 } }
grafanaDashboard: { enabled: true }
Alerts¶
| Alert | Fires when | Runbook |
|---|---|---|
KeeperBackupOverdue |
No successful full backup for overdueHours (per env label) |
BackupOverdue |
KeeperBackupNeverSucceeded |
A target has never had a successful backup | BackupOverdue |
KeeperStreamLagHigh |
The change stream is behind | StreamLagHigh |
KeeperStreamChainBroken |
The point-in-time chain broke (slot lost, binlog purged) | StreamChainBroken |
KeeperVerificationFailed |
The last verification restore failed | VerificationFailed |
KeeperVerificationOverdue |
No successful verification for 8 days | VerificationFailed |
KeeperSlotWALRetentionHigh |
A replication slot retains a lot of WAL on the database | SlotWALRetentionHigh |
KeeperStoreUnreachable |
A BackupStore cannot be reached | StoreUnreachable |
KeeperGCFailing |
Garbage collection has not succeeded recently | GCFailing |
KeeperSandboxExpiryFailing |
Sandboxes past their TTL are not being deleted | SandboxExpiryFailing |
The rules ship in charts/keeper/files/prometheus-rules.yaml and are unit-tested with promtool. Every metric is
listed in Metrics.
Dashboard¶
The Keeper Grafana dashboard shows, per target: last success, backup duration and size (raw and stored), stream lag, point-in-time window, verification results and restore time, slot WAL retention, store size, queue depth and active sandboxes.