0011: Change stream lifecycle: chain start, liveness and slot cleanup¶
Date: 2026-10-07 · Status: accepted
Context¶
Running the PITR e2e test in a real cluster showed three gaps in the M2 streamer:
- A new chain was persisted only after its first complete WAL segment. On a quiet database that can take hours, so the controller never saw the chain start and never took the base backup the chain needs.
- On an idle database, Postgres sends no keepalives while the client keeps sending status updates. The streamer saw no traffic, and its liveness probe restarted a healthy pod every few minutes. Each restart began a new chain.
- Nothing dropped a target's replication slot when PITR was turned off or the target was deleted. A leftover physical slot retains WAL on the database forever and can fill its disk.
Decision¶
- Chain start. The streamer writes
stream/position.json(chain id and start, no position yet) as soon as the stream is connected and the slot exists. A restart with such a position continues the same chain from the slot. The controller's chain-restart backup therefore always starts after the slot reserves WAL. - Liveness. Every 30 s the receiver asks the server for a reply (standby status update with reply requested). Only messages from the server count as stream activity, so liveness still detects a dead connection or an upload stuck past the watchdog.
- Slot cleanup.
- PITR targets carry the finalizer
keeper.republic.global/stream-cleanup. - When PITR is turned off, or the target is deleted, the controller removes the streamer and runs a short Job,
keeper streamer --cleanup, which drops the slot over the replication protocol. It waits up to 2 min for the slot to become inactive, and a missing slot counts as success. - The finalizer goes once the Job finishes. It also goes after 15 min whatever happens: finalizers never block,
and a leftover slot is covered by the
KeeperSlotWALRetentionHighalert and the documented manual drop.
Consequences¶
The database keeps no WAL for targets that no longer stream. Deleting a PITR target takes a few seconds longer (the cleanup Job). If the credentials secret is deleted together with the target, the cleanup fails and the operator drops the slot by hand (docs/ONCALL.md).