Skip to content

0011: Change stream lifecycle: chain start, liveness and slot cleanup

Date: 2026-10-07 · Status: accepted

Context

Running the PITR e2e test in a real cluster showed three gaps in the M2 streamer:

  • A new chain was persisted only after its first complete WAL segment. On a quiet database that can take hours, so the controller never saw the chain start and never took the base backup the chain needs.
  • On an idle database, Postgres sends no keepalives while the client keeps sending status updates. The streamer saw no traffic, and its liveness probe restarted a healthy pod every few minutes. Each restart began a new chain.
  • Nothing dropped a target's replication slot when PITR was turned off or the target was deleted. A leftover physical slot retains WAL on the database forever and can fill its disk.

Decision

  • Chain start. The streamer writes stream/position.json (chain id and start, no position yet) as soon as the stream is connected and the slot exists. A restart with such a position continues the same chain from the slot. The controller's chain-restart backup therefore always starts after the slot reserves WAL.
  • Liveness. Every 30 s the receiver asks the server for a reply (standby status update with reply requested). Only messages from the server count as stream activity, so liveness still detects a dead connection or an upload stuck past the watchdog.
  • Slot cleanup.
  • PITR targets carry the finalizer keeper.republic.global/stream-cleanup.
  • When PITR is turned off, or the target is deleted, the controller removes the streamer and runs a short Job, keeper streamer --cleanup, which drops the slot over the replication protocol. It waits up to 2 min for the slot to become inactive, and a missing slot counts as success.
  • The finalizer goes once the Job finishes. It also goes after 15 min whatever happens: finalizers never block, and a leftover slot is covered by the KeeperSlotWALRetentionHigh alert and the documented manual drop.

Consequences

The database keeps no WAL for targets that no longer stream. Deleting a PITR target takes a few seconds longer (the cleanup Job). If the credentials secret is deleted together with the target, the cleanup fails and the operator drops the slot by hand (docs/ONCALL.md).