Skip to content

0013: Configurable compression, on by default

Date: 2026-10-07 · Status: accepted

Context

Every data object has always been compressed with zstd (klauspost/compress, multi-threaded) before age encryption. The level was fixed at default, and compression could not be turned off. Most backups compress well: SQL text and dump formats often shrink 3–10×, and Postgres pages 2–4×. Payloads that are already compressed (images, archives, encrypted columns) gain nothing and only cost CPU.

Decision

  • On by default. A policy without compression uses zstd at level default, the same as before.
  • Policy setting. BackupPolicy.spec.compression:
  • algorithm: zstd (default) or none.
  • level: fastest, default, better or best.
  • threads: zstd threads per mover. The default is one per CPU of the Job's limit; streamers use one thread unless this is set.
  • It applies to full backups (movers) and change segments (streamers). Small metadata objects (inventories, previews) always use zstd, because their cost is negligible.
  • Compression happens in one place: the Keeper pipeline, before encryption.
  • Dump tools run uncompressed (pg_dump --compress=0, plain mysqldump, pg_basebackup -Ft without compression), so data is never compressed twice.
  • Compression never runs on the database server: no server-zstd for pg_basebackup, no protocol compression. CPU is spent in the mover pod, where the policy's Job limits bound it, and never on production databases.
  • Encrypted bytes do not compress, so compressing after encryption would gain nothing.
  • Readers detect the format. After decryption, a reader checks for the zstd frame magic and decompresses only when it is there. Old objects, new objects and objects written with none all read through the same path, and changing a policy never breaks restores of older backups. The plaintext of every format Keeper writes (tar, PGDMP, SQL text, WAL pages, binlog events, JSON) never starts with the zstd magic.
  • Manifests also record compression per object, for display and audits.
  • Object names keep their .zst.age suffix whatever the setting, so the storage layout and listings are unchanged.

Consequences

  • There is no real drawback to the default. zstd default compresses at several hundred MB/s per core, faster than an S3 upload, and decompression speed does not depend on the level. Restores stay as fast as before.
  • better and best trade mover CPU for less storage and egress, a good fit for large, rarely restored databases.
  • none exists for data that is already compressed, or for a mover whose CPU is the bottleneck.