skip to content

How does compression.type work at the record-batch level, and what happens when producer, topic, and broker compression settings differ?

level: seniorimportance: should knowfreq 28%

answer

  1. codec in batch attributes bits; whole records section
  2. none/gzip/snappy/lz4/zstd
  3. topic compression.type default = producer (store as-is)
  4. topic codec ≠ producer codec → broker recompress, lose zero-copy
  5. bigger batch.size/linger.ms → better ratio; consumer decompresses

basics

~20 s

compression.type sets the codec (none, gzip, snappy, lz4, zstd) applied to the whole batch's record section, recorded in the batch header. The producer normally compresses; the broker keeps it as-is unless the topic's compression.type forces a different codec, which makes the broker recompress.

solid answer

~50 s

Compression in Kafka is **per batch**: the producer's **compression.type** (none/gzip/snappy/lz4/zstd) compresses the entire records section of a batch, and the codec is encoded in the batch header **attributes** bits. The whole batch travels and is stored compressed; the **consumer decompresses**. The topic also has a **compression.type** (default `producer`, meaning "keep whatever the producer used"). If the topic specifies a concrete codec that **differs** from the batch's codec, the broker must **decompress and recompress** on append — losing zero-copy on that path and adding CPU. If topic `compression.type=producer`, the broker stores the batch unchanged (cheap, zero-copy-friendly). zstd is often the best ratio/CPU trade-off but requires message format v2 and reasonably modern clients. Compressing larger batches (tune `batch.size`/`linger.ms`) yields better ratios because there's more redundancy per batch. Note: `gzip` is CPU-heavy; `lz4`/`snappy` are fast/lower ratio; `zstd` tunable and usually the modern default choice.

go deeper

for a junior

Know compression.type picks a codec (gzip/snappy/lz4/zstd) and applies to batches.

for a middle

Explain producer vs topic compression.type and that consumers decompress.

for a senior

Detail the broker recompression path when codecs differ and the zero-copy/CPU consequences.

for a principal

Reason about codec selection trade-offs, batch tuning for ratio, format-version constraints, and standardizing compression.type=producer across a platform.

## Compression is a batch-level property In RecordBatch v2, **compression applies to the entire records section of a batch**, not to individual records and not to the batch header. The chosen **codec** is stored in the header's **attributes** bitfield. Supported codecs: - **none** — no compression. - **gzip** — highest ratio historically, but CPU-heavy. - **snappy** — fast, moderate ratio. - **lz4** — fast, moderate ratio. - **zstd** — tunable, typically best ratio-per-CPU; requires message format v2 (KIP-110, Kafka 2.1+). Because it's per batch, **bigger batches compress better** (more repeated keys/fields to exploit). So tuning **`batch.size`** and **`linger.ms`** upward improves compression as a side effect. ## Producer vs topic vs broker settings There are two main `compression.type` knobs: 1. **Producer `compression.type`** (default `none`): the producer compresses each batch with this codec before sending. 2. **Topic `compression.type`** (default **`producer`**): controls what the broker stores. - **`producer`** → the broker **keeps the producer's codec as-is**. The batch is stored and later streamed **unchanged**, preserving **zero-copy** and avoiding broker CPU. - a **specific codec** (e.g. `zstd`) → if it **matches** the producer's codec, no work; if it **differs**, the broker must **decompress the batch and recompress** it with the topic codec on append. This costs CPU and **breaks zero-copy** for that write (the broker had to touch the bytes). - **`uncompressed`** → broker decompresses and stores uncompressed. There is also a **cluster/broker default** (`compression.type` in broker config) that supplies the topic default when a topic doesn't override it. ## Consumer side The **consumer always decompresses** — the broker never decompresses to serve reads (that's what keeps the fetch path zero-copy when the codec isn't being changed). The consumer reads the codec from the batch attributes and inflates the records. ## Practical guidance - Prefer topic `compression.type=producer` to keep brokers cheap and preserve zero-copy; pick the actual codec on the producer. - Use **zstd** for the best ratio/CPU balance on modern clients; **lz4**/**snappy** when CPU/latency matter more than ratio; avoid gzip unless ratio is paramount. - Raise `batch.size`/`linger.ms` to fill batches so compression has more to work with. - Forcing a topic codec that differs from producers is usually an anti-pattern: it doubles compression work (producer compresses, broker recompresses) and kills zero-copy. ## Edge cases - **End-to-end format version:** zstd needs v2 batches and clients that understand it; very old consumers trigger **down-conversion**, which also breaks zero-copy. - **Re-compression on append** also recomputes the batch CRC (the records changed), unlike a plain offset re-stamp. - Compression interacts with `max.message.bytes`: limits apply to the **compressed** batch size on the broker but the producer also checks uncompressed bounds via `max.request.size`.

  • What is the effect of setting a topic's compression.type to a specific codec that differs from what producers send?
    The broker must decompress each incoming batch and recompress it with the topic's codec on append — extra CPU, a recomputed CRC, and loss of zero-copy. Leaving it as 'producer' avoids all that by storing batches unchanged.
  • Why do larger batches compress better, and which configs increase batch size?
    Compression exploits redundancy across records (repeated keys, field names); more records per batch means more redundancy to remove. Raise batch.size and linger.ms to let batches fill before sending.

saying these in an interview costs you the question

  • Saying compression is per record rather than per batch.
  • Claiming the broker always decompresses to serve consumers (only the consumer decompresses).
  • Thinking topic compression.type=producer makes the broker recompress (it stores as-is).
  • Forgetting that a mismatched topic codec breaks zero-copy and adds CPU.

context