When does a Kafka broker recompress (or decompress) batches instead of storing them as-is, and why does that matter?
answer
- default = pass-through + sendfile zero-copy
- topic compression.type != producer forces recompress
- old-consumer down-conversion breaks zero-copy
- LogAppendTime touches batch header
- watch FetchMessageConversionsPerSec
basics
~20 sNormally a broker stores the producer's compressed batch untouched (zero-copy friendly). But if the topic's compression.type forces a different codec, or older message-format conversion / timestamp validation / offset assignment requires it, the broker must decompress and recompress, costing CPU and breaking the zero-copy path.
solid answer
~50 sBy default a broker keeps producer-compressed batches as-is: it appends them to the log in their original compressed form and serves them to consumers via the zero-copy sendfile path, paying no (de)compression cost. The broker recompresses only when forced. The main triggers: (1) the topic-level compression.type is set to a specific codec different from the producer's, so the broker decompresses and recompresses to that codec; (2) message format down-conversion for old consumers (e.g. v2 to v1) requires unwrapping and rewrapping batches, defeating zero-copy; (3) certain validation paths — log append time timestamps (message.timestamp.type=LogAppendTime) or non-idempotent offset/sequence rewriting in older formats — require touching batch contents. Recompression burns broker CPU, increases append latency, and disables sendfile. Best practice: set topic compression.type=producer so the broker passes the codec through, and keep all clients on a modern message format to avoid down-conversion.
go deeper
Know that brokers usually store the producer's compressed batch untouched.
Explain that topic compression.type=producer avoids recompression and that mismatched codecs force it.
Detail zero-copy/sendfile, down-conversion triggers, and the CPU/latency/memory cost of recompression.
Design client-upgrade and topic-config policy to guarantee pass-through; capacity-plan brokers for conversion edge cases and monitor conversion metrics.
## The default: broker is a pass-through Kafka's performance story leans heavily on the broker being **cheap**. When a producer sends a compressed batch, the broker normally: 1. Validates the batch (CRC, format) without fully decompressing payloads where possible. 2. Appends the batch to the partition log **in the same compressed bytes**. 3. On fetch, serves those bytes to consumers using the **zero-copy `sendfile` syscall** — data goes from page cache straight to the socket without entering the JVM heap or being decompressed. The consumer, not the broker, decompresses. This is why brokers can push enormous throughput. ## What forces decompress + recompress The broker must crack open and re-encode batches when: ### 1. Topic-level compression.type mismatch The topic config `compression.type` defaults to `producer` (keep whatever the producer used). If an operator sets it to a specific codec (e.g. `gzip`) and the producer sent `lz4`, the broker **decompresses the lz4 batch and recompresses it as gzip** before writing. Setting it to `uncompressed` forces decompression. Only `producer` avoids this. ### 2. Message format down-conversion Kafka's on-disk record format has versions (v0/v1 legacy, v2 = the modern batch format from KIP-98/KIP-32 era). If a **consumer too old** to understand v2 fetches data stored as v2, the broker must **down-convert** — unwrap the batch, rewrite records in the old format, recompress. This defeats zero-copy (data must go through the heap) and spikes CPU/memory. The metric `FetchMessageConversionsPerSec` exposes this. ### 3. Timestamp / validation rewrites If the topic uses `message.timestamp.type=LogAppendTime`, the broker stamps each batch with its own time, requiring it to touch the batch header (cheaper than full recompression but still a write). In older formats, assigning offsets/sequence numbers could force payload rewrites. ## Why it matters - **CPU**: (de)compression on the broker steals cycles from the I/O-bound fast path; gzip recompression in particular is expensive. - **Latency**: append path lengthens; p99 produce latency rises. - **Zero-copy lost**: down-conversion routes bytes through the JVM heap, increasing GC pressure and memory use, and can cause OOM under heavy old-consumer fetch. - **Throughput cliff**: a cluster sized assuming pass-through can fall over when a fleet of legacy consumers triggers conversions. ## Best practices 1. Keep topic `compression.type=producer` so the broker passes the codec through. 2. Upgrade all clients so everyone speaks the modern message format — no down-conversion. 3. Monitor `FetchMessageConversionsPerSec` and `ProduceMessageConversionsPerSec`. 4. Prefer `CreateTime` timestamps unless you specifically need `LogAppendTime`. 5. If you must standardize a codec cluster-wide, accept the recompression CPU cost knowingly and size brokers for it.
- What topic setting avoids broker recompression, and what's its default?compression.type at the topic level. Its default is 'producer', meaning the broker keeps whatever codec the producer used and stores the batch as-is. Setting it to a specific codec (gzip, etc.) or 'uncompressed' forces decompress/recompress. Leave it 'producer' to preserve the zero-copy pass-through path.
- Why does an old consumer fetching modern-format data hurt the broker?The broker must down-convert v2 batches to the older message format the consumer understands. That requires unwrapping/recompressing batches and routing bytes through the JVM heap, defeating zero-copy sendfile, spiking CPU and GC, and risking OOM. The FetchMessageConversionsPerSec metric flags it.
saying these in an interview costs you the question
- Claiming the broker always decompresses every batch (it normally doesn't)
- Not knowing topic compression.type defaults to 'producer'
- Forgetting that down-conversion for old clients breaks zero-copy
- Saying consumers never decompress (they normally do — the broker passes through)