skip to content

How does compression.type affect producer throughput, and how do lz4, zstd, and snappy compare?

level: middleimportance: should knowfreq 62%

answer

  1. compress per BATCH, not per record
  2. bigger batch = better ratio
  3. snappy/lz4 = fast, modest; zstd = harder, more CPU; gzip = max ratio, slow
  4. stored compressed end-to-end
  5. topic compression.type=producer avoids recompression

basics

~10 s

compression.type compresses each batch before sending, cutting network and disk usage and raising throughput. snappy/lz4 are fast with modest ratios; zstd compresses harder (smaller payloads) at more CPU; gzip is highest ratio but slowest.

solid answer

~50 s

compression.type (none|gzip|snappy|lz4|zstd, default none) tells the producer to compress each batch as a unit before it leaves the client. Because compression is per-batch, larger batches (bigger batch.size / higher linger.ms) compress better. Compressed batches reduce network bytes, broker disk usage, and replication traffic, and are stored compressed on the broker and decompressed by consumers — so the codec spans the whole pipeline. lz4 and snappy favor speed with moderate ratios (good default for high throughput, low CPU); zstd gives notably better ratios at tunable, higher CPU cost — excellent when network or storage is the bottleneck and you have CPU headroom. gzip has the best ratio but the worst speed. Set the broker topic's compression.type to 'producer' to keep the producer codec end-to-end and avoid broker-side recompression. zstd requires message format v2 and broker/consumer support (Kafka 2.1+).

go deeper

for a junior

Know compression cuts bytes sent and is set via compression.type; snappy/lz4 fast, gzip slow.

for a middle

Explain per-batch compression, the batch-size/ratio link, and the CPU vs ratio tradeoff across codecs.

for a senior

Reason about end-to-end storage/replication savings, topic compression.type=producer, and zstd version constraints.

for a principal

Choose codecs against the actual bottleneck (CPU vs network vs storage), validate with real payloads, and govern broker recompression policy fleet-wide.

## What compression.type does `compression.type` is a producer config (`none`, `gzip`, `snappy`, `lz4`, `zstd`; default `none`). When set, the producer compresses **each batch as a single unit** right before sending. The unit of compression is the batch — never a single record in isolation — which is why batching and compression reinforce each other: a bigger batch gives the compressor more redundant data to exploit, yielding a better ratio. ## Why it raises throughput Throughput here usually means bytes of useful data per second. Compression shrinks the on-the-wire payload, so: - **Network**: fewer bytes per produce request → more records fit in the same bandwidth and the same `max.request.size`. - **Broker disk**: Kafka stores batches **still compressed**, so log segments are smaller and the page cache holds more. - **Replication**: followers replicate the compressed bytes, multiplying the savings by the replication factor. - **Consumers**: read compressed data and decompress client-side, so the codec is effectively end-to-end. The trade is **CPU**: producers spend cycles compressing, consumers spend cycles decompressing. ## Codec comparison - **snappy**: very fast, low CPU, modest ratio. Good all-rounder for high-volume, CPU-sensitive pipelines. - **lz4**: similar speed class to snappy, often slightly better ratio/speed balance; a common modern default for throughput. - **zstd**: meaningfully better compression ratio, with speed that is still good (and tunable via level on some clients). Best when **network or storage is the bottleneck** and CPU is available. Requires message format v2 and Kafka 2.1+ on broker and consumers. - **gzip**: highest ratio, lowest speed/highest CPU. Use when bandwidth/storage is scarce and latency/CPU is not critical. ## Broker-side behavior A topic has its own `compression.type`. If it is set to `producer` (recommended), the broker keeps whatever the producer sent and does **no recompression** — cheapest and preserves the chosen codec end-to-end. If the topic specifies a different codec, the broker **decompresses and recompresses**, burning broker CPU and discarding the producer's choice. Mismatches can also force recompression when the broker must re-batch (e.g., for down-conversion to older consumers). ## Edge cases - Compression of already-compressed payloads (images, video, gzipped JSON) yields little gain and just wastes CPU — measure first. - Very small batches compress poorly; pair compression with reasonable `batch.size`/`linger.ms`. - zstd against old consumers that lack support causes errors or expensive down-conversion. - Compression ratio and codec choice should be validated with your real payload, not assumed from benchmarks.

  • Why does a larger batch.size or higher linger.ms improve the compression ratio?
    Compression operates on the whole batch, so more records per batch give the compressor more repeated/similar data to deduplicate, producing a better ratio than compressing many tiny batches independently.
  • What does setting a topic's compression.type to 'producer' achieve?
    The broker stores and serves whatever codec the producer used without decompressing/recompressing, saving broker CPU and keeping the codec consistent end-to-end through storage, replication, and consumers.

saying these in an interview costs you the question

  • Saying Kafka compresses each record individually (it's per batch)
  • Claiming the broker always decompresses to store data (it stores compressed)
  • Asserting zstd is always best — it costs more CPU and needs 2.1+ support
  • Forgetting consumers must support the codec (esp. zstd)
  • Recommending compression for already-compressed payloads

context