skip to content

Batching, Linger and Compression

Trading latency for throughput with batch.size, linger.ms, and batch-level compression codecs. Interviewers ask because these three settings move producer throughput more than anything else on the client.

part ofApache Kafkaoverview, primer and where to startread it →
on this pageshow

questions

5

What do batch.size and linger.ms control in a Kafka producer, and how do they trade latency for throughput?

level: juniorimportance: must knowfreq 78%

answer

  1. batch.size = byte ceiling, linger.ms = time ceiling
  2. send when EITHER fills first
  3. defaults: 16384 bytes, 0 ms
  4. linger.ms=0 still batches under load
  5. bigger/longer = throughput, costs latency

basics

~20 s

batch.size caps how many bytes a producer collects per partition before sending; linger.ms tells it to wait up to that many milliseconds for more records to fill a batch. Bigger batches and more lingering raise throughput but add latency.

solid answer

~40 s

A Kafka producer groups records destined for the same partition into batches. batch.size (default 16384 bytes) is the maximum size of a single batch; once a batch fills, it is sent immediately. linger.ms (default 0) is how long the producer will wait, after the first record is added, for more records to accumulate before sending — even if batch.size isn't reached. With linger.ms=0 the producer sends as soon as a sender thread is free, giving lowest latency but small, frequent batches. Raising linger.ms (e.g. 5-100ms) lets batches grow, improving compression ratios and throughput and reducing request overhead, at the cost of added per-record latency. A batch is sent when EITHER batch.size is hit OR linger.ms elapses, whichever comes first. Tuning both together is the core latency-vs-throughput knob.

go deeper

for a junior

Know the two knobs and that bigger/longer = more throughput, more latency.

for a middle

Explain the OR-trigger semantics and that linger.ms=0 still batches under load.

for a senior

Tie batching to RecordAccumulator/Sender, buffer.memory back-pressure, and compression-ratio effects; give concrete tuning ranges.

for a principal

Reason about end-to-end latency budgets, choosing linger per workload SLA, and the interaction with idempotence, in-flight requests, and broker throughput.

## What problem batching solves A Kafka producer does not send every `send()` call to the broker as its own network request — that would be wasteful (per-request overhead, TCP round trips, broker-side work). Instead it **accumulates records into batches**, one batch per topic-partition, inside an in-memory structure called the **RecordAccumulator**. A background **Sender** thread drains ready batches and ships them to brokers. ## batch.size `batch.size` (default `16384` = 16 KB) is the **maximum number of bytes for a single batch** for one partition. Important nuances: - It is a per-partition cap, not a global one. - A batch is allocated from a buffer pool sized by `buffer.memory` (default 32 MB total). - When a batch reaches `batch.size`, it becomes **ready to send immediately**. - A single record larger than `batch.size` still gets its own batch (it isn't rejected for exceeding batch.size; that's `max.request.size`'s job). ## linger.ms `linger.ms` (default `0`) is the **maximum time the producer waits** for a batch to fill before sending it, measured from when the first record lands in that batch. It is an artificial delay introduced to let more records accumulate. - With `linger.ms=0`, the producer does NOT busy-wait — it sends a batch as soon as a sender thread is available. Under load, batches still form naturally because records queue while the sender is busy with prior requests. - With `linger.ms>0`, the producer deliberately waits, trading latency for fuller batches. ## The send trigger: OR semantics A batch is dispatched when **either** condition is met first: 1. The batch reaches `batch.size`, OR 2. `linger.ms` elapses since the first record was added. So `batch.size` is a size ceiling and `linger.ms` is a time ceiling. ## The latency/throughput trade-off - **Low latency**: `linger.ms=0`, small batches. Records leave fast but you do many small requests — more overhead, weaker compression. - **High throughput**: larger `batch.size` (e.g. 32-256 KB) and `linger.ms` of a few ms to tens of ms. Fewer, fatter requests; better compression ratios; less broker CPU per record. Each record waits longer. ## Edge cases - If your produce rate is high, even `linger.ms=0` yields large batches because records pile up behind in-flight requests. - If your rate is low, `linger.ms` is what creates batching at all — without it, you'd send tiny one-record batches. - `buffer.memory` exhaustion blocks `send()` (or throws after `max.block.ms`); batching tuning interacts with this back-pressure.

  • If linger.ms is 0, does the producer never batch?
    No — it still batches under load. With linger.ms=0 the producer sends as soon as a sender thread is free, but while a prior request is in flight, new records accumulate into the next batch. Batching emerges naturally from throughput; linger.ms only forces extra waiting when the rate is low.
  • What happens to a record larger than batch.size?
    It still gets sent — it forms its own batch of that larger size. batch.size doesn't reject big records; the hard upper limit is max.request.size (default 1 MB), which WILL reject records exceeding it.

saying these in an interview costs you the question

  • Saying linger.ms=0 means no batching ever (wrong — it batches under load)
  • Thinking batch.size is a global limit rather than per-partition
  • Claiming a record bigger than batch.size is rejected (that's max.request.size)
  • Believing the producer always waits the full linger.ms even when batch.size is reached first

context

open as a page

How does compression.type work in a Kafka producer, and how do gzip, snappy, lz4, and zstd compare?

level: middleimportance: must knowfreq 70%

basics

~20 s

compression.type tells the producer to compress each batch as a whole using a codec: gzip (best ratio, slow), snappy (fast, modest ratio), lz4 (fast), or zstd (great ratio and good speed). Compression is per-batch, so bigger batches compress better.

open as a page

How do max.request.size (producer), message.max.bytes (broker), and replica.fetch.max.bytes interact, and what breaks if they're misaligned?

level: seniorimportance: must knowfreq 55%

basics

~20 s

max.request.size limits how big a single producer request/record can be (default ~1 MB). message.max.bytes is the broker's per-batch ceiling, and max.message.bytes is its per-topic version. If the producer allows bigger messages than the broker accepts, the broker rejects them; replicas also need replica.fetch.max.bytes large enough or replication stalls.

open as a page

When does a Kafka broker recompress (or decompress) batches instead of storing them as-is, and why does that matter?

level: seniorimportance: should knowfreq 45%

basics

~20 s

Normally a broker stores the producer's compressed batch untouched (zero-copy friendly). But if the topic's compression.type forces a different codec, or older message-format conversion / timestamp validation / offset assignment requires it, the broker must decompress and recompress, costing CPU and breaking the zero-copy path.

open as a page

You need to maximize producer throughput for a high-volume Kafka pipeline. Which producer configs do you tune together, and what are the trade-offs and failure modes?

level: principalimportance: should knowfreq 40%

basics

~20 s

Raise batch.size and linger.ms so batches fill, enable a fast codec like lz4 or zstd via compression.type, and increase buffer.memory so the producer doesn't block. Accept added per-record latency and watch for buffer exhaustion, ordering, and broker size limits.

open as a page