skip to content

Producer Throughput vs Latency Tuning

Pushing producer throughput with batching, linger, compression and buffer sizing, and knowing what each costs in latency. Interviewers ask you to walk the curve rather than recite a magic value.

part ofApache Kafkaoverview, primer and where to startread it →
on this pageshow

questions

6

What do batch.size and linger.ms do on a Kafka producer, and how do they interact to trade latency for throughput?

level: juniorimportance: must knowfreq 78%

answer

  1. batch.size = bytes per partition batch (16KB default)
  2. linger.ms = wait time, default 0
  3. send when batch full OR linger elapsed
  4. bigger both = throughput up, latency up
  5. linger barely matters under high volume

basics

~10 s

batch.size is the max bytes per partition batch; linger.ms is how long the producer waits to fill a batch before sending. Raising either groups more records per request, boosting throughput but adding latency.

solid answer

~40 s

A Kafka producer buffers records per topic-partition into batches. batch.size (default 16384 bytes) caps a single batch's size — when it fills, the batch is sent immediately. linger.ms (default 0) is an artificial delay the producer waits, even when a batch isn't full, to give more records a chance to join. With linger.ms=0 a record can ship the moment a sender thread is free, minimizing latency but producing many small requests. Setting linger.ms=5-100 and a larger batch.size (e.g. 64KB-256KB) lets batches fill, so each produce request carries more records — fewer requests, better compression ratios, higher throughput — at the cost of up to linger.ms extra latency. A batch is sent when EITHER it reaches batch.size OR linger.ms elapses, whichever comes first.

go deeper

for a junior

Know the units (batch.size=bytes, linger.ms=ms) and the core tradeoff: bigger = more throughput, more latency.

for a middle

Explain the 'full OR linger elapsed' send rule and why linger barely affects high-volume streams.

for a senior

Reason about the latency/throughput curve, compression interaction, and per-partition batching memory implications.

for a principal

Frame defaults as SLA-driven choices, model the curve under bursty vs steady load, and coordinate batch.size with buffer.memory and partition count.

## The problem A Kafka producer can send one record per network request, but that is wasteful: each request has fixed overhead (TCP, request headers, broker processing, an ack round-trip). Batching many records into one request amortizes that overhead and dramatically raises throughput. The two knobs that control batching are `batch.size` and `linger.ms`. ## How the producer buffers When you call `producer.send(record)`, the record is NOT sent immediately. It is appended to an in-memory buffer (a `RecordAccumulator`) keyed by **topic-partition**. Each partition has its own queue of batches. A background **sender thread** drains ready batches and sends them to brokers. Batches are formed and compressed per partition. ## batch.size - Unit: **bytes**, default **16384 (16KB)**. It is the maximum size of a single batch for one partition. - When a partition's current batch reaches `batch.size`, it becomes 'ready' and is eligible to send immediately — you do not wait for `linger.ms`. - A single record larger than `batch.size` still gets its own batch (it isn't rejected by this setting; `max.request.size` governs the hard limit). - Larger `batch.size` = more records per request and usually better compression, but more memory held per partition and potentially higher latency if traffic is low (the batch takes longer to fill). ## linger.ms - Unit: **milliseconds**, default **0**. It is how long the producer will deliberately wait for more records to accumulate before sending a not-yet-full batch. - With `linger.ms=0`, a batch is sent as soon as a sender thread is available — even with a single record. This minimizes latency but tends to create many small batches under bursty load. - With `linger.ms=20`, the producer holds the batch up to 20ms hoping more records arrive so the batch ships fuller. ## The send rule A batch is dispatched when **EITHER** condition is met, whichever happens first: 1. The batch reaches `batch.size`, OR 2. `linger.ms` has elapsed since the batch's first record. So under high volume, batches fill before `linger.ms` and the linger barely matters; under low volume, `linger.ms` caps the added latency. This is why even a small linger (5-10ms) can sharply cut request count on bursty workloads without hurting tail latency much. ## The latency/throughput curve - Low latency config: `linger.ms=0`, modest `batch.size`. Records go out ASAP, more, smaller requests, lower throughput ceiling. - High throughput config: `linger.ms=20-100`, large `batch.size` (64KB-1MB). Fuller batches, fewer requests, better compression, higher max throughput, but each record waits up to `linger.ms` longer. ## Edge cases - `batch.size=0` disables batching entirely (each record sent separately) — almost never desirable. - `buffer.memory` must be large enough to hold batches for all partitions; if it fills, `send()` blocks up to `max.block.ms`. - Batches also count against `max.in.flight.requests.per.connection` and broker-side `max.request.size`. - Increasing `linger.ms` improves compression because the compressor sees more data per batch.

  • If linger.ms is 0, does that mean batching never happens?
    No. Even at linger.ms=0, records that arrive while the sender thread is busy with the previous request accumulate into a batch, so natural batching still occurs under load. linger.ms=0 just means the producer never adds artificial delay.
  • What happens if a single record exceeds batch.size?
    It still gets sent in its own batch; batch.size is not a hard reject. The hard ceiling is max.request.size (default 1MB) and the broker's message.max.bytes; exceeding those throws RecordTooLargeException.

saying these in an interview costs you the question

  • Saying linger.ms is the time a record always waits (it's a max; a full batch ships earlier)
  • Claiming batch.size is per-producer rather than per-topic-partition
  • Thinking linger.ms=0 disables all batching
  • Confusing batch.size (bytes) with a record count

context

open as a page

How does the acks setting trade durability against producer throughput and latency?

level: middleimportance: must knowfreq 80%

basics

~20 s

acks controls how many brokers must confirm a write. acks=0 (no wait) is fastest but can lose data; acks=1 waits for the leader; acks=all waits for all in-sync replicas — most durable but highest latency.

open as a page

How does compression.type affect producer throughput, and how do lz4, zstd, and snappy compare?

level: middleimportance: should knowfreq 62%

basics

~10 s

compression.type compresses each batch before sending, cutting network and disk usage and raising throughput. snappy/lz4 are fast with modest ratios; zstd compresses harder (smaller payloads) at more CPU; gzip is highest ratio but slowest.

open as a page

What is buffer.memory, and what happens when a high-volume producer outpaces the brokers?

level: seniorimportance: should knowfreq 48%

basics

~10 s

buffer.memory is the total bytes the producer can use to buffer unsent records. When it fills (the producer outpaces brokers), send() blocks up to max.block.ms, then throws TimeoutException. It's the producer's backpressure mechanism.

open as a page

What does max.in.flight.requests.per.connection control, and how does it interact with retries, ordering, and idempotence?

level: seniorimportance: should knowfreq 55%

basics

~20 s

It's the number of unacknowledged produce requests the producer allows per broker connection at once. Higher values pipeline more for throughput; with retries and no idempotence, values above 1 can reorder records on a failed-then-retried batch.

open as a page

Design producer configs for two workloads: (a) lowest p99 latency for small synchronous events, and (b) maximum throughput for a high-volume bulk ingest. Justify each setting.

level: principalimportance: should knowfreq 42%

basics

~10 s

Low latency: linger.ms=0, small batches, compression none/lz4, acks=1, modest buffer. High throughput: linger.ms=20-100, large batch.size, zstd/lz4 compression, acks=all with pipelining, big buffer.memory. The two sit at opposite ends of the batching/latency curve.

open as a page