Explain how fetch.min.bytes and fetch.max.wait.ms work together to trade off latency vs throughput in consumer fetches.
answer
- min.bytes = wait for this much
- max.wait.ms = but no longer than this
- respond on EITHER condition
- min.bytes=1 default = lowest latency
- raise min.bytes = throughput, +latency
basics
~20 sfetch.min.bytes tells the broker the minimum data to accumulate before answering a fetch; the broker waits up to fetch.max.wait.ms for that much to arrive, then responds anyway. Bigger min.bytes = better throughput but higher latency.
solid answer
~40 sWhen a consumer sends a Fetch request, the broker doesn't have to respond instantly. fetch.min.bytes (default 1) is the minimum number of bytes the broker should accumulate before sending a response. fetch.max.wait.ms (default 500ms) caps how long the broker will hold the request waiting for that much data. So the broker responds when EITHER it has at least fetch.min.bytes available OR fetch.max.wait.ms elapses — whichever comes first. With the default fetch.min.bytes=1, the broker replies as soon as any data exists, minimizing latency. Raising fetch.min.bytes (e.g. to 64KB) makes the broker batch more before responding: fewer, larger fetch responses improve throughput and reduce broker CPU/network overhead, at the cost of up to fetch.max.wait.ms extra latency when traffic is light. It's the classic latency-vs-throughput knob, server-side batching analogous to Nagle's algorithm.
go deeper
Know fetch.min.bytes is a batching threshold and fetch.max.wait.ms caps the wait.
Explain the either/or trigger and the latency-vs-throughput trade-off with defaults.
Tune both for a workload (quiet vs busy topics) and predict worst-case latency from the wait window.
Set fetch sizing as part of a latency budget across producer linger, broker batching, and consumer prefetch.
## Server-side fetch batching A consumer asks a broker for data via a **Fetch request**. Naively the broker could answer immediately with whatever exists, but that produces many tiny responses under low-volume topics, wasting CPU and network. Kafka adds a server-side batching window controlled by two settings. ### fetch.min.bytes (default 1) The minimum amount of data (in bytes) the broker should gather before replying to a fetch. If set to 1, the broker replies the instant any record is available (lowest latency). If set to, say, 65536, the broker waits until at least ~64KB has accumulated across the requested partitions. ### fetch.max.wait.ms (default 500) The maximum time the broker will hold the fetch request waiting to reach fetch.min.bytes. This is the safety valve: even if data trickles in slowly, the broker will not delay the response beyond this window. ### The combined rule The broker responds when **either** condition is met: - accumulated bytes ≥ `fetch.min.bytes`, **or** - elapsed time ≥ `fetch.max.wait.ms`. So `fetch.max.wait.ms` bounds the worst-case added latency from this batching. ### Tuning trade-off - **Low latency** (default): `fetch.min.bytes=1` → respond ASAP. - **High throughput / fewer requests**: raise `fetch.min.bytes` so each response carries more data; fewer round trips, less per-request overhead on both ends, better compression batching. Downside: on a quiet topic, a consumer may wait up to `fetch.max.wait.ms` for the response. ### Worked example Topic gets ~1 record/sec, each 200 bytes. - `fetch.min.bytes=1`: each fetch returns ~immediately with one record → ~1ms latency, ~1 fetch/record. - `fetch.min.bytes=10000`, `fetch.max.wait.ms=500`: broker waits, never reaches 10KB in 500ms, so responds at the 500ms timeout with whatever it has (~a few records) → up to 500ms latency, far fewer fetches. ### Relationship to other settings These govern *when* the broker responds and *minimum* size; `fetch.max.bytes` and `max.partition.fetch.bytes` cap the *maximum* size of the response. `max.poll.records` is unrelated — it shapes the client-side hand-off after data arrives. ### Analogy Like a shuttle bus that leaves when it's full (`fetch.min.bytes`) or when a fixed timer expires (`fetch.max.wait.ms`), whichever comes first — full buses are efficient, the timer caps how long passengers wait.
- If you set fetch.min.bytes very high on a low-traffic topic, what's the user-visible effect?Added latency: the broker holds each fetch up to fetch.max.wait.ms (default 500ms) before responding because the byte threshold is rarely reached, so consumers see records delayed by up to that window.
saying these in an interview costs you the question
- Saying the broker always waits the full fetch.max.wait.ms regardless of data volume.
- Claiming fetch.min.bytes counts records rather than bytes.
- Forgetting that the broker responds on whichever condition is met first.
- Confusing fetch.min.bytes (a minimum/threshold) with fetch.max.bytes (a maximum cap).