What is max.partition.fetch.bytes, how does it relate to fetch.max.bytes, and what failure mode arises if a record exceeds it?
answer
- per-partition cap ~1MB
- fetch.max.bytes total ~50MB (KIP-74)
- both SOFT — first batch always returned
- no permanent stall on big record
- set >= broker max.message.bytes
basics
~20 smax.partition.fetch.bytes caps bytes returned per partition per fetch (default ~1MB); fetch.max.bytes caps total bytes across all partitions (~50MB). Both are soft limits: to avoid stalling, a broker always returns at least one full record batch even if it exceeds the cap.
solid answer
~50 smax.partition.fetch.bytes (default 1,048,576) limits how many bytes a single partition can contribute to one fetch response; fetch.max.bytes (default ~50MB) limits the total across all partitions in the response. Both are SOFT limits with a critical safety rule: to guarantee progress, the broker will always return at least the first record batch of a partition even if it is larger than the cap. So a single oversized message does NOT permanently stall the consumer — modern Kafka makes forward progress. The historical failure mode (older clients/strict interpretation) was a stuck consumer if a record exceeded max.partition.fetch.bytes; since KIP-74 the per-partition limit and overall fetch.max.bytes coexist and the oversized-batch guarantee prevents the stall. Practically, max.partition.fetch.bytes also bounds per-partition memory and influences fairness across many assigned partitions, and it should be set at least as large as the broker's message.max.bytes / topic max.message.bytes to comfortably handle big records.
go deeper
Know there is a per-partition byte cap on fetches, default about 1MB.
Distinguish per-partition vs total fetch caps and that both are about bytes.
Explain the soft-limit progress guarantee, KIP-74's fetch.max.bytes, and the misconfiguration failure mode.
Design fetch sizing for large-message topics and high partition counts, balancing fairness, memory, and the message.max.bytes alignment.
## Per-partition and total fetch byte caps A fetch response can pull from many partitions at once. Two settings bound its size. ### max.partition.fetch.bytes (default 1 MB) The maximum bytes returned **per partition** in a single fetch. With N assigned partitions, the naive worst case is roughly N × max.partition.fetch.bytes — which is why a second, overall cap exists. ### fetch.max.bytes (default ~50 MB) The maximum bytes returned **across all partitions** in one fetch response. Introduced by **KIP-74** precisely so that consumers subscribed to many partitions don't blow up memory: it bounds the whole response, while max.partition.fetch.bytes still keeps any single partition from monopolizing the response. ### Both are SOFT limits — the progress guarantee Kafka stores records in compressed **record batches**. A limit could be smaller than the first batch on a partition, which would mean the broker returns zero usable data → the consumer can never advance → permanent stall. To prevent this, the rule is: **the broker always returns at least one full record batch for the first non-empty partition**, even if that batch exceeds the byte cap. Hence the caps are soft. This guarantees liveness regardless of message size. ### Failure mode, then and now - **Historically / conceptually**, people feared: "if a record is bigger than max.partition.fetch.bytes the consumer gets stuck." In strict early implementations this was a real risk. - **Today**, thanks to the oversized-batch guarantee, a single large message is still delivered — the consumer makes progress, just with a response larger than the soft cap for that fetch. - The remaining real-world failure is **misconfiguration**: setting max.partition.fetch.bytes well *below* the topic/broker `message.max.bytes` (a.k.a. `max.message.bytes`) wastes the soft-cap intent and can cause large per-fetch responses or memory pressure. Best practice: max.partition.fetch.bytes ≥ broker max message size. ### Why both settings exist (interplay) - `max.partition.fetch.bytes` ensures **fairness** — one hot partition can't starve others by filling the whole response. - `fetch.max.bytes` ensures **bounded total memory** — protects the client heap across all partitions. - Together they shape per-fetch memory: total response ≤ ~fetch.max.bytes, each partition ≤ ~max.partition.fetch.bytes (both modulo the oversized-batch exception). ### Tuning notes - Many partitions + small max.partition.fetch.bytes → more even round-robin draining across partitions per fetch. - Large records → raise max.partition.fetch.bytes to keep responses efficient and avoid relying on the single-batch override every time. - Watch client heap: prefetch buffers hold up to these byte sizes per in-flight fetch per broker.
- Does a message larger than max.partition.fetch.bytes permanently stall a modern consumer?No. The limit is soft: the broker always returns at least the first full record batch for a partition even if it exceeds the cap (the progress guarantee), so the oversized message is delivered and the consumer advances.
- Why introduce fetch.max.bytes when max.partition.fetch.bytes already exists?Because with many assigned partitions, per-partition caps multiply (N × cap) and can balloon memory. KIP-74's fetch.max.bytes bounds the total response size across all partitions, protecting the client heap while the per-partition cap preserves fairness.
saying these in an interview costs you the question
- Asserting a record bigger than max.partition.fetch.bytes permanently blocks the consumer (modern clients make progress).
- Treating the caps as hard limits with no oversized-batch exception.
- Confusing per-partition (max.partition.fetch.bytes) with total (fetch.max.bytes).
- Ignoring the need to align max.partition.fetch.bytes with the broker's max.message.bytes.