skip to content

What is max.partition.fetch.bytes versus fetch.max.bytes, and what breaks if max.partition.fetch.bytes is smaller than your largest message?

level: middleimportance: should knowfreq 50%

answer

  1. fetch.max.bytes = whole response (~50MB)
  2. max.partition.fetch.bytes = per partition (~1MB)
  3. KIP-74 -> limits soft, first record always returned
  4. old stall: msg > limit -> no progress
  5. set per-partition >= max message size

basics

~20 s

fetch.max.bytes caps the total size of a whole fetch response across all partitions; max.partition.fetch.bytes caps how much comes from each single partition. Historically, if a message was bigger than max.partition.fetch.bytes the consumer could stall — modern clients still return that oversized record so the consumer can progress.

solid answer

~40 s

fetch.max.bytes (default ~50MB) is the soft cap on the total bytes a single fetch response can return across all partitions. max.partition.fetch.bytes (default ~1MB) is the per-partition cap within that response. The consumer divides its fetch budget across the partitions it owns. The historical gotcha: if a single record on a partition exceeded max.partition.fetch.bytes, older consumers (pre-0.10.1 / KIP-74) could get stuck — the broker would never return that partition's data, and the consumer made no progress. Since KIP-74, both limits are 'soft': if the first record in a partition is larger than the limit, the broker still returns it so the consumer doesn't deadlock. So with modern brokers/clients you won't stall, but undersized max.partition.fetch.bytes still hurts throughput and can cause memory spikes when a large record forces an oversized response.

go deeper

for a junior

Know there's a per-partition cap and a total cap, and roughly their defaults.

for a middle

Explain the historical hard-limit stall and the KIP-74 soft-limit fix.

for a senior

Coordinate consumer-side limits with broker message.max.bytes and reason about memory = partitions x per-partition cap, bounded by total.

for a principal

Define platform limits end-to-end (producer/broker/replication/consumer) so large messages never silently break any tier.

## Two nested size limits A fetch response can pull from many partitions at once. Two configs bound it: - **fetch.max.bytes** (default ~**52428800** = 50MB): maximum total data the broker returns for the *entire* fetch response. - **max.partition.fetch.bytes** (default ~**1048576** = 1MB): maximum data returned *per partition* in that response. Think of fetch.max.bytes as the truck's total capacity and max.partition.fetch.bytes as the per-crate limit. The consumer spreads its budget across the partitions assigned to it. ## The classic stall (pre-KIP-74) Before Kafka 0.10.1, both limits were **hard**. If a producer wrote a message larger than max.partition.fetch.bytes, the broker could not fit that record under the limit, so it returned nothing for that partition. The consumer was permanently stuck at that offset — a silent poison-record stall. The standard fix was to raise max.partition.fetch.bytes above your largest possible message. ## KIP-74 made the limits soft Since 0.10.1, if the **first** record of a partition exceeds the limit, the broker returns it anyway, guaranteeing forward progress. So a modern consumer won't deadlock on a large record. But the limits still matter: - They cap **memory**: the consumer buffers fetched bytes; oversized responses spike heap. - They shape **throughput** and **fairness**: too-small per-partition limits mean many partitions return little data per fetch, increasing round trips. ## Tuning guidance - Set max.partition.fetch.bytes >= max.message.bytes (the broker/topic max record size) to avoid oversized single-record fetches surprising you, and to keep per-partition throughput healthy. - Keep fetch.max.bytes coordinated with how many partitions a consumer owns: total memory roughly scales with (partitions x max.partition.fetch.bytes) but is capped by fetch.max.bytes. - On the broker side, message.max.bytes and replica.fetch.max.bytes must also accommodate large messages, or production/replication breaks first. ## Edge cases - A consumer assigned 200 partitions with max.partition.fetch.bytes=10MB could theoretically want 2GB; fetch.max.bytes caps the actual response so memory stays bounded, but per-poll partition fairness suffers. - Compression is applied to the on-disk batch; the limits count compressed bytes returned over the wire.

  • Before KIP-74, how would you diagnose a consumer stuck because a message exceeded max.partition.fetch.bytes?
    Consumer lag on one partition would grow without bound while the consumer appeared healthy (still polling, no errors). You'd correlate the stuck offset with a known large message and confirm max.partition.fetch.bytes < that record size, then raise the limit above max.message.bytes.
  • If a consumer owns 100 partitions with max.partition.fetch.bytes=1MB but fetch.max.bytes=50MB, how is memory bounded?
    The per-partition limit would allow up to 100MB, but fetch.max.bytes caps the actual response at ~50MB. The broker fills partitions up to the total budget, so memory per fetch stays around 50MB rather than 100MB, at the cost of some partitions getting less data per round.

saying these in an interview costs you the question

  • Claiming modern consumers still permanently stall on an oversized record (KIP-74 fixed this).
  • Swapping the two: thinking fetch.max.bytes is per-partition.
  • Ignoring broker-side message.max.bytes / replica.fetch.max.bytes which must also fit large messages.

context