How do max.request.size (producer), message.max.bytes (broker), and replica.fetch.max.bytes interact, and what breaks if they're misaligned?
answer
- producer max.request.size ~1MB default
- broker message.max.bytes / topic max.message.bytes
- replica.fetch.max.bytes >= broker limit or replication stalls
- consumer fetch.max.bytes too
- limits apply to COMPRESSED batch; bump all in lockstep
basics
~20 smax.request.size limits how big a single producer request/record can be (default ~1 MB). message.max.bytes is the broker's per-batch ceiling, and max.message.bytes is its per-topic version. If the producer allows bigger messages than the broker accepts, the broker rejects them; replicas also need replica.fetch.max.bytes large enough or replication stalls.
solid answer
~50 sThese three limits must be aligned for large messages to flow end to end. max.request.size (producer, default 1048576 / 1 MB) caps the size of a single produce request — effectively the largest record (or batch for one partition) the client will send. message.max.bytes (broker, default ~1 MB) and its per-topic override max.message.bytes cap the largest batch the broker will accept and store. replica.fetch.max.bytes (and the broker-side fetch settings) must be at least as large so followers can replicate those big batches. Misalignment failures: if max.request.size > broker max.message.bytes, the broker returns RecordTooLargeException / MESSAGE_TOO_LARGE and the produce fails. If broker accepts a batch larger than replica.fetch.max.bytes, followers can't fetch it and replication stalls, blocking ISR progress — historically a cluster-wedging bug. Consumers also need fetch.max.bytes / max.partition.fetch.bytes large enough to read them. So raising max message size means coordinated bumps across producer, broker (topic), replication, and consumer.
go deeper
Know max.request.size (producer) and message.max.bytes (broker) must roughly match to send big messages.
Explain the rejection (RecordTooLargeException) when producer limit exceeds broker limit.
Cover the full chain including replica.fetch.max.bytes and consumer fetch limits, plus compressed-byte semantics.
Set org-wide message-size policy, prefer claim-check for large payloads, and reason about replication/ISR impact of oversized batches.
## The chain a large message must pass through A record produced to Kafka traverses several size gates. Each has its own limit and they must be **consistent** or the record is rejected or stuck somewhere in the pipeline. ### 1. Producer: max.request.size `max.request.size` (default `1048576` = 1 MB) caps the size of a **single produce request**. In practice it bounds the largest individual record the producer will send (a record can't exceed the request that carries it). The producer checks this client-side and throws `RecordTooLargeException` locally before sending if a record exceeds it. Note this is distinct from `batch.size` — `batch.size` is a target for batching efficiency; `max.request.size` is a hard ceiling. ### 2. Broker / topic: message.max.bytes (and max.message.bytes) `message.max.bytes` is the broker-default maximum size of a **record batch** the broker will accept (default ~1 MB; the exact default has shifted across versions). The per-topic override is `max.message.bytes`. If an incoming batch exceeds this, the broker rejects the produce with `MESSAGE_TOO_LARGE` and the client surfaces a `RecordTooLargeException`. Because compression happens before this check on the producer and the broker validates the (possibly compressed) batch size, **compression can let a logically large payload fit under the limit**. ### 3. Replication: replica.fetch.max.bytes Follower replicas fetch batches from the leader to replicate. `replica.fetch.max.bytes` must be **>= the largest batch the broker accepts**. If a leader accepts a 5 MB batch but `replica.fetch.max.bytes` is 1 MB, followers cannot fetch that batch, fall out of the ISR, and replication stalls — historically this could wedge a partition. Modern Kafka mitigates by always allowing at least one batch, but you should still size it correctly. ### 4. Consumer: fetch.max.bytes / max.partition.fetch.bytes Consumers must be able to fetch large batches too. `max.partition.fetch.bytes` and `fetch.max.bytes` need to accommodate them, or consumers can't make progress on partitions holding oversized batches (again, modern clients guarantee at least one batch to avoid total stall). ## What breaks on misalignment | Misalignment | Symptom | |---|---| | producer max.request.size > broker max.message.bytes | Broker rejects: MESSAGE_TOO_LARGE / RecordTooLargeException | | broker max.message.bytes > replica.fetch.max.bytes | Followers can't replicate; ISR shrinks; under-replicated partitions | | broker/topic limit > consumer fetch limits | Consumer can't read oversized records (mitigated by min-one-batch rule) | | producer limit < actual record size | Client-side RecordTooLargeException before send | ## How compression interacts The size checks generally apply to the **compressed batch bytes**. So enabling `compression.type` can shrink a batch below `max.message.bytes`, letting larger logical payloads through. But never rely on compression to dodge a hard limit for variable data — worst-case incompressible payloads can blow past it. ## Operational guidance To support larger messages, bump all four in lockstep: `max.request.size` (producers), `message.max.bytes`/topic `max.message.bytes` (brokers/topics), `replica.fetch.max.bytes` (replication), and `fetch.max.bytes`/`max.partition.fetch.bytes` (consumers). For very large payloads, prefer the **claim-check pattern** (store the blob in object storage, send a reference) instead of pushing limits sky-high.
- A producer's max.request.size is 5 MB but the topic max.message.bytes is 1 MB. What happens?The producer sends the large batch, but the broker rejects it with MESSAGE_TOO_LARGE, surfaced to the client as RecordTooLargeException. The client-side check passes (record is under 5 MB) so the failure happens at the broker. You must raise the topic/broker limit to match.
- Why must replica.fetch.max.bytes track message.max.bytes?Follower replicas fetch batches from the leader. If the broker accepts a batch larger than replica.fetch.max.bytes, followers can't fetch it, fall out of the ISR, and replication stalls — leaving under-replicated partitions. It must be at least as large as the biggest batch the broker accepts.
- Does compression help you fit under these limits?Yes — the size checks apply to the compressed batch bytes, so enabling compression.type can bring a large logical payload under message.max.bytes. But don't rely on it for incompressible data; worst case it won't shrink and you'll hit the limit. For truly large blobs use the claim-check pattern.
saying these in an interview costs you the question
- Confusing max.request.size with batch.size (one is a hard ceiling, the other a batching target)
- Forgetting replica.fetch.max.bytes, causing silent replication stalls
- Assuming raising only the producer or only the broker limit is enough
- Not realizing the limits apply to compressed bytes