Walk through replica.fetch.max.bytes, replica.fetch.wait.max.ms, and num.replica.fetchers — what each controls and how you'd tune them.
answer
- max.bytes = per-partition batch cap, >= max.message.bytes
- wait.max.ms = long-poll timeout, < replica.lag.time.max.ms
- num.replica.fetchers = threads per remote leader, default 1
- One slow partition HOL-blocks its fetcher thread
- Tune for lag: more fetchers / bigger batches
basics
~10 sreplica.fetch.max.bytes caps bytes per partition per fetch; replica.fetch.wait.max.ms is the max long-poll wait when no data is ready; num.replica.fetchers sets how many fetcher threads a follower runs per source broker to parallelize replication.
solid answer
~50 sThese three broker configs shape follower replication throughput and latency. `replica.fetch.max.bytes` (default ~1 MB) limits how many bytes the leader returns *per partition* in one fetch response — it must be at least as large as the largest message (`max.message.bytes`) or replication of that record stalls; raising it lets a lagging follower pull bigger batches and catch up faster at the cost of memory. `replica.fetch.wait.max.ms` (default 500 ms) is the long-poll timeout: when a caught-up follower has nothing new to fetch, the leader holds the request open up to this long, then returns empty — lower it for tighter replication latency, but very low values increase request rate. `num.replica.fetchers` (default 1) is how many `ReplicaFetcherThread`s a follower broker spins up *per remote leader broker*; with one thread multiplexing many partitions, a single slow/large partition blocks the rest, so on high-partition-count or high-throughput clusters you raise it (e.g. 4-8) to parallelize and reduce replication lag.
go deeper
Recognize these are replication tuning knobs for batch size, wait time, and thread count.
State each default and its direct effect on follower fetch behavior.
Tune them together against lag, memory, max message size, and replica.lag.time.max.ms; know per-remote-leader thread semantics.
Diagnose replication-lag root causes across the whole pipeline (throttles, NIC, disk, request handlers) and set fleet-wide defaults.
These are all **broker-level** configs governing the follower side of replication. ### `replica.fetch.max.bytes` (default 1048576 = 1 MiB) - **What it does:** The maximum number of bytes the leader will return **per partition** in a single fetch response to a follower. It bounds the batch size pulled for each partition. - **Hard constraint:** It must be **>= the broker's `message.max.bytes` / topic `max.message.bytes`**. Kafka has a safety override so a fetch always returns at least one full record even if it exceeds the limit (otherwise a large message could never replicate), but you should still size this config above your max message to avoid pathological single-record fetches. - **Companion:** `replica.fetch.response.max.bytes` caps the *whole* response across all partitions; `replica.fetch.max.bytes` is per-partition within it. - **Tuning:** Larger values let a behind follower drain more per round-trip (fewer round-trips, faster catch-up) but increase per-fetch memory on both leader and follower. On fat, high-throughput partitions, raising it reduces replication lag. ### `replica.fetch.wait.max.ms` (default 500) - **What it does:** The maximum time the leader will block a follower's fetch request waiting for data when fewer than `replica.fetch.min.bytes` (default 1) bytes are available. This is the **long-poll** timeout for replication. - **Interaction:** Works with `replica.fetch.min.bytes`. With min.bytes=1, the leader returns as soon as any byte is available, so the wait mostly matters when the partition is idle — the request returns empty after the timeout. - **Tuning:** Lower it to reduce the worst-case staleness of a follower (and thus tighten how quickly the high watermark can advance under bursty load), at the cost of more frequent (possibly empty) fetch requests and CPU. Raising it batches more but adds latency. Must stay comfortably below `replica.lag.time.max.ms` so an idle-but-healthy follower isn't dropped from the ISR. ### `num.replica.fetchers` (default 1) - **What it does:** The number of `ReplicaFetcherThread`s each follower broker runs **toward each remote leader broker**. Partitions whose leader is broker X are distributed across these threads. - **Why it matters:** A single fetcher thread processes its assigned partitions sequentially; one large or slow partition can head-of-line-block replication of the others sharing that thread. More threads = more parallelism toward the same leader. - **Tuning:** On clusters with many partitions per broker and/or high write throughput, increasing to 4-8 (sometimes more) markedly reduces replication lag and speeds partition reassignment. The ceiling is CPU/network on the broker; too many threads add context-switching overhead and can overwhelm the leader's request handling. ### Putting it together A follower lagging behind is usually network/round-trip bound (raise `num.replica.fetchers` and/or `replica.fetch.max.bytes`) or, if idle-latency sensitive, you trim `replica.fetch.wait.max.ms`. Always validate against memory (`fetchers * partitions * fetch.max.bytes`), `replica.lag.time.max.ms`, and the max message size constraint.
- Why must replica.fetch.max.bytes be at least as large as the maximum message size?If a single message is larger than the per-partition fetch limit, the follower could never pull it and replication of that partition would stall. Kafka mitigates with a 'return at least one record' rule, but you should still set replica.fetch.max.bytes >= max.message.bytes to avoid inefficient single-record fetches and ISR shrinkage.
- You increase num.replica.fetchers from 1 to 8 and replication lag barely improves. What else might be the bottleneck?The leader's network/disk throughput, NIC saturation, replication throttles (leader/follower.replication.throttled.rate), small replica.fetch.max.bytes forcing many round-trips, or simply having few partitions per leader so extra threads have nothing to parallelize. Also check broker CPU and request-handler thread saturation.
saying these in an interview costs you the question
- Saying num.replica.fetchers is per-partition or per-broker globally — it's per remote leader broker
- Confusing replica.fetch.max.bytes (per-partition) with replica.fetch.response.max.bytes (whole response)
- Setting replica.fetch.wait.max.ms above replica.lag.time.max.ms
- Setting replica.fetch.max.bytes below max.message.bytes and expecting smooth replication