skip to content

Your monitoring shows UnderReplicatedPartitions spiking and IsrShrinksPerSec elevated, but brokers are all up. How do you diagnose and act?

level: seniorimportance: should knowfreq 42%

answer

  1. Brokers up but followers slow ≠ dead broker
  2. kafka-topics --describe --under-replicated-partitions
  3. Find common broker / leader bandwidth
  4. Disk, network, GC, num.replica.fetchers
  5. Don't just raise lag.time.max.ms to hide it

basics

~20 s

Under-replicated partitions with all brokers up means followers can't keep up, not that brokers are down. Check follower I/O (disk, network, GC), look at IsrShrinks/ExpandsPerSec for flapping, inspect which brokers are dropping out, and either fix the bottleneck or tune replica.lag.time.max.ms / replica fetcher threads.

solid answer

~40 s

UnderReplicatedPartitions > 0 with healthy brokers indicates followers are lagging past `replica.lag.time.max.ms`, so leaders shrank the ISR. First confirm scope: `kafka-topics --describe --under-replicated-partitions` shows which partitions and which replicas are missing from the ISR. Correlate the dropped replicas to specific brokers — usually a few overloaded or slow brokers. Then look at root causes on those brokers: disk saturation, network bandwidth limits, long GC pauses, or too few `num.replica.fetchers` to keep up with write volume. High `IsrShrinksPerSec`/`IsrExpandsPerSec` together signal flapping near the threshold. Remediation: relieve the bottleneck (rebalance partitions, add fetcher threads, fix disk/network), and if the lag is transient and benign, raise `replica.lag.time.max.ms` to reduce churn. Always check `UnderMinIsrPartitionCount` to see if any writes are actually failing.

go deeper

for a junior

Know that under-replicated partitions mean some replicas are not in the ISR.

for a middle

Use kafka-topics --describe to find which replicas and brokers are lagging.

for a senior

Diagnose root cause (disk/network/GC/fetchers/throttles) and pick the right remediation.

for a principal

Set up alerting thresholds, capacity planning, and reassignment strategy to keep ISR healthy under load.

## Reading the signal correctly **`UnderReplicatedPartitions`** is the count of partitions where the current ISR size is **less than** the replication factor — i.e., at least one assigned replica is not in-sync. Crucially, **brokers being up does not mean replication is healthy**: a broker can be running yet too slow to keep its followers caught up, which still shrinks the ISR. **`IsrShrinksPerSec`** and **`IsrExpandsPerSec`** together describe *churn*. A burst of shrinks that quickly re-expands means followers are **flapping** around the `replica.lag.time.max.ms` boundary — they fall behind, get evicted, catch up, rejoin, repeat. ## Diagnostic workflow 1. **Scope it.** Run `kafka-topics.sh --bootstrap-server ... --describe --under-replicated-partitions`. Each line shows Leader, Replicas (AR), and Isr. The replicas in AR but missing from Isr are the lagging ones. 2. **Find the common broker.** If the same broker id is the missing replica across many partitions, that broker is the bottleneck (its fetchers can't keep up, or it's the slow follower). If the missing replica is always a *follower of a particular leader*, the **leader** may be saturating outbound bandwidth. 3. **Inspect broker health on the suspects:** - **Disk:** is the log directory's disk at 100% util / high await? Slow appends keep followers behind. - **Network:** is inter-broker bandwidth saturated? Replication competes with client traffic. - **GC:** long stop-the-world pauses stall the fetcher threads; check GC logs. - **Fetcher threads:** `num.replica.fetchers` (default 1) may be too low for the partition count / throughput; followers fetch serially and fall behind. - **Throttles:** an active replication quota (`replica.alter.log.dirs.io.max.bytes.per.second` or leader/follower replication throttles set during a reassignment) can artificially slow replication. 4. **Check write impact.** `UnderMinIsrPartitionCount` and `UnderMinIsrPartitionCount > 0` tells you whether the shrink has crossed `min.insync.replicas`, meaning acks=all producers are now failing — that escalates urgency. ## Remediation levers - **Relieve load:** reassign partitions off the hot broker (`kafka-reassign-partitions.sh`), or add brokers and rebalance. - **Increase `num.replica.fetchers`** so followers replicate from multiple leaders in parallel. - **Fix the resource bottleneck:** faster disks, more network, GC tuning (G1/ZGC, heap sizing). - **Remove stale throttles** left over from a prior reassignment. - **Tune `replica.lag.time.max.ms`:** if the lag is genuinely transient and harmless, a larger window stops the flapping; but don't mask a real, sustained problem this way — it just delays detection. ## What NOT to conclude - Don't assume a dead broker — all brokers being up is the whole point of this scenario. - Don't immediately blow up `replica.lag.time.max.ms` as a first move; that hides the symptom. Find the cause first. - A single brief spike that self-heals may be a routine event (e.g., a follower restart catching up); persistent or growing under-replication is the real alert. ## Key metrics cheat sheet - `kafka.server:type=ReplicaManager,name=UnderReplicatedPartitions` - `kafka.server:type=ReplicaManager,name=UnderMinIsrPartitionCount` - `kafka.server:type=ReplicaManager,name=IsrShrinksPerSec` / `IsrExpandsPerSec` - Per-broker `RequestHandlerAvgIdlePercent`, network/io thread idle, disk util, GC pause time.

  • How do you tell whether the bottleneck is a slow follower broker or a saturated leader?
    Look at which replicas are missing from the ISR across partitions. If one broker is consistently the missing follower, it is the slow one. If a particular leader's followers all lag, the leader's outbound network or disk is likely saturated, throttling replication for everyone fetching from it.
  • Why is reflexively increasing replica.lag.time.max.ms a poor first response?
    It widens the window so lagging followers stay counted as in-sync longer, suppressing the alert without fixing the underlying slowness. It also weakens acks=all durability, since a genuinely behind follower now lingers in the ISR. Diagnose the resource bottleneck first.

saying these in an interview costs you the question

  • Assuming under-replicated partitions imply a broker is down.
  • Jumping to raise replica.lag.time.max.ms before diagnosing the cause.
  • Ignoring leftover replication throttles from a prior reassignment.
  • Confusing UnderReplicatedPartitions (ISR < RF) with UnderMinIsrPartitionCount (ISR < min.insync.replicas).

context