skip to content

Producers using acks=all see p99 latency spikes, and you notice ISR shrinking on some partitions. Explain the chain of causation and which metrics you'd correlate to confirm a slow follower is the culprit.

level: seniorimportance: must knowfreq 62%

answer

  1. acks=all waits on slowest ISR follower
  2. slow follower -> RemoteTimeMs up -> p99 up
  3. lag > replica.lag.time.max.ms (30s) -> ISR shrink
  4. IsrShrinksPerSec / UnderReplicatedPartitions / ISR flap
  5. pin culprit via that follower's GC/disk/network

basics

~20 s

With acks=all a produce request can't complete until all in-sync replicas (the ISR) catch up. A slow follower lags, so the produce waits longer (high RemoteTimeMs); if it lags past replica.lag.time.max.ms the leader drops it from ISR (ISR shrink). Correlate RemoteTimeMs, replica lag, ISR-shrink rate, and the follower's own health.

solid answer

~40 s

acks=all means the leader only acknowledges a produce once every replica in the ISR has replicated the records. A slow follower (GC, disk, network, or overloaded broker) replicates late, so the produce sits in purgatory longer — visible as rising Produce RemoteTimeMs and p99. If the follower's lag exceeds replica.lag.time.max.ms (default 30s), the leader ejects it from the ISR, raising IsrShrinksPerSec and lowering that partition's effective replication. To confirm, I correlate: Produce RemoteTimeMs p99 on the leader; per-replica lag via UnderReplicatedPartitions and the follower's FetchFollower request rate/latency; IsrShrinksPerSec / IsrExpandsPerSec churn; and the suspect follower's own broker signals (GC pauses, RequestHandlerAvgIdlePercent, disk await, network). I also check replica.fetch threads (num.replica.fetchers) and whether min.insync.replicas is forcing waits. A flapping ISR (shrink/expand churn) plus one broker with bad local metrics pins the slow follower.

go deeper

for a junior

Know that acks=all makes producers wait for replicas and that a slow replica makes producing slower and can shrink the ISR.

for a middle

Explain the lag-to-shrink chain, name replica.lag.time.max.ms, and know RemoteTimeMs and ISR-shrink metrics are the signals.

for a senior

Trace the full causation, correlate leader RemoteTime with per-replica lag and the follower's local health, and recognize the post-shrink durability trap and min.insync.replicas interaction.

for a principal

Set durability/latency policy (acks, min.insync.replicas, replica.lag.time.max.ms, fetcher sizing), design alerting on ISR churn, and architect to isolate slow-broker blast radius.

## Definitions first - **Replica / ISR.** Each partition has a leader and follower replicas. The **ISR (in-sync replica set)** is the subset of replicas currently caught up with the leader. A follower stays in ISR as long as it has fetched up to the leader's log end within `replica.lag.time.max.ms` (default 30 000 ms). - **acks=all (acks=-1).** The producer setting that makes the leader wait until **all replicas in the ISR** have replicated the batch before acknowledging. This gives durability but couples producer latency to the slowest in-sync follower. - **min.insync.replicas.** The minimum ISR size required for an acks=all produce to succeed. If ISR shrinks below it, acks=all produces fail with NotEnoughReplicas rather than just being slow. ## The causation chain 1. A follower broker gets slow — GC pauses, disk I/O stalls, network saturation, or general overload — so its **ReplicaFetcher** threads pull from the leader more slowly. 2. That follower's replication **lag** grows (its log-end offset falls behind the leader's). 3. Because the producer used **acks=all**, the leader can't ack until that follower (still in ISR) catches up → the produce request waits in purgatory → **Produce RemoteTimeMs** rises → producer-observed p99 spikes. 4. If the follower lags beyond `replica.lag.time.max.ms`, the leader **removes it from the ISR** → **IsrShrinksPerSec** ticks up, the partition may become **under-replicated** (UnderReplicatedPartitions > 0). Now acks=all waits only on the remaining (faster) ISR members, so latency may *recover* — but durability dropped and you risk falling below min.insync.replicas. 5. The follower recovers, catches up, rejoins ISR (**IsrExpandsPerSec**) — and the cycle can repeat, producing **ISR flapping** and intermittent p99 spikes. ## Metrics to correlate (and what each proves) - **kafka.network RequestMetrics RemoteTimeMs (Produce, p99/p999)** on the leader — confirms the latency is replication-wait, not local. - **kafka.server:type=ReplicaManager,name=UnderReplicatedPartitions** and **IsrShrinksPerSec / IsrExpandsPerSec** — confirm ISR churn and which partitions. - **Per-replica lag** — replica fetcher lag / max lag; identifies *which follower* is behind. - **The suspect follower's local health**: GC pause time, RequestHandlerAvgIdlePercent / NetworkProcessorAvgIdlePercent, disk await/util, NIC saturation, and its FetchFollower LocalTimeMs. A single broker with bad local signals that matches the lagging follower pins the culprit. - **num.replica.fetchers** and replica fetch config — too few fetcher threads can itself cause lag under high partition counts. ## Edge cases & nuances - **Latency 'recovering' after ISR shrink is a trap**: the spike stops because the slow replica was dropped, but you've silently lost a replica's worth of durability. Alert on IsrShrinks, not just latency. - **min.insync.replicas interaction**: if shrink takes ISR below min.insync.replicas, acks=all produces start *failing*, not just slowing — a much louder symptom. - **Network vs broker**: a slow link between leader and one follower (not the follower's CPU) can cause the same lag; check network between the specific pair. - **Leader-side cause**: occasionally the *leader* is slow serving FetchFollower requests, making all its followers lag — distinguish by whether one follower or all followers of a leader are behind. - Don't confuse this with consumer lag; this is *replica* (follower) lag, an internal replication metric.

  • After ISR shrinks, the producer p99 recovers. Why is that not actually good news?
    Latency recovered only because the slow replica was ejected from the ISR, so acks=all now waits on fewer replicas. You've lost a replica's durability and moved closer to min.insync.replicas; if you fall below it, acks=all produces will start failing. Alert on ISR shrink, not just latency.
  • How would you tell whether the lag is caused by the follower being slow versus the leader being slow to serve fetches?
    If only one follower of a leader lags, the follower is likely slow (check its GC/disk/network). If all followers of the same leader lag together, suspect the leader's FetchFollower handling, network egress, or disk read path. Per-replica lag plus per-broker local metrics separate the two.
  • What config governs how long a follower can lag before being removed from the ISR?
    replica.lag.time.max.ms (default 30000 ms). A follower that hasn't caught up to the leader's log-end offset within that window is dropped from the in-sync replica set.

saying these in an interview costs you the question

  • Saying acks=all waits on all replicas (it waits on all replicas in the ISR, which can be a shrunken subset)
  • Treating post-shrink latency recovery as a resolution rather than a durability loss
  • Confusing follower replication lag with consumer-group lag
  • Forgetting min.insync.replicas, so missing that shrink can turn slow produces into failed produces

context