skip to content

What do IsrShrinksPerSec and IsrExpandsPerSec measure, and what does a high or flapping rate tell you?

level: seniorimportance: should knowfreq 40%

answer

  1. shrink = follower removed (lagged > replica.lag.time.max.ms)
  2. expand = follower caught up, rejoined
  3. meters -> watch OneMinuteRate
  4. flapping = overloaded/unstable broker (disk/GC/network/fetchers)
  5. correlates with UnderReplicatedPartitions

basics

~20 s

They are rate meters counting how often replicas leave the in-sync replica set (shrink) or rejoin it (expand) per second. Occasional events are normal; a high or oscillating rate means replicas are repeatedly falling behind and catching up — a sign of an overloaded or unstable broker.

solid answer

~40 s

IsrShrinksPerSec (kafka.server:type=ReplicaManager,name=IsrShrinksPerSec) and IsrExpandsPerSec (kafka.server:type=ReplicaManager,name=IsrExpandsPerSec) are meter metrics: a shrink event fires when the leader removes a follower from the ISR because it fell behind beyond replica.lag.time.max.ms (default 30s); an expand fires when a follower catches up and rejoins. A few during normal operation or a restart is fine. A sustained high rate, or shrink-then-expand flapping, means replicas keep dropping out and rejoining — caused by an overloaded follower (disk/network saturation), long GC pauses, undersized replica fetcher threads (num.replica.fetchers), or network instability. Flapping ISR correlates with rising UnderReplicatedPartitions and degrades durability because the effective ISR shrinks. You read these as rates (OneMinuteRate) and alert on a threshold sustained over time, then drill into the affected broker's resource and GC metrics.

go deeper

for a junior

Knows shrink = replica fell behind/left ISR, expand = it caught up; occasional is fine.

for a middle

Connects to replica.lag.time.max.ms and reads the OneMinuteRate, knows flapping is bad.

for a senior

Diagnoses flapping causes (disk/GC/network/fetchers) and correlates with UnderReplicated and resource metrics.

for a principal

Defines durability-stability SLOs, fetcher/JVM tuning standards, and capacity policies that keep ISR stable under burst.

## ISR recap The **ISR (In-Sync Replica set)** is the set of replicas currently caught up to the partition leader. A follower stays in the ISR as long as it has fetched up to the leader's log end offset within **`replica.lag.time.max.ms`** (default **30000 ms**). If a follower hasn't caught up within that window, the leader **shrinks** the ISR by removing it. When that follower catches up again, the leader **expands** the ISR by re-adding it. ## The metrics - `kafka.server:type=ReplicaManager,name=IsrShrinksPerSec` — a **meter** (rate) counting ISR-shrink events per second. Exposes `Count`, `OneMinuteRate`, `FiveMinuteRate`, `MeanRate`. - `kafka.server:type=ReplicaManager,name=IsrExpandsPerSec` — the matching meter for expand events. Because they're meters you usually watch the **OneMinuteRate** and alert on a sustained elevated value rather than a single spike. ## Reading the signal - **Occasional shrink/expand:** normal — happens during rolling restarts, brief load bursts, or reassignments. - **Sustained high rate / flapping (shrink quickly followed by expand, repeating):** replicas are **chronically falling behind and catching up**. Root causes: - An **overloaded follower broker**: disk I/O saturation (can't write fetched data fast enough) or network saturation. - **Long GC pauses** on a broker stalling the fetcher threads. - **Too few replica fetcher threads** (`num.replica.fetchers`) for the partition count / throughput. - **Network instability / packet loss** between brokers. - Throughput spikes that push followers past `replica.lag.time.max.ms` momentarily. ## Why it matters Every shrink reduces the effective ISR, which: - Increases `UnderReplicatedPartitions` while shrunk. - Reduces durability margin; if it drops below `min.insync.replicas`, **acks=all producers start failing**. - Flapping also adds controller/metadata churn. So IsrShrinks/Expands are an **early-warning, root-cause-flavored** signal: UnderReplicatedPartitions tells you *that* you're under-replicated; the shrink/expand rate tells you it's happening *repeatedly and dynamically*, pointing at an unstable or overloaded broker rather than a cleanly-down one. ## Remediation path 1. Identify which broker is the lagging follower (per-partition replica state, or correlate with that broker's CPU/disk/network/GC). 2. If disk/network bound: rebalance partitions off the broker, add capacity, or tune. 3. If GC bound: tune the JVM / heap. 4. If fetcher-bound: raise `num.replica.fetchers`. 5. If chronic momentary lag under bursty load: consider whether `replica.lag.time.max.ms` is appropriate, but prefer fixing the underlying resource constraint over loosening the lag window. ## Edge cases - A single clean broker shutdown causes a one-time shrink (not flapping) — distinguish a step change from oscillation. - Loosening `replica.lag.time.max.ms` hides flapping but masks a real problem and lets followers fall further behind before removal, weakening durability guarantees. - These are leader-reported, so aggregate across brokers and correlate with the follower's own resource metrics.

  • You see IsrShrinksPerSec and IsrExpandsPerSec both elevated and oscillating on one broker. Walk through your investigation.
    Flapping means a follower keeps falling outside replica.lag.time.max.ms and recovering. Identify the lagging follower, then inspect that broker's disk utilization, network throughput, GC pause times, and replica fetcher saturation (num.replica.fetchers). Likely causes are disk/network saturation or long GC. Remediate by rebalancing partitions off it, adding fetchers, or tuning the JVM — not by simply raising replica.lag.time.max.ms, which masks the issue.
  • How do IsrShrinks relate to UnderReplicatedPartitions?
    When the ISR shrinks below the replication factor, those partitions become under-replicated, so a high shrink rate drives UnderReplicatedPartitions up. Shrinks/expands reveal the dynamic flapping behavior; UnderReplicatedPartitions shows the current static count of affected partitions.

saying these in an interview costs you the question

  • Saying any ISR shrink is an emergency (occasional shrinks are normal)
  • Treating a single shutdown step-change as flapping
  • Recommending to just raise replica.lag.time.max.ms to silence it (masks the real cause)
  • Confusing ISR shrink/expand with consumer rebalances — different mechanism

context