Walk through what happens, step by step, when a follower falls out of the ISR and later rejoins. Who decides, and how is the change propagated?
answer
- Leader decides, controller persists
- AlterPartition (KRaft) vs ZooKeeper znode (legacy)
- Leader epoch fences stale proposals
- Rejoin = catch up to moving LEO
- IsrShrinks/ExpandsPerSec, UnderReplicatedPartitions
basics
~20 sThe leader detects a follower exceeding replica.lag.time.max.ms, shrinks the ISR, and persists the new ISR via the controller (AlterPartition in KRaft, ZooKeeper in older versions). When the follower catches back up to the leader's LEO, the leader expands the ISR and re-propagates it.
solid answer
~40 sThe partition leader owns ISR decisions. It periodically scans followers; any follower whose last caught-up-to-LEO time exceeds `replica.lag.time.max.ms` is dropped. The leader then submits the proposed ISR change to the controller — an `AlterPartition` request in KRaft mode (or a ZooKeeper write to the partition state znode in ZK mode). The controller validates and commits it, bumping the partition leader epoch and broadcasting the updated metadata. While shrunk, the high watermark is computed over the smaller ISR, so acks=all writes commit faster but with weaker redundancy, and min.insync.replicas may start rejecting writes. When the lagging follower resumes fetching and reaches the leader's LEO, the leader proposes an ISR *expansion* the same way. Metrics like IsrShrinksPerSec, IsrExpandsPerSec, and UnderReplicatedPartitions expose this churn.
go deeper
Know that the leader removes lagging followers and adds them back when caught up.
Describe shrink/expand triggers and the high-watermark impact.
Detail the leader→controller propagation path, AlterPartition vs ZooKeeper, and relevant metrics.
Reason about leader-epoch fencing, flapping mitigation, and ISR churn impact on durability/availability SLOs.
## Who owns ISR membership: the leader The **partition leader** is the authority on which followers are in-sync. It is the only replica that sees every follower's fetch progress, so it is the natural decision-maker. The **controller** (cluster coordinator) is the authority on *persisting and broadcasting* the ISR, not on computing it. ## Step-by-step: shrink (follower falls out) 1. **Tracking.** For each follower, the leader records the last time that follower's fetch reached the leader's **log end offset (LEO)** — the caught-up timestamp. 2. **Periodic check.** The leader runs a scheduled task (default every `replica.lag.time.max.ms / 2`) checking each follower's timestamp against `now`. 3. **Decision.** A follower whose gap exceeds `replica.lag.time.max.ms` (default 30s) is marked out of sync. 4. **Propose the new ISR.** The leader sends an **`AlterPartition`** request to the **controller** (this is the KRaft path; in pre-KRaft/ZooKeeper deployments the leader wrote the new ISR to the partition's state znode in ZooKeeper). 5. **Controller commits.** The controller validates the request against the current **leader epoch** (a monotonically increasing fencing number that rejects stale proposals from a deposed leader), commits the new ISR to the cluster metadata log, and propagates updated `LeaderAndIsr`/metadata to brokers. 6. **High watermark recomputed.** The leader now computes the high watermark over the *smaller* ISR. Consequences: acks=all requests commit once the remaining (fewer) ISR members ack — potentially faster but less redundant; if ISR size drops below `min.insync.replicas`, acks=all producers receive `NotEnoughReplicasException` / `NotEnoughReplicasAfterAppendException`. ## Step-by-step: expand (follower rejoins) 1. The ejected follower keeps (or resumes) fetching from the leader. 2. It replicates the backlog until its fetch finally requests the leader's **current** LEO — i.e., it is fully caught up. 3. The leader proposes an **ISR expansion** via the same `AlterPartition` path; the controller commits it and bumps metadata. 4. The high watermark and durability guarantees return to full redundancy. Important: a rejoining follower must catch up to the *moving* LEO, not a frozen snapshot — under heavy write load this can take a while, and the follower stays out until it converges. ## Fencing with leader epoch Every ISR change is tagged with the **leader epoch**. If a stale leader (e.g., one that was network-partitioned and superseded) tries to alter the ISR, the controller rejects the request because its epoch is behind. This prevents split-brain ISR corruption. ## Observability Key JMX metrics: - **`IsrShrinksPerSec`** / **`IsrExpandsPerSec`** — rate of ISR membership changes; sustained nonzero values signal flapping followers. - **`UnderReplicatedPartitions`** — partitions where ISR size < replication factor; should be 0 in steady state. - **`UnderMinIsrPartitionCount`** — partitions whose ISR is below min.insync.replicas (writes failing). ## Edge cases - **Flapping:** a follower hovering near the threshold can repeatedly shrink/expand, churning metadata. Tune `replica.lag.time.max.ms` up or fix the underlying slowness (disk, network, GC). - **Leader failure during shrink:** if the leader dies mid-decision, the controller elects a new leader from the *committed* ISR, never from a follower that was never in-sync (assuming unclean election is disabled).
- In KRaft mode, how does a leader publish an ISR change?It sends an AlterPartition request to the active controller, which validates the leader epoch, commits the new ISR to the metadata log, and broadcasts updated LeaderAndIsr metadata. The leader does not write to ZooKeeper, which KRaft removes entirely.
- What stops a stale, deposed leader from corrupting the ISR?The leader epoch. Every AlterPartition/ISR update carries the proposing leader's epoch; the controller rejects any proposal whose epoch is behind the current one, fencing out a partitioned old leader.
saying these in an interview costs you the question
- Saying the controller decides ISR membership — the leader decides; the controller persists/broadcasts.
- Claiming KRaft still writes ISR to ZooKeeper.
- Forgetting the leader-epoch fencing that prevents split-brain.
- Assuming a rejoining follower catches up to a fixed offset rather than the moving LEO.