Walk through what happens when the ISR shrinks below min.insync.replicas while producers are writing with acks=all.
answer
- follower lags past replica.lag.time.max.ms -> dropped from ISR
- before append: NotEnoughReplicasException
- after append: NotEnoughReplicasAfterAppendException
- both retriable -> bounded by delivery.timeout.ms
- reads OK, writes blocked, self-heals
basics
~10 sThe leader rejects new acks=all writes with NotEnoughReplicasException, so producers fail and retry. The partition becomes read-still-available but write-blocked until enough followers rejoin the ISR.
solid answer
~50 sWhen a follower falls behind (lagging past replica.lag.time.max.ms) or a broker dies, the leader removes it from the ISR. If the ISR shrinks below min.insync.replicas, the leader stops accepting acks=all writes: it rejects the produce request before appending with NotEnoughReplicasException. If the ISR was sufficient at append time but drops below the floor before the write is fully replicated, the producer instead gets NotEnoughReplicasAfterAppendException — the record may be on the leader's log but is not considered committed. Both are retriable on the producer (they extend RetriableException), so a producer with retries configured will keep trying; if the floor isn't restored within delivery.timeout.ms, delivery fails. Consumers can still read previously committed records — only writes are blocked. The partition recovers automatically when lagging followers catch up and the controller re-expands the ISR above the floor.
go deeper
Know that breaching the floor blocks writes with an error and reads still work.
Distinguish the before-append vs after-append exceptions and that both are retriable.
Explain the availability-vs-durability tradeoff and the delivery.timeout.ms retry window behavior.
Design alerting on ISR shrink, reason about blast radius of mixed acks producers, and set recovery runbooks.
## Setup Assume the standard hardened config: **RF=3, min.insync.replicas=2, acks=all**. The partition has a leader plus two followers, ISR size 3. ## How the ISR shrinks A follower is dropped from the ISR when it stops fetching or lags behind the leader's log end offset for longer than **`replica.lag.time.max.ms`** (default 30000 ms). Causes: the follower's broker crashed, a network partition, GC pause, disk saturation, or the broker being rolling-restarted. The leader (with the controller) shrinks the ISR. - ISR 3 -> 2: still **>= min.insync.replicas (2)**. Writes continue normally. You have lost your redundancy buffer but durability is intact. - ISR 2 -> 1: now **< 2**. The floor is breached. ## What the producer sees Once ISR < min.insync.replicas, the leader refuses `acks=all` writes: - **`NotEnoughReplicasException`** — thrown *before* the append, when the leader checks the ISR size and finds it too small. Nothing is written. - **`NotEnoughReplicasAfterAppendException`** — thrown when the ISR was adequate at the moment of append but shrank below the floor before the write committed across the ISR. The bytes may sit in the leader's local log but the offset is **not committed** (not exposed to consumers, not durable per the contract). Both extend `RetriableException`. With default producer retries (effectively `Integer.MAX_VALUE` retries bounded by `delivery.timeout.ms`, default 120000 ms), the producer keeps retrying. If the ISR is not restored above the floor within that window, the send's future completes exceptionally and the application sees the failure. ## Availability impact This is an intentional **availability sacrifice to preserve durability**. The partition is effectively **write-unavailable** for `acks=all` producers, while: - `acks=0`/`acks=1` producers to the same partition would still succeed (they don't honor the floor) — usually you don't mix them. - Consumers keep reading committed data; reads are unaffected. ## Recovery When a lagging follower catches up (re-fetches to within the lag window) or a downed broker restarts and re-replicates, the controller expands the ISR back to >= min.insync.replicas. The leader resumes accepting writes; queued producer retries then succeed. No manual intervention is needed for a transient outage — the system self-heals. ## Why this is the right default behavior Refusing the write surfaces the problem loudly (producer errors, alertable) instead of silently degrading to single-copy durability. You trade a window of write unavailability for the guarantee that any acknowledged write survived to at least `min.insync.replicas` brokers.
- What is the difference between NotEnoughReplicasException and NotEnoughReplicasAfterAppendException?The first is thrown before the record is appended (the leader checks the ISR size up front and finds it below the floor, nothing is written). The second is thrown when the ISR was adequate at append time but shrank below the floor before the write committed across the ISR — the bytes may be in the leader's local log but the offset is not committed.
- Are these exceptions fatal to the producer or retriable?Retriable — both extend RetriableException. The producer keeps retrying until the ISR recovers or delivery.timeout.ms (default 120s) elapses, after which the send future fails.
saying these in an interview costs you the question
- Saying consumers also get blocked — only acks=all writes are blocked; reads of committed data continue.
- Treating NotEnoughReplicasException as a fatal/non-retriable error.
- Claiming the broker silently drops or accepts the write at lower durability instead of rejecting it.
- Assuming manual recovery is required — the ISR re-expands automatically when followers catch up.