Explain why setting min.insync.replicas equal to the replication factor hurts availability, and how to pick the value to hit a specific durability/availability target.
answer
- Survive MISR-1 losses
- Available while ISR >= MISR
- MISR=RF => any blip stops writes
- MISR = RF-1 is the sweet spot
- watch UnderMinIsr metric
basics
~20 sIf min.insync.replicas equals RF, every replica must be in sync to accept writes, so losing or lagging just one broker stops all writes. Setting it to RF-1 (e.g. 2 with RF=3) keeps you durable while tolerating one failure.
solid answer
~50 smin.insync.replicas (MISR) is the minimum ISR size for an acks=all write to be accepted. The general rule: an acknowledged write survives up to (MISR - 1) simultaneous replica losses, and writes stay available as long as at least MISR replicas are in sync. So MISR trades the two off. With RF=3: MISR=2 survives one broker loss for durability AND stays writable through one failure — the sweet spot. MISR=3 (equal to RF) maximizes durability (acknowledged write on all 3) but any single replica leaving the ISR — a rolling restart, GC pause, slow disk, or failure — drops the ISR below MISR and halts all producers with NotEnoughReplicas. To hit a target: choose RF for how many failures the data must survive, then set MISR = RF - 1 for the standard 'durable and one-fault-tolerant' point, or MISR = RF if you must never acknowledge a write that isn't on every replica and can accept write outages.
go deeper
Know that MISR=RF means one slow/restarting broker stops writes; MISR=RF-1 is safer.
State the two formulas: survive MISR-1 losses, available while ISR>=MISR, and pick RF/MISR for a failure budget.
Connect to operational realities (rolling deploys, GC, replica.lag.time.max.ms) and monitoring UnderMinIsrPartitionCount.
Define org standards per workload tier and reason about the cost/availability frontier across many topics and DCs.
**Setup.** With `acks=all`, a produce request is acknowledged only after every member of the current ISR has written the record. `min.insync.replicas` (MISR) gates whether the write is even allowed: if the live ISR is smaller than MISR, the leader returns `NotEnoughReplicasException` and the producer cannot make progress (it retries until the ISR recovers or `delivery.timeout.ms` expires). **The two guarantees MISR controls:** 1. *Durability:* because a write is acknowledged only when at least MISR replicas hold it, an acknowledged record can survive up to MISR-1 of those replicas dying simultaneously and still exist somewhere. 2. *Write availability:* the partition accepts writes only while >= MISR replicas are in the ISR. The number of replicas you can lose before writes stall is RF - MISR. **Why MISR = RF is fragile.** Suppose RF=3, MISR=3. All three must be in the ISR to write. But the ISR shrinks for many benign reasons: a follower GC-pauses past `replica.lag.time.max.ms` (30s default), a broker is restarted during a rolling deploy, a disk gets slow, or the network blips. The moment the ISR drops to 2 — even with zero data loss — producers using acks=all get NotEnoughReplicas and the partition is read-only for writes until the follower rejoins. In a cluster of any size this happens routinely, so MISR=RF causes frequent self-inflicted write outages. **Why MISR = RF - 1 is the standard.** RF=3, MISR=2: you can lose one replica (ISR drops from 3 to 2) and still write, and every acknowledged write is on >= 2 brokers, so a subsequent single failure cannot lose it. You only lose write availability if a *second* replica also drops (ISR -> 1 < MISR). This 'tolerate one fault' point is what most production guidance recommends. **Picking values for a target.** Decide failure budget first: - 'Survive one broker loss, stay writable through it, never lose acknowledged data': RF=3, MISR=2, acks=all. - 'Survive two simultaneous losses': RF=5, MISR=3 (tolerates losing 2 replicas and still writes; acknowledged data on >= 3). - 'Strictest durability, write outages acceptable': MISR=RF. **Interactions / edge cases.** MISR is meaningless without acks=all. RF must be >= MISR or the topic can never accept writes. MISR can be set at broker level (default) and overridden per topic. Also note: when the ISR is *exactly* MISR, the partition is durable but has zero availability headroom — a single further failure stops writes, so monitoring `UnderMinIsrPartitionCount` is essential.
- What error does a producer see when the ISR falls below min.insync.replicas, and is the data lost?The leader rejects with NotEnoughReplicasException (or NotEnoughReplicasAfterAppendException). No data is lost or acknowledged — the producer retries until the ISR recovers or delivery.timeout.ms expires, then fails the send.
- For a topic that must survive two simultaneous broker failures with no acknowledged-data loss, what RF and MISR would you choose?RF=5 with min.insync.replicas=3, acks=all. Acknowledged writes live on >=3 replicas (survive losing 2), and writes remain available while at least 3 are in sync.
saying these in an interview costs you the question
- Saying MISR=RF is 'safest' without noting it destroys write availability on any single ISR shrink.
- Claiming a NotEnoughReplicas rejection means data loss (it means the write was refused, not lost).
- Forgetting that ISR shrinks for benign reasons (rolling restarts, GC) not just hard failures.