How do ISR shrink dynamics interact with min.insync.replicas, acks=all, and unclean.leader.election.enable to shape the durability vs. availability tradeoff?
answer
- RF=3, min.insync.replicas=2, acks=all = gold standard
- ISR < min.insync → NotEnoughReplicasException
- min.insync only bites with acks=all
- unclean=false → offline; true → data loss
- Watch UnderMinIsrPartitionCount
basics
~20 sWhen the ISR shrinks, fewer replicas hold committed data. min.insync.replicas sets the floor: below it, acks=all writes are rejected (availability lost to protect durability). unclean.leader.election lets an out-of-sync replica become leader if ISR is empty — restoring availability but risking data loss.
solid answer
~40 sThese four settings jointly tune the CAP-style tradeoff. With `acks=all`, a write commits only when every current ISR member has it. `min.insync.replicas` (e.g., 2) is the minimum ISR size for such writes to be accepted; if ISR shrinks below it, producers get `NotEnoughReplicasException` — Kafka sacrifices write availability to guarantee that no acknowledged write is under-replicated. The classic safe config is RF=3, min.insync.replicas=2, acks=all: tolerates one broker loss with zero data loss and continued writes. `unclean.leader.election.enable=false` (the safe default) means if the ISR becomes empty, the partition goes offline rather than electing a stale replica — preserving durability at the cost of availability. Setting it true lets a lagging out-of-sync replica become leader, restoring availability but silently dropping records that were committed only on the failed leader.
go deeper
Know that min.insync.replicas plus acks=all means a write needs that many in-sync copies.
Explain the RF=3/min.insync=2/acks=all gold standard and why min.insync=RF is brittle.
Detail NotEnoughReplicas errors, high-watermark effects, and the unclean-election tradeoff.
Design a coherent durability/availability policy across RF, min.ISR, acks, unclean election, AZ placement, and alerting.
## The four interacting levers ### 1. acks (producer-side) - **acks=0:** fire-and-forget; no durability tie to ISR. - **acks=1:** leader-only ack; data on followers not guaranteed — leader crash can lose acknowledged writes. - **acks=all (a.k.a. acks=-1):** the leader acknowledges only after **every member of the current ISR** has appended the record. This is where ISR dynamics directly govern durability. ### 2. min.insync.replicas (topic/broker-side) The **minimum number of in-sync replicas** that must be present for an `acks=all` write to be accepted. It only takes effect with `acks=all`. Mechanics: - If the current ISR size ≥ `min.insync.replicas`, the write proceeds. - If ISR has shrunk *below* it, the leader rejects the produce with **`NotEnoughReplicasException`** (before append) or **`NotEnoughReplicasAfterAppendException`** (if it shrank mid-append). Consumers can still read committed data; only new acks=all writes fail. **This is a deliberate availability sacrifice to protect durability:** rather than accept a write that would survive on too few replicas, Kafka refuses it. ### 3. Replication factor (RF) The number of assigned replicas. The canonical durable setup is **RF=3, min.insync.replicas=2, acks=all**: - Steady state ISR = 3. - Lose one broker → ISR = 2, still ≥ min.insync.replicas → writes continue, **zero acknowledged data loss**. - Lose a second → ISR = 1 < 2 → writes rejected, but committed data is still safe on the survivor. Anti-pattern: **RF=3, min.insync.replicas=3** — any single follower hiccup that shrinks ISR to 2 halts all writes. Too brittle. And **min.insync.replicas=1 with acks=all** gives no more durability than acks=1. ### 4. unclean.leader.election.enable (broker/topic-side) When the **entire ISR is lost** (e.g., all in-sync replicas crash): - **`false` (default, safe):** the partition becomes **offline** — no leader, no reads, no writes — until an in-sync replica returns. Durability preserved; availability sacrificed. - **`true`:** the controller may elect a leader from a **non-ISR** replica that is behind. Availability restored immediately, but any records that were committed only on the now-dead in-sync replicas are **permanently lost**, and consumers may see the log truncate (offsets regress relative to what acks=all had promised). ## How a shrink cascades through durability 1. A follower exceeds `replica.lag.time.max.ms` → leader shrinks ISR. 2. The **high watermark** is recomputed over the smaller ISR — committing can actually speed up (fewer acks needed) but redundancy drops. 3. If the shrink crosses `min.insync.replicas`, acks=all writes start failing — a self-protecting backpressure that surfaces the durability risk to producers instead of hiding it. 4. If shrink reaches an empty ISR, the unclean-election policy decides offline-vs-dataloss. ## Architectural guidance - **Default durable profile:** RF=3, min.insync.replicas=2, acks=all, unclean.leader.election=false. Survives one failure with no data loss and no write outage. - **Tune `replica.lag.time.max.ms`** to avoid flapping followers crossing the min.ISR floor under transient load. - **Monitor** `UnderMinIsrPartitionCount` (partitions currently rejecting writes) and `UnderReplicatedPartitions`; alert before the floor is hit. - **Multi-AZ:** spread replicas across availability zones so a zone failure removes at most one replica per partition, keeping ISR ≥ min.insync.replicas. - **Idempotent/transactional producers** require acks=all and benefit from min.insync.replicas=2 for exactly-once durability guarantees. The overarching principle: ISR is the *consistency boundary*; min.insync.replicas is the *durability floor*; unclean election is the *availability escape hatch*. Configure them as one coherent policy, not in isolation.
- Why is RF=3 with min.insync.replicas=3 considered a brittle anti-pattern?Because any transient event that shrinks the ISR to 2 (a follower GC pause, brief network blip) immediately violates the floor and halts all acks=all writes. There is zero tolerance for follower hiccups, trading away availability with no durability gain over min.insync.replicas=2.
- What does setting min.insync.replicas=1 with acks=all actually guarantee?Effectively nothing more than acks=1. With ISR floor of 1, the leader alone satisfies the requirement, so an acknowledged write can exist on only the leader and be lost if it crashes before any follower replicates it.
- When does unclean.leader.election=true cause visible data loss?When the entire ISR is lost and an out-of-sync replica is elected leader. Records committed only on the failed in-sync replicas are gone; the new leader's log is shorter, so the partition's high watermark effectively regresses and consumers may re-read or never see those offsets.
saying these in an interview costs you the question
- Saying min.insync.replicas applies to acks=1 or acks=0 — it only constrains acks=all.
- Recommending min.insync.replicas equal to RF as a 'safe' default (it's brittle).
- Claiming unclean leader election never loses data.
- Believing a shrunk ISR blocks reads of already-committed data — only new acks=all writes are affected.