Walk through what happens to a Kafka topic (RF=3 across 3 AZs, min.insync.replicas=2, acks=all) when one entire AZ fails. What survives, and what would break this guarantee?
answer
- 1 replica lost per partition, 2 survive
- ISR 3→2, still ≥ min.ISR=2 → writes continue
- controller elects new leader from in-sync survivor
- under-replicated but no data loss
- breakers: RF<racks, min.ISR=3, unclean election, mislabeled rack
basics
~20 sEach partition keeps 2 of its 3 replicas because they're in different AZs. Leaders that were in the dead AZ fail over to surviving replicas, ISR drops to 2 which still meets min.insync.replicas=2, so acks=all producers and consumers keep working with no data loss.
solid answer
~50 sWith RF=3 spread one-replica-per-AZ, losing a full AZ removes exactly one replica from every partition. For partitions whose leader was in the dead AZ, the controller elects a new leader from the two surviving in-sync replicas. The ISR shrinks from 3 to 2 — still ≥ min.insync.replicas=2 — so acks=all producers continue to be acknowledged and there's no data loss or write stall. Consumers keep reading from surviving leaders/followers. The cluster runs under-replicated (no new third replica until the AZ returns or you reassign). Things that break this: RF<3 or RF spread across <3 racks (some partition loses 2 replicas → below ISR floor → acks=all stalls); brokers not actually mapped to distinct AZs (mislabeled broker.rack); or unclean.leader.election.enable=true causing data loss if only an out-of-sync replica survives. So the guarantee depends on correct broker.rack + RF=racks + min.insync.replicas=RF-1.
go deeper
Know RF=3 across 3 AZs means losing one AZ still leaves 2 copies, so data survives.
Add the ISR 3→2 vs min.ISR=2 interaction and that writes keep flowing.
Walk the full failover sequence and enumerate the misconfigs (RF<racks, min.ISR=RF, unclean election) that break it.
Generalize to an F-failure durability formula, reason about network partitions vs clean failures, and set cluster-wide defaults for the guarantee.
## Setup recap - **RF=3, 3 AZs, rack-aware**: each partition has one replica in each AZ. One leader + two followers, all in ISR. - **min.insync.replicas (min.ISR)=2**: the durability floor — `acks=all` writes are only acknowledged if at least 2 replicas (including the leader) have the record. - **acks=all**: the producer waits for all *in-sync* replicas (bounded below by min.ISR) to confirm before considering a write durable. ## Failure sequence when AZ-2 dies 1. **Brokers in AZ-2 go down.** Every partition that had a replica there loses exactly that one replica. Two replicas remain (in AZ-1 and AZ-3). 2. **Leadership.** Partitions whose leader was in AZ-2 are now leaderless. The **controller** detects the brokers left the cluster and **elects a new leader** from the surviving ISR members (AZ-1 or AZ-3). This is a *clean* election — the new leader was in-sync, so no committed data is lost. 3. **ISR shrinks 3→2.** Each affected partition's ISR now has 2 members. Because **2 ≥ min.ISR (2)**, the partition is still *writable*: `acks=all` producers keep getting acknowledgments. No stall. 4. **Under-replicated.** The cluster reports `UnderReplicatedPartitions > 0`. There's no third replica until AZ-2 brokers return (they rejoin and catch up) or you reassign to other brokers. 5. **Consumers** keep reading; fetch-from-follower consumers whose preferred replica was in AZ-2 fall back to the leader or another in-AZ follower. **Net result: no data loss, no write outage, just reduced redundancy.** ## What breaks the guarantee - **RF < 3 or fewer than 3 racks.** If a partition had 2 replicas in the failed AZ, only 1 survives → ISR=1 < min.ISR=2 → the partition becomes **read-only for acks=all** (producers stall) until a replica recovers. This is the classic anti-pattern. - **min.insync.replicas misconfigured.** With min.ISR=3 and RF=3, losing one replica drops ISR to 2 < 3 → writes stall even though data is safe. min.ISR should be RF-1 to tolerate one failure. - **Mislabeled broker.rack.** If two replicas of a partition are physically in the same AZ (because labels lied or RF>racks), the math above fails silently. - **Unclean leader election.** If `unclean.leader.election.enable=true` and the only survivor is *out of sync*, Kafka may elect it as leader and **lose committed data**. Keep it false for durability. - **Network partition vs failure.** A partial partition (not a clean AZ death) can cause flapping ISR; quorum/controller behavior differs and may temporarily reduce availability. ## The durability formula To survive **F** simultaneous rack failures with continued `acks=all` writes: `RF ≥ (F+1) × ?` — practically: spread across at least `F + min.ISR` racks so a surviving ISR ≥ min.ISR remains. The common one-AZ-tolerant recipe is **RF=3, racks=3, min.ISR=2, unclean.leader.election.enable=false**.
- Why must min.insync.replicas be 2 (not 3) for this to tolerate an AZ loss?With RF=3 and min.ISR=3, losing any single replica drops ISR to 2 < 3, which stalls acks=all producers even though the data is perfectly safe. Setting min.ISR=RF-1=2 keeps writes flowing through a single-AZ failure while still requiring a 2-replica quorum for durability.
- What role does unclean.leader.election.enable play here?If left false (recommended), Kafka only elects in-sync replicas as leaders, guaranteeing no committed data is lost. If true, an out-of-sync replica can become leader when no in-sync one survives, trading durability for availability and potentially losing data.
saying these in an interview costs you the question
- Claiming the cluster goes down or loses data on a single AZ loss when RF=3/min.ISR=2 — it survives cleanly.
- Setting min.insync.replicas equal to RF, which stalls writes on any single replica loss.
- Assuming rack-awareness alone guarantees survival without correct min.ISR and unclean.leader.election settings.
- Ignoring that RF spread across <3 racks lets one rack hold a majority of replicas.