skip to content

Walk through what happens to a Kafka topic (RF=3 across 3 AZs, min.insync.replicas=2, acks=all) when one entire AZ fails. What survives, and what would break this guarantee?

level: seniorimportance: must knowfreq 48%

answer

  1. 1 replica lost per partition, 2 survive
  2. ISR 3→2, still ≥ min.ISR=2 → writes continue
  3. controller elects new leader from in-sync survivor
  4. under-replicated but no data loss
  5. breakers: RF<racks, min.ISR=3, unclean election, mislabeled rack

basics

~20 s

Each partition keeps 2 of its 3 replicas because they're in different AZs. Leaders that were in the dead AZ fail over to surviving replicas, ISR drops to 2 which still meets min.insync.replicas=2, so acks=all producers and consumers keep working with no data loss.

solid answer

~50 s

With RF=3 spread one-replica-per-AZ, losing a full AZ removes exactly one replica from every partition. For partitions whose leader was in the dead AZ, the controller elects a new leader from the two surviving in-sync replicas. The ISR shrinks from 3 to 2 — still ≥ min.insync.replicas=2 — so acks=all producers continue to be acknowledged and there's no data loss or write stall. Consumers keep reading from surviving leaders/followers. The cluster runs under-replicated (no new third replica until the AZ returns or you reassign). Things that break this: RF<3 or RF spread across <3 racks (some partition loses 2 replicas → below ISR floor → acks=all stalls); brokers not actually mapped to distinct AZs (mislabeled broker.rack); or unclean.leader.election.enable=true causing data loss if only an out-of-sync replica survives. So the guarantee depends on correct broker.rack + RF=racks + min.insync.replicas=RF-1.

go deeper

for a junior

Know RF=3 across 3 AZs means losing one AZ still leaves 2 copies, so data survives.

for a middle

Add the ISR 3→2 vs min.ISR=2 interaction and that writes keep flowing.

for a senior

Walk the full failover sequence and enumerate the misconfigs (RF<racks, min.ISR=RF, unclean election) that break it.

for a principal

Generalize to an F-failure durability formula, reason about network partitions vs clean failures, and set cluster-wide defaults for the guarantee.

## Setup recap - **RF=3, 3 AZs, rack-aware**: each partition has one replica in each AZ. One leader + two followers, all in ISR. - **min.insync.replicas (min.ISR)=2**: the durability floor — `acks=all` writes are only acknowledged if at least 2 replicas (including the leader) have the record. - **acks=all**: the producer waits for all *in-sync* replicas (bounded below by min.ISR) to confirm before considering a write durable. ## Failure sequence when AZ-2 dies 1. **Brokers in AZ-2 go down.** Every partition that had a replica there loses exactly that one replica. Two replicas remain (in AZ-1 and AZ-3). 2. **Leadership.** Partitions whose leader was in AZ-2 are now leaderless. The **controller** detects the brokers left the cluster and **elects a new leader** from the surviving ISR members (AZ-1 or AZ-3). This is a *clean* election — the new leader was in-sync, so no committed data is lost. 3. **ISR shrinks 3→2.** Each affected partition's ISR now has 2 members. Because **2 ≥ min.ISR (2)**, the partition is still *writable*: `acks=all` producers keep getting acknowledgments. No stall. 4. **Under-replicated.** The cluster reports `UnderReplicatedPartitions > 0`. There's no third replica until AZ-2 brokers return (they rejoin and catch up) or you reassign to other brokers. 5. **Consumers** keep reading; fetch-from-follower consumers whose preferred replica was in AZ-2 fall back to the leader or another in-AZ follower. **Net result: no data loss, no write outage, just reduced redundancy.** ## What breaks the guarantee - **RF < 3 or fewer than 3 racks.** If a partition had 2 replicas in the failed AZ, only 1 survives → ISR=1 < min.ISR=2 → the partition becomes **read-only for acks=all** (producers stall) until a replica recovers. This is the classic anti-pattern. - **min.insync.replicas misconfigured.** With min.ISR=3 and RF=3, losing one replica drops ISR to 2 < 3 → writes stall even though data is safe. min.ISR should be RF-1 to tolerate one failure. - **Mislabeled broker.rack.** If two replicas of a partition are physically in the same AZ (because labels lied or RF>racks), the math above fails silently. - **Unclean leader election.** If `unclean.leader.election.enable=true` and the only survivor is *out of sync*, Kafka may elect it as leader and **lose committed data**. Keep it false for durability. - **Network partition vs failure.** A partial partition (not a clean AZ death) can cause flapping ISR; quorum/controller behavior differs and may temporarily reduce availability. ## The durability formula To survive **F** simultaneous rack failures with continued `acks=all` writes: `RF ≥ (F+1) × ?` — practically: spread across at least `F + min.ISR` racks so a surviving ISR ≥ min.ISR remains. The common one-AZ-tolerant recipe is **RF=3, racks=3, min.ISR=2, unclean.leader.election.enable=false**.

  • Why must min.insync.replicas be 2 (not 3) for this to tolerate an AZ loss?
    With RF=3 and min.ISR=3, losing any single replica drops ISR to 2 < 3, which stalls acks=all producers even though the data is perfectly safe. Setting min.ISR=RF-1=2 keeps writes flowing through a single-AZ failure while still requiring a 2-replica quorum for durability.
  • What role does unclean.leader.election.enable play here?
    If left false (recommended), Kafka only elects in-sync replicas as leaders, guaranteeing no committed data is lost. If true, an out-of-sync replica can become leader when no in-sync one survives, trading durability for availability and potentially losing data.

saying these in an interview costs you the question

  • Claiming the cluster goes down or loses data on a single AZ loss when RF=3/min.ISR=2 — it survives cleanly.
  • Setting min.insync.replicas equal to RF, which stalls writes on any single replica loss.
  • Assuming rack-awareness alone guarantees survival without correct min.ISR and unclean.leader.election settings.
  • Ignoring that RF spread across <3 racks lets one rack hold a majority of replicas.

context