In a stretch cluster across 3 AZs, how should replication.factor, min.insync.replicas, and acks be set so the cluster survives the loss of one AZ without data loss?
answer
- RF=3, minISR=2, acks=all
- minISR only with acks=all
- minISR=3 = brittle, no extra durability
- RF=2 → unwritable on AZ loss
- keep unclean election off
basics
~20 sUse replication factor 3 (one replica per AZ), min.insync.replicas=2, and producers with acks=all. Then a write is only acknowledged after two zones have it, so losing one AZ still leaves a committed copy and writes keep working.
solid answer
~40 sFor a 3-AZ stretch cluster the standard recipe is replication.factor=3 with rack-aware placement (one replica per AZ), min.insync.replicas=2, and producers set to acks=all. acks=all means the leader only acknowledges once all in-sync replicas (ISR) have the record; min.insync.replicas=2 means the leader rejects writes (NotEnoughReplicas) unless at least 2 replicas are in sync. Together this guarantees every acknowledged write exists in at least two AZs, so one AZ failing still leaves a committed copy and the partition stays writable (2 ISR remain → 1 can be lost). Setting min.insync.replicas=3 would block all writes the moment any one replica/AZ is down, trading availability for nothing extra on durability. min.insync.replicas=1 with acks=all gives no cross-AZ durability guarantee. The config must be paired with acks=all on the producer — min.insync.replicas does nothing for acks=0/1.
go deeper
Memorize the recipe: replication factor 3, min.insync.replicas 2, acks=all to survive one zone failing.
Explain how acks + min ISR interact and why min ISR 3 or RF 2 break the guarantee.
Discuss high watermark, unclean leader election, the writable-vs-durable distinction, and producer error handling.
Frame the trade space (availability vs durability vs cost) and reason about quorum sizing for multi-AZ/multi-region designs.
## The three knobs - **replication.factor** — number of copies of each partition. Set per topic (`--replication-factor` or the `replication.factor` config). - **min.insync.replicas** (min ISR) — a topic/broker config: the minimum number of replicas that must be **in-sync** for a write to be accepted. The **ISR** (in-sync replica set) is the set of replicas currently caught up to the leader. - **acks** — a *producer* setting: `acks=0` (fire-and-forget), `acks=1` (leader only), `acks=all` / `acks=-1` (all in-sync replicas must persist before the leader acks). These only combine into a durability guarantee when used **together**: `min.insync.replicas` is *only consulted when `acks=all`*. With `acks=1` the leader acks alone and min ISR is irrelevant. ## The 3-AZ recipe With 3 AZs you want **one replica per AZ**, so: - `replication.factor = 3` + rack-aware placement (`broker.rack` per AZ). - `min.insync.replicas = 2`. - producer `acks = all`. **Why these numbers.** With acks=all + min ISR 2, the leader will not acknowledge a record until at least **2 replicas in 2 different AZs** hold it. So every acknowledged record is durably in two zones. If one AZ dies, one replica is lost but two remain in ISR — that is still ≥ min ISR, so: - No acknowledged data is lost (it was in ≥2 zones). - The partition remains **writable** (2 ≥ 2). ## Why not other values - **min ISR = 3** — now *all three* replicas must be in sync to accept writes. The instant any broker or AZ blips, writes halt (`NotEnoughReplicasException` / `NotEnoughReplicasAfterAppend`). You get no extra durability over min ISR 2 but lose single-AZ-failure availability. This is the classic over-tightening mistake. - **min ISR = 1** (or acks=1) — the leader can ack a write that exists in only one AZ. If that AZ dies before a follower copies it, the write is lost. No cross-AZ durability. - **RF = 2** across 2 AZs with min ISR 2 — losing one AZ drops ISR to 1 < 2, so the partition becomes **unwritable** even though data is safe. RF=3/min-ISR=2 is what keeps you both durable *and* available under one-AZ loss. ## Edge cases - When ISR shrinks below min ISR, the leader **rejects produces** but still serves consumers; reads of already-committed data continue. - `unclean.leader.election.enable` should stay **false** in this design — allowing an out-of-sync replica to become leader can resurrect a stale log and lose acknowledged data. - The **high watermark** (the offset up to which all ISR have replicated) is the boundary consumers can read; acks=all + min ISR controls how far that advances. - This protects against *one* AZ loss. Surviving *two* simultaneous AZ losses out of three is not possible while staying writable with RF=3.
- Why is min.insync.replicas=3 with RF=3 usually a bad idea for a stretch cluster?It requires all three replicas in sync to accept any write, so the loss of a single broker or AZ immediately blocks producers (NotEnoughReplicas). It gives no durability beyond min ISR=2 (which already guarantees two-zone persistence) but sacrifices single-AZ-failure availability.
- What error does a producer see when ISR drops below min.insync.replicas?The leader rejects the produce with NotEnoughReplicasException (or NotEnoughReplicasAfterAppend if it fails after the local append). Reads of already-committed data continue; only new writes are blocked until ISR recovers.
- Does min.insync.replicas do anything if the producer uses acks=1?No. min.insync.replicas is only enforced when acks=all. With acks=1 the leader acknowledges on its own, so the min ISR threshold is never checked and you have no cross-replica durability guarantee.
saying these in an interview costs you the question
- Setting min.insync.replicas equal to replication.factor and thinking it improves durability (it only hurts availability)
- Believing min.insync.replicas protects writes under acks=1
- Forgetting that RF=2 across 2 AZs becomes unwritable when one AZ is lost
- Leaving unclean.leader.election.enable=true, which can lose acknowledged data