skip to content

Why is RF=3 with min.insync.replicas=2 the standard durable configuration, rather than RF=3/min.isr=3 or RF=2/min.isr=2?

level: seniorimportance: must knowfreq 65%

answer

  1. goal: survive 1 broker loss, keep writing, no data loss
  2. floor = RF - 1 rule
  3. RF3/min2 healthy ISR=3, one down -> 2 = floor, writes continue
  4. min.isr=RF -> zero headroom, rolling restart blocks writes
  5. needs unclean.leader.election.enable=false

basics

~20 s

RF=3/min.isr=2 lets you survive one broker failure while still accepting writes and keeping two durable copies. RF=3/min.isr=3 blocks writes on any single failure; RF=2/min.isr=2 blocks writes on any single failure and has no spare copy.

solid answer

~50 s

The goal is to tolerate one broker failure without losing data and without losing write availability. With RF=3 and min.insync.replicas=2, a healthy ISR is 3; if one broker dies the ISR drops to 2, which still meets the floor, so acks=all writes continue and every committed record still lives on two brokers. RF=3/min.isr=3 gives the same durability when healthy but zero fault tolerance for writes — any single replica loss (even a rolling restart) drops the ISR below the floor and blocks all producers. RF=2/min.isr=2 also blocks writes on a single failure and leaves no headroom: losing one of two copies breaches the floor immediately. So RF=3/min.isr=2 is the sweet spot: N=3 copies for fault tolerance, floor of N-1 so one node can go down for maintenance or failure while writes keep flowing with guaranteed dual-copy durability.

go deeper

for a junior

Memorize RF=3/min.isr=2/acks=all as the safe default trio.

for a middle

Explain why one broker can fail and writes still flow with this config.

for a senior

Compare alternatives, derive the floor = RF-1 principle, and tie in unclean leader election.

for a principal

Set cluster-wide durability standards, choose RF=5 for critical paths, and weigh cost vs. fault domains/rack awareness.

## The design goal A production Kafka cluster wants three things simultaneously: (1) **no data loss** when a single broker fails, (2) **continued write availability** during that single failure or a routine rolling restart, and (3) reasonable cost (replicas consume disk and network). The standard answer that balances all three is **RF=3, min.insync.replicas=2, acks=all**. ## Definitions - **Replication factor (RF)**: total copies of each partition. RF=3 means three brokers each hold a full copy. - **min.insync.replicas (floor)**: minimum in-sync copies required for an `acks=all` write to be accepted. - A write is **committed** (durable, visible to consumers) only once it is on every member of the current ISR, and the ISR must be at least the floor. ## Comparing the configurations **RF=3 / min.isr=2 (the standard):** - Healthy ISR = 3. - One broker fails or is restarted -> ISR = 2 = floor -> writes continue, every committed record is on >= 2 brokers. - Two brokers fail -> ISR = 1 < floor -> writes block (durability preserved by refusing), reads continue. - Tolerates **one** failure for writes, **two** failures without data loss for already-committed records (they survive on the one remaining copy, but new writes are blocked). **RF=3 / min.isr=3:** - Same 3 copies, but the floor equals RF. The ISR has **no headroom**. - Any single replica leaving the ISR (a crash, GC pause, or even a planned rolling restart) drops ISR to 2 < 3 -> all `acks=all` writes block. - Maximum durability per committed write, but operationally fragile: you cannot patch or restart a broker without write downtime. Generally over-strict. **RF=2 / min.isr=2:** - Only 2 copies, floor = RF, again no headroom. - One broker down -> ISR = 1 < 2 -> writes block. So you've paid for replication but still lose write availability on a single failure. - And because there are only two copies total, you're one bad disk away from a single-copy situation. **RF=2 / min.isr=1** is sometimes seen but is weak: it's effectively `acks=1` durability since the floor allows the ISR to be just the leader. ## The N / N-1 principle The rule of thumb is **floor = RF - 1**. With RF=3 that's a floor of 2. This gives you exactly one unit of slack: one replica can be absent (failure or maintenance) while writes continue with the durability guarantee intact. Going to RF=5/min.isr=3 follows the same spirit with more slack (tolerate two failures) at higher cost — common for the most critical clusters and for internal topics like the consumer-offsets topic. ## Interaction with unclean leader election For this contract to actually hold, **`unclean.leader.election.enable=false`** (the modern default) must be set. If unclean election were enabled, an out-of-sync replica could become leader and silently truncate committed records, defeating the whole point of the floor. RF=3/min.isr=2 assumes clean elections only. ## Summary table - RF=3/min.isr=2: 1 failure tolerated for writes, dual-copy durability, supports rolling restarts. Standard. - RF=3/min.isr=3: 0 failures tolerated for writes, fragile, over-strict. - RF=2/min.isr=2: 0 failures tolerated, no spare copy. Wasteful.

  • What breaks if you set min.insync.replicas equal to the replication factor?
    You lose all write fault tolerance: any single replica leaving the ISR — a crash or even a planned rolling restart — drops the ISR below the floor and blocks every acks=all producer. Durability per committed write is maximal but availability is fragile.
  • Why does RF=3/min.isr=2 also depend on unclean.leader.election.enable=false?
    If unclean leader election were enabled, an out-of-sync replica could be elected leader and truncate committed records, silently losing data the floor was meant to protect. The durability contract only holds with clean elections (the modern default).

saying these in an interview costs you the question

  • Recommending min.insync.replicas = replication factor as 'most durable' without noting it kills write availability and rolling restarts.
  • Claiming RF=2/min.isr=2 tolerates a broker failure (it blocks writes immediately).
  • Forgetting that the contract requires unclean.leader.election.enable=false.
  • Confusing surviving a failure for reads vs. for writes — committed reads survive longer than write availability.

context