How do acks=all, min.insync.replicas, and replication.factor work together to define a durability guarantee?
answer
- RF=3, min.isr=2, acks=all
- min.isr only with acks=all
- min.isr=RF → brittle availability
- ISR check before AND after append
- lose 1 broker still write, lose 2 halt
basics
~20 sreplication.factor sets how many copies of each partition exist. min.insync.replicas sets how many in-sync copies must acknowledge a write under acks=all. Only acks=all enforces min.insync.replicas; together they define how many failures the system tolerates without losing acknowledged data.
solid answer
~40 sThree settings combine to set durability. `replication.factor` (topic-level) is how many replicas each partition has — the upper bound on redundancy. `min.insync.replicas` (broker/topic-level) is the minimum number of in-sync replicas that must successfully write a record for an `acks=all` produce to succeed; if fewer ISR members are available, the leader rejects the write with `NotEnoughReplicas`/`NotEnoughReplicasAfterAppend`. Crucially, `min.insync.replicas` is only enforced when `acks=all` — with `acks=0/1` it is ignored. The canonical durable setup is `replication.factor=3`, `min.insync.replicas=2`, `acks=all`: every acknowledged record lives on at least 2 brokers, you tolerate 1 broker failure with continued availability, and you tolerate the loss of all but 1 ISR member without losing acknowledged data. Setting min.insync.replicas equal to replication.factor maximizes durability but loses availability the moment any one replica falls out of the ISR.
go deeper
Know the standard combo RF=3 / min.isr=2 / acks=all and that more replicas means more durability.
Explain that min.isr is only enforced under acks=all and why min.isr=2 with RF=3 balances durability and write availability.
Reason about the ISR-size check before and after append, NotEnoughReplicas vs ...AfterAppend, and the consistency-over-availability tradeoff when ISR shrinks.
Set org-wide topic defaults, weigh min.isr against rolling-restart/GC availability, and combine with unclean.leader.election.enable=false for a complete no-loss contract.
## The three knobs **`replication.factor`** — a per-topic setting: how many copies (replicas) of each partition exist across brokers. With RF=3, each partition has one leader and two followers on three different brokers. This is the *ceiling* on how many copies can exist; it doesn't by itself force writes to reach all of them. **`min.insync.replicas`** (often abbreviated `min.insync.replicas` / `min.isr`) — a broker-default or per-topic setting: the minimum number of replicas in the **in-sync replica set (ISR)** that must acknowledge an `acks=all` write for it to be accepted. The ISR is the set of replicas currently caught up to the leader (within `replica.lag.time.max.ms`). **`acks`** — the producer setting (see related question). `min.insync.replicas` is **only consulted when `acks=all`**. Under `acks=1` or `acks=0` the broker ignores it entirely — a common trap. ## How they interact Under `acks=all`, when a producer sends a record the leader: 1. Checks the current ISR size. If `|ISR| < min.insync.replicas`, it **rejects** the write before appending with `NotEnoughReplicasException` (error code 19). 2. Otherwise it appends to its log and waits for the followers in the ISR to fetch and append. If the ISR shrinks below `min.insync.replicas` *after* the append but before full replication, it returns `NotEnoughReplicasAfterAppendException` (error code 20) — the record may be in the leader's log but is not considered committed. 3. Once all current ISR members have the record, the **high watermark** advances and the record is **committed** and acknowledged. ## The canonical recipe: RF=3, min.isr=2, acks=all - Each committed record is on **at least 2** brokers. - **Availability**: you can lose 1 broker and still produce (ISR=2 ≥ min.isr=2). - **Durability**: you can lose any 1 of the 2 ISR copies and still not lose an acknowledged record (1 ISR copy remains). - Lose 2 brokers and producing **halts** (ISR=1 < 2) — Kafka chooses consistency over availability here. This is intentional: it's better to reject writes than to silently accept un-replicated ones. ## Why not min.isr = RF? Setting `min.insync.replicas=3` with RF=3 means **any** single replica falling out of the ISR (even briefly, e.g. during a rolling restart or GC pause) blocks all producing. You get maximum durability but brittle availability. RF=3/min.isr=2 is the standard balance. ## Tradeoff summary | Setting | Durability | Availability for writes | |---|---|---| | min.isr = 1 | weak (acks=all ≈ acks=1) | high | | min.isr = 2 (RF=3) | strong | tolerates 1 failure | | min.isr = RF | strongest | brittle, no failure headroom | ## Edge cases - `min.insync.replicas` set higher than `replication.factor` makes the topic permanently un-writable under acks=all. - It does not affect consumers — they read up to the high watermark regardless. - It's enforced on the **current** ISR, so a recently-restarted follower not yet back in ISR doesn't count toward the minimum.
- If acks=1 is set, does min.insync.replicas do anything?No. min.insync.replicas is only enforced for acks=all. With acks=0 or acks=1 the broker ignores it, so you can silently lose the durability guarantee you thought the topic had.
- Why is min.insync.replicas=2 (not 3) the recommendation with replication.factor=3?It keeps writes flowing through a single broker failure (ISR can drop to 2 and still meet the minimum) while still requiring 2 copies per record. min.isr=3 would block all writes the instant any replica lags.
- What happens to producing if 2 of 3 brokers in an RF=3 topic go down with min.isr=2?The ISR shrinks to 1, which is below min.insync.replicas=2, so acks=all produces are rejected with NotEnoughReplicas. Writes halt until a replica rejoins — Kafka prioritizes consistency over availability.
saying these in an interview costs you the question
- Believing min.insync.replicas is enforced regardless of acks (it requires acks=all).
- Saying RF and min.isr are the same thing or that min.isr must equal RF.
- Claiming a higher min.isr improves availability (it reduces it).
- Thinking min.insync.replicas affects consumer reads.