skip to content

A team reports intermittent producer failures during routine broker maintenance. Their topic is RF=3, min.insync.replicas=3, acks=all. What is wrong and how would you reason about the fix?

level: seniorimportance: should knowfreq 40%

answer

  1. min.isr = RF -> zero headroom
  2. one broker down -> ISR 3->2 < floor 3 -> writes fail
  3. fix: min.insync.replicas=2 (RF-1), dynamic, no restart
  4. check unclean.leader.election.enable=false
  5. rack-aware placement + watch UnderMinIsrPartitionCount

basics

~20 s

With min.insync.replicas equal to RF, taking any one broker down for maintenance drops the ISR from 3 to 2, below the floor of 3, so all acks=all writes fail. Lowering min.insync.replicas to 2 fixes it while keeping dual-copy durability.

solid answer

~50 s

The root cause is min.insync.replicas=3 on an RF=3 topic — the floor equals the replication factor, leaving zero ISR headroom. Any rolling restart or maintenance that takes one broker offline shrinks the ISR to 2, below the floor of 3, and the leader rejects acks=all writes with NotEnoughReplicasException. The producers retry but during the maintenance window the floor isn't restored, so sends fail. The fix is to set min.insync.replicas=2 (the RF-1 standard): now one broker can be down for maintenance while writes continue with every committed record still on at least two brokers. This is a dynamic config change, no restart needed. I'd also confirm unclean.leader.election.enable=false so durability is preserved, and consider rack/zone-aware replica placement so maintenance domains don't overlap with replica placement. The deeper lesson: min.isr=RF maximizes per-write durability but sacrifices the operability the cluster needs.

go deeper

for a junior

Recognize that min.isr=RF means any one broker down blocks writes.

for a middle

Apply the RF-1 fix and know it's a dynamic config change.

for a senior

Reason through the durability/operability tradeoff, secondary checks, and metrics to monitor.

for a principal

Establish maintenance procedures, fault-domain placement, and org defaults that prevent the class of incident.

## Diagnosis The symptom — producer failures correlated with maintenance — plus the config **RF=3, min.insync.replicas=3** points to a single cause: the **floor equals the replication factor**, so the ISR has **no headroom**. Walk the mechanism: 1. Healthy state: ISR = {leader, follower1, follower2}, size 3, equals the floor. Writes succeed but the system is already at the edge. 2. Maintenance takes one broker offline (or a rolling restart bounces it). That replica stops fetching and is removed from the ISR after `replica.lag.time.max.ms`. 3. ISR shrinks to 2, which is **< min.insync.replicas (3)**. 4. The leader rejects every `acks=all` produce request with **`NotEnoughReplicasException`** (or `...AfterAppendException` for in-flight writes). 5. Producers retry, but the floor stays breached for the whole maintenance window, so sends ultimately fail once `delivery.timeout.ms` elapses. ## The fix Lower the floor to **`min.insync.replicas=2`** (the RF-1 standard). This is a **dynamic topic config**: ``` kafka-configs.sh --alter --entity-type topics --entity-name <topic> \ --add-config min.insync.replicas=2 --bootstrap-server broker:9092 ``` Now during single-broker maintenance the ISR is 2, which meets the floor, so writes continue — and every committed record is still on at least two brokers, so durability is intact. No restart required; the change applies to new produce requests immediately. ## Reasoning about the tradeoff - **min.isr=3 (RF)**: every committed write is on all 3 brokers. Strongest per-write durability, but zero operational slack — you cannot patch, restart, or tolerate even a transient GC pause on any broker without blocking writes. Almost never worth it. - **min.isr=2 (RF-1)**: every committed write is on >= 2 brokers. Tolerates one broker down (failure or maintenance) with writes flowing. The industry standard. - **min.isr=1**: 'all' can mean just the leader -> effectively acks=1, weak durability. ## Secondary checks - **`unclean.leader.election.enable=false`** must hold, or an out-of-sync replica could be elected and truncate committed data, undoing the floor's guarantee. - **Rack/zone awareness** (`broker.rack` + rack-aware assignment): ensure the three replicas live in distinct fault/maintenance domains so a single maintenance action never removes two replicas at once — which would breach even a min.isr=2 floor. - **Maintenance procedure**: rolling restarts should wait for the ISR to fully re-expand (all partitions back to full ISR / no under-replicated partitions) before moving to the next broker. Monitoring `UnderReplicatedPartitions` and `UnderMinIsrPartitionCount` JMX metrics catches this. ## The lesson Maximizing one dimension (per-write durability via min.isr=RF) silently sacrifices another (operability). The RF-1 floor is the deliberate balance, and the broader fix is to make maintenance fault-domain-aware so the floor is never breached by routine operations.

  • Which JMX metric would you alert on to catch this class of problem proactively?
    UnderMinIsrPartitionCount (partitions whose ISR is below min.insync.replicas — these are write-blocked) and UnderReplicatedPartitions (ISR below RF — losing redundancy). Alert on the former as a write-availability incident and the latter as an early warning.
  • Would lowering min.insync.replicas to 1 also fix the failures? Why is that a bad idea?
    Yes, it would stop the failures, but it's bad: with a floor of 1, acks=all can be satisfied by the leader alone, so a single broker/disk failure after acknowledgement loses data. It reduces durability to effectively acks=1. Use 2 (RF-1), not 1.

saying these in an interview costs you the question

  • Recommending lowering min.insync.replicas to 1 as the fix (restores availability but destroys durability).
  • Blaming the producer's retry config instead of the structural min.isr=RF problem.
  • Suggesting a broker restart to fix it (the floor change is a dynamic config, no restart needed).
  • Ignoring fault-domain/rack placement, so maintenance can still take out multiple replicas at once.

context