skip to content

As a principal engineer, how do acks, min.insync.replicas, the high watermark, and leader epochs together determine whether an acknowledged write can ever be lost? What configuration gives the strongest durability and what are the trade-offs?

level: principalimportance: should knowfreq 35%

answer

  1. acks=all → acked when HW passes offset
  2. min.insync.replicas=2 + RF=3 survives 1 broker
  3. HW = commit boundary; epochs = keep it across failover
  4. unclean.leader.election=false = no out-of-ISR leader
  5. CP trade-off: ISR<minISR → writes rejected

basics

~20 s

Strongest durability: acks=all, min.insync.replicas=2 (RF=3), and unclean.leader.election.enable=false. Then a write is acked only after the HW advances past it (committed on >=2 in-sync replicas), and leader-epoch truncation prevents divergence on failover. The trade-off is higher latency and reduced availability when replicas fall behind.

solid answer

~50 s

Durability is a chain. acks=all means the producer is acked only once the record's offset is below the HW — i.e., replicated to all current ISR members. min.insync.replicas sets the floor on the ISR size required for a write to succeed; with min.insync.replicas=2 and RF=3, a record is committed only when at least two replicas hold it, so one broker can fail without losing it. The HW guarantees consumers never see un-committed records. Leader epochs (KIP-101/279) ensure that on failover the new leader and rejoining followers truncate to the true divergence point, so committed records aren't silently dropped or overwritten. Finally, unclean.leader.election.enable=false forbids electing an out-of-ISR replica as leader — the one remaining way to lose committed data. Trade-offs: higher write latency (wait for replication), and the partition becomes unavailable for writes if the ISR shrinks below min.insync.replicas (a deliberate CP choice over availability).

go deeper

for a junior

Knows acks=all is safer than acks=1.

for a middle

Can pair acks=all with min.insync.replicas and explain the HW commit point.

for a senior

Reasons about the full recipe including unclean leader election and the availability trade-off.

for a principal

Integrates HW, ISR, epochs, and CP/AP trade-offs into a coherent durability architecture and explains failure modes.

## The durability chain Whether an *acknowledged* write can be lost depends on four interacting controls: ### 1. acks (producer) - `acks=0`: fire and forget — lost on any failure. - `acks=1`: acked when the **leader** appends (before the HW advances). If the leader dies before a follower replicates that record, it is lost even though the producer got a success. - `acks=all` (a.k.a. `-1`): the producer is acked only when the record's offset has been **committed**, i.e., is below the HW — replicated to **all replicas currently in the ISR**. This is the prerequisite for no-loss. ### 2. min.insync.replicas (topic/broker) `acks=all` alone is weak if the ISR has shrunk to just the leader (then 'all ISR members' = 1). `min.insync.replicas=N` makes a Produce with acks=all **fail with NotEnoughReplicas** unless at least N replicas are in sync. With **RF=3 and min.insync.replicas=2**, every committed record lives on >=2 brokers, so any single broker loss is survivable. Setting it equal to RF maximizes durability but means *any* replica outage stops writes. ### 3. High watermark The HW is the mechanism that *defines* 'committed' for acks=all: a record is acknowledged precisely when the HW passes its offset. It also guarantees consumers never read past it, so no consumer ever observes a record that could later be lost. ### 4. Leader epochs (KIP-101/279) Even with acks=all and min.insync.replicas=2, the **pre-KIP-101** HW-based truncation could silently lose or diverge committed records during failover. Leader-epoch-based truncation closes that hole by truncating to the actual lineage divergence point. This is the part many candidates forget: the producer/ack settings guarantee a record is *committed*, but **epoch-based recovery is what keeps it committed across leader changes.** ### 5. unclean.leader.election.enable The last gap: if all in-sync replicas are down and `unclean.leader.election.enable=true`, Kafka may elect an **out-of-sync** replica (one missing committed records) as leader to restore availability — losing committed data. Set it to **false** to forbid this (the partition stays offline until an in-sync replica returns). This is the CP-vs-AP knob. ## The strongest-durability recipe - Replication factor 3 - `acks=all` - `min.insync.replicas=2` - `unclean.leader.election.enable=false` - `enable.idempotence=true` (avoids duplicates from retries; default true in modern clients) and bounded/inf retries With this, a committed (acked) record survives any single-broker failure, is never re-ordered or duplicated, and is never dropped or diverged on failover. ## Trade-offs - **Latency:** acks=all waits for replication round-trips; p99 rises, especially across racks/AZs. - **Availability:** if the ISR falls below min.insync.replicas (e.g., two of three brokers slow/down), the partition rejects writes (NotEnoughReplicas). You traded availability for consistency. - **Throughput:** stronger settings + idempotence cap in-flight tuning (max.in.flight.requests.per.connection<=5 with idempotence). - **Operational:** min.insync.replicas=RF is fragile — a single rolling restart can stall producers; min.insync.replicas=2 with RF=3 is the common balance. ## Common pitfalls - Setting acks=all but leaving min.insync.replicas=1 → 'all ISR' can be just the leader → false sense of safety. - Leaving unclean.leader.election.enable=true on a critical topic. - Assuming consumer-visible (below HW) implies producer-durable under acks=1 — it does not; acks=1 records can be lost while still never being consumer-visible before commit.

  • With acks=all but min.insync.replicas=1, where is the durability hole?
    The ISR can shrink to just the leader, so 'all in-sync replicas' = the leader alone. The record is acked while living on one broker; if that broker dies before a follower catches up, the committed record is lost. min.insync.replicas>=2 closes this.
  • Why is unclean.leader.election.enable=false necessary even with acks=all and min.insync.replicas=2?
    If all in-sync replicas are unavailable, allowing an out-of-sync replica to become leader restores availability but at the cost of dropping committed records it never received. Disabling it keeps the partition offline rather than losing data — a consistency-over-availability choice.
  • What role do leader epochs play in this durability story specifically?
    They ensure that on leader change, committed records (below the HW) are not silently truncated or overwritten by the old HW-based recovery. Without epoch truncation, acks=all could still lose committed data during failover sequences.

saying these in an interview costs you the question

  • Saying acks=all alone guarantees no data loss (needs min.insync.replicas and unclean.leader.election=false too).
  • Claiming min.insync.replicas=1 is safe with acks=all.
  • Ignoring leader epochs as part of the durability guarantee across failover.
  • Asserting the strongest config has no availability cost (it intentionally trades availability for consistency).

context