skip to content

Why does idempotence require acks=all, and how do durability settings (min.insync.replicas) interact with the exactly-once-per-partition guarantee?

level: seniorimportance: should knowfreq 42%

answer

  1. exactly-once = no dup AND no loss
  2. acks=all waits for full ISR
  3. min.insync.replicas = durability floor
  4. RF=3, min.isr=2, acks=all classic
  5. too few ISR -> NotEnoughReplicas -> safe retry

basics

~20 s

acks=all means a write is acknowledged only after all in-sync replicas store it, so an acked record can't be silently lost on leader failover. Idempotence dedups retries, but only acks=all makes the underlying write durable enough for the guarantee to mean exactly-once.

solid answer

~40 s

Idempotence stops duplicate writes from retries, but exactly-once also implies no lost writes. acks=all (the default with idempotence) requires every in-sync replica (ISR) to persist the batch before the leader acks, so a record the producer considers committed survives leader failover. With acks=1 the leader could ack and then crash before replicating, losing the record while the producer thinks it's done — that breaks once-and-only-once. The durability floor is set by the topic/broker config min.insync.replicas: with acks=all and min.insync.replicas=2 on a replication.factor=3 topic, a write needs at least 2 in-sync replicas or the producer gets NotEnoughReplicasException and retries (safely, thanks to idempotence) rather than losing data. So idempotence handles the duplicate side and acks=all + min.insync.replicas handle the loss side; together they give exactly-once-per-partition with real durability.

go deeper

for a junior

Know acks=all is required and means replicas confirm the write.

for a middle

Explain why acks=1 can lose data and why idempotence forces acks=all.

for a senior

Combine acks=all with min.insync.replicas and explain the reject-and-retry path.

for a principal

Reason about availability/durability trade-offs, ISR shrink scenarios, and delivery.timeout.ms fail-closed semantics.

## Exactly-once = no duplicates AND no losses "Exactly once" has two failure modes to defeat: - **Duplicates** — handled by the **idempotent producer** (PID + sequence dedup). - **Losses** — handled by **durability settings**: `acks` and `min.insync.replicas`. If you only solved duplicates but a confirmed write could vanish on failover, you'd have at-most-once-ish behavior, not exactly-once. That's why `enable.idempotence=true` forces `acks=all`. ## acks levels - `acks=0` — fire and forget; no ack. Records can be lost freely. (Illegal with idempotence.) - `acks=1` — leader writes to its log and acks **before** replicating to followers. If the leader crashes before a follower replicates, the record is **lost** even though the producer saw success. (Illegal with idempotence.) - `acks=all` (`-1`) — leader waits until **all in-sync replicas (ISR)** have the record, then acks. Survives leader failover because a new leader is chosen from the ISR, which already has the record. ## min.insync.replicas — the durability floor `acks=all` alone isn't enough if the ISR has shrunk to just the leader (e.g., followers fell behind). The broker/topic config **`min.insync.replicas`** sets the minimum ISR size required to accept an `acks=all` write: - Typical safe setup: `replication.factor=3`, `min.insync.replicas=2`, `acks=all`. A write needs at least 2 replicas in sync. - If fewer than `min.insync.replicas` are in sync, the leader **rejects** the produce with **NotEnoughReplicasException / NotEnoughReplicasAfterAppendException**. The idempotent producer **retries** (safely — no duplicates), turning a potential data-loss window into a temporary unavailability. This is the classic **availability vs. durability** trade: `min.insync.replicas=2` prefers refusing writes over risking loss. ## How idempotence and durability compose | Failure | Defended by | |---|---| | Lost ack -> producer retries -> duplicate | Idempotence (sequence dedup) | | Leader crashes after acking but before replicating | acks=all (wait for ISR) | | ISR too small, leader alone | min.insync.replicas (reject + retry) | ## Subtlety: retries can still observe transient errors Even with everything set correctly, the producer may see retriable exceptions (NotEnoughReplicas, NotLeaderForPartition, network timeouts). Idempotence makes those retries safe. The terminal bound is `delivery.timeout.ms`; if durability can't be achieved within it, `send()` fails rather than risking loss — fail-closed behavior appropriate for exactly-once. ## Common misread acks=all does **not** mean "all replicas" — it means "all replicas **currently in the ISR**." `min.insync.replicas` is what guarantees the ISR is large enough to be meaningful.

  • Does acks=all mean every replica including out-of-sync ones?
    No. It means all replicas currently in the in-sync replica set (ISR). min.insync.replicas sets how large that set must be to accept the write.
  • What happens to an idempotent producer when the ISR drops below min.insync.replicas?
    The leader rejects the produce with NotEnoughReplicas(AfterAppend)Exception; the idempotent producer retries safely until ISR recovers or delivery.timeout.ms expires.

saying these in an interview costs you the question

  • Saying acks=all waits for ALL replicas rather than all in-sync replicas.
  • Claiming idempotence alone prevents data loss — it only prevents duplicates; acks=all + min.insync.replicas prevent loss.
  • Setting acks=1 with idempotence (it's rejected) and assuming it's still safe.
  • Ignoring min.insync.replicas and assuming RF=3 with acks=all is automatically durable when ISR can shrink to 1.

context