On the producer side, how do acks and retries determine whether you get at-most-once or at-least-once, and what failure produces a duplicate write?
answer
- acks: 0 / 1 / all
- lost ack → retry → dup
- no retries → loss, no dup
- delivery.timeout.ms bounds retries
- max.in.flight>1 → reorder risk
basics
~20 sRetries cause duplicates when a write succeeds on the broker but the ack is lost — the producer resends, so the record is stored twice (at-least-once). Disabling retries (or acks=0) avoids duplicates but risks losing records that weren't acked (at-most-once).
solid answer
~50 sA producer sends a batch and waits for an **ack** controlled by `acks` (0 = fire-and-forget, 1 = leader persisted, all = all in-sync replicas persisted). If a send fails or times out, the producer **retries** (`retries`, `delivery.timeout.ms`). The duplicate-producing failure is the **lost-ack scenario**: the broker actually persisted the write, but the ack was lost on the network, so the producer believes it failed and resends — the record is now in the log twice → **at-least-once**. To get **at-most-once** on the producer side you avoid retries (and/or use `acks=0`), accepting that any send whose ack never arrives is simply dropped → possible **loss**, no duplicates. `acks=all` plus retries gives durability but, without idempotence, still risks duplicates. The fix for producer-side duplicates is the **idempotent producer** (`enable.idempotence=true`), which tags records with a producer ID + sequence number so the broker dedups retries — but that mechanism is owned by a sibling topic.
go deeper
Know that acks=0/no-retries risks loss and retries can cause duplicates.
Explain the lost-ack duplicate window and map acks/retries to semantics.
Discuss reordering under in-flight>1 and how idempotence closes both the dup and reorder gaps.
Reason about acks/min.insync.replicas/idempotence together for a durability + correctness SLA.
## The producer's two knobs - **`acks`** — how many replicas must persist a record before the broker acknowledges it: - `acks=0`: producer doesn't wait at all (fire-and-forget). Highest throughput, weakest durability — a record dropped before reaching the leader is silently lost. - `acks=1`: the partition **leader** must write it to its log. Lost if the leader fails before a follower replicates. - `acks=all` (a.k.a. `-1`): all **in-sync replicas** (ISR) must persist it. Strongest durability; combined with `min.insync.replicas` it survives broker failures. - **`retries` / `delivery.timeout.ms`** — whether and how long the producer re-sends a batch that failed or timed out. ## How these map to semantics ### At-most-once (no retries / acks=0) If the producer does not retry, then any send whose acknowledgement never arrives is given up on. If the broker never got the record → **loss**, but there is no chance of a second copy → **no duplicates**. `acks=0` amplifies this: the producer assumes success immediately, so any in-flight loss is invisible. ### At-least-once (retries on, acks≥1) With retries enabled and a meaningful ack, the producer keeps re-sending until it gets a success or exhausts `delivery.timeout.ms`. No record is abandoned → **no loss**. But this opens the **duplicate window**. ## The duplicate window (lost-ack scenario) 1. Producer sends record R, broker leader appends R to the log durably. 2. The broker's ack to the producer is **lost** (network drop, broker GC pause, timeout). 3. The producer times out, concludes the send failed, and **retries** R. 4. The broker appends R **again** → the partition now contains R twice. The producer did everything right; the duplicate is inherent to retry-on-uncertain-outcome. This is the producer-side root of at-least-once. ## A subtler hazard: reordering with retries With `max.in.flight.requests.per.connection > 1` and retries, a retried batch can land *after* a later batch, causing **reordering** within a partition. Historically you set `max.in.flight=1` to preserve order; the idempotent producer relaxes this (allows up to 5 while preserving order via sequence numbers). ## The fix (owned elsewhere) `enable.idempotence=true` makes the broker dedup retries using a **producer ID (PID) + per-partition sequence number**, collapsing the duplicate window so retries are safe and order-preserving. The detailed mechanism belongs to the idempotent-producer topic; here the point is only that producer retries are *why* duplicates arise without it. ## Summary | Config | Loss? | Duplicates? | Semantic | |---|---|---|---| | acks=0 / no retries | yes | no | at-most-once | | acks=all + retries (no idempotence) | no | yes (lost-ack) | at-least-once | | acks=all + retries + idempotence | no | no | exactly-once (producer side) |
- Exactly what failure makes acks=all + retries produce a duplicate?The broker persists the record but its ack is lost in transit; the producer times out, retries, and the broker appends a second copy of the same record.
- Why can retries reorder records within a partition, and how is that mitigated?With max.in.flight.requests.per.connection > 1, a retried earlier batch can be appended after a later one. Mitigate by setting max.in.flight=1, or enable.idempotence=true which preserves order via sequence numbers up to 5 in-flight.
saying these in an interview costs you the question
- Saying acks=all alone gives exactly-once (it gives at-least-once; you still need idempotence to dedup retries).
- Claiming retries cause data loss (they prevent loss; they risk duplicates).
- Thinking acks=0 can produce duplicates (it produces loss, not duplicates).
- Confusing min.insync.replicas with acks (they work together but are different knobs).