You must guarantee no acknowledged message is ever lost or silently duplicated end-to-end. Which producer, topic, and broker settings do you combine, and why is acks=all/MISR=2 alone insufficient?
answer
- acks=all + MISR=2 = broker durability only
- unclean=false closes truncation hole
- enable.idempotence = PID + sequence dedup
- retries needed => dups without idempotence
- transactions + read_committed for EOS
basics
~20 sUse RF=3, min.insync.replicas=2, acks=all, unclean.leader.election.enable=false on the broker/topic, plus enable.idempotence=true on the producer (with retries and proper delivery.timeout.ms). acks=all/MISR=2 stops loss on the broker side but doesn't prevent producer retries from creating duplicates or unclean election from truncating data.
solid answer
~40 sEnd-to-end no-loss/no-duplicate needs all layers aligned. On the topic/broker: RF=3, min.insync.replicas=2, acks=all so a committed write lives on >=2 replicas, and unclean.leader.election.enable=false so a stale replica can never be elected and truncate committed records. On the producer: enable.idempotence=true (default in modern clients), which assigns a producer ID and per-partition sequence numbers so the broker deduplicates retried batches and preserves order — without it, the retries that durability requires can silently produce duplicates. Keep retries high (effectively MAX_VALUE under idempotence) and delivery.timeout.ms generous so transient ISR shrink doesn't drop messages. acks=all + MISR=2 alone covers only broker-side durability of an accepted write; it says nothing about duplicate suppression (producer concern) or the unclean-election truncation hole. For atomic multi-partition/exactly-once across read-process-write, layer in transactions (transactional.id, read_committed consumers).
go deeper
Recognize the durable baseline plus 'turn on idempotence' for no duplicates.
List the producer/topic/broker settings and explain idempotence dedup via PID + sequence numbers.
Explain why each layer is necessary and where acks=all/MISR=2 falls short (unclean election, retries).
Design the full end-to-end guarantee including transactions/EOS, error-handling contracts, and the availability cost, and codify it as org standard.
**Goal.** 'No acknowledged message lost, none silently duplicated' is a *pipeline* property; each component has a failure mode that the others don't cover. **Layer 1 — broker/topic durability (prevents loss of accepted writes).** - `acks=all`: producer waits for the full ISR before considering the write done. - `min.insync.replicas=2` (with RF=3): a write is only accepted while >=2 replicas are in sync, so a committed record is on at least two brokers and survives one broker death. - `unclean.leader.election.enable=false`: when all ISR members for a partition are down, the partition goes offline rather than electing a behind replica. This closes the truncation hole — otherwise a stale leader could discard records that acks=all had already acknowledged. **This is why acks=all + MISR=2 alone is insufficient: with unclean=true, the worst case still loses committed data.** **Layer 2 — producer idempotence (prevents silent duplicates).** Durability *requires* the producer to retry on transient failures (NotEnoughReplicas, leader change, timeout). A naive retry can write the same record twice if the first attempt actually succeeded but the ack was lost. `enable.idempotence=true` solves this: the broker assigns the producer a **Producer ID (PID)** and tracks a monotonic **sequence number** per (PID, partition). Duplicate or out-of-order batches are rejected by the broker, giving exactly-once *delivery to a partition* and preserving ordering even with retries. Idempotence also implicitly sets acks=all, retries>0, and max.in.flight.requests.per.connection<=5 — so enabling it is the second half of 'no loss, no dup.' It is the default in clients since Kafka 3.0. **Layer 3 — config consistency.** Keep `retries` effectively unbounded (idempotence sets Integer.MAX_VALUE) and tune `delivery.timeout.ms` to bound total retry time; if it expires during an ISR outage the send fails (surfaced to the app) rather than being silently dropped — the app must handle that error to remain truly lossless. **Layer 4 — exactly-once across stages (optional, for read-process-write).** Idempotence covers a single producer to single partitions. For atomic writes spanning multiple partitions/topics, or to tie consuming + producing into one atomic unit (e.g. Kafka Streams EOS), use **transactions**: set a `transactional.id`, wrap sends in beginTransaction/commitTransaction, commit consumer offsets within the transaction, and have downstream consumers use `isolation.level=read_committed` so they never see aborted records. **Putting it together (typical lossless producer):** - Broker/topic: RF=3, min.insync.replicas=2, unclean.leader.election.enable=false. - Producer: acks=all, enable.idempotence=true, retries=MAX, delivery.timeout.ms sized for tolerable retry window, max.in.flight<=5. - Consumer (for EOS): read_committed; for read-process-write, transactions. **Edge cases.** (1) Idempotence is per producer *session*; a producer restart gets a new PID, so cross-restart dedup needs transactions or app-level keys. (2) If the app ignores a failed send after delivery.timeout.ms, you reintroduce loss — error handling is part of the guarantee. (3) Disk-level durability (`flush`) is usually left to replication rather than fsync-per-message; relying on replication is the standard model, so losing power to multiple brokers simultaneously is the residual risk RF mitigates. (4) MISR=2 with RF=3 means a second concurrent failure stops writes (availability), which is the deliberate price of the durability guarantee.
- What exactly does enable.idempotence add on top of acks=all, and what does it implicitly configure?It makes the broker assign a producer ID and per-partition sequence numbers so retried batches are deduplicated and ordering is preserved, turning at-least-once retries into exactly-once delivery to each partition. It implicitly sets acks=all, retries>0 (MAX), and max.in.flight.requests.per.connection<=5.
- Why isn't acks=all + min.insync.replicas=2 sufficient on its own for a no-loss guarantee?It only ensures an accepted write lands on >=2 in-sync replicas. It doesn't prevent unclean leader election from truncating committed records when all ISR members fail, and it doesn't deduplicate the retries that durability forces — so you also need unclean.leader.election.enable=false and producer idempotence.
- When would you escalate from idempotence to full transactions?When you need atomicity across multiple partitions/topics or must tie consuming and producing into one atomic commit (read-process-write / Kafka Streams exactly-once), using a transactional.id and read_committed consumers.
saying these in an interview costs you the question
- Claiming acks=all alone gives exactly-once or no-duplicates (it's at-least-once without idempotence).
- Forgetting unclean.leader.election.enable=false, leaving the truncation hole open.
- Thinking idempotence survives producer restarts without transactions (PID changes on restart).
- Ignoring that a failed send after delivery.timeout.ms reintroduces loss unless the app handles it.