Design an end-to-end no-data-loss producer/topic configuration. Beyond acks=all, what settings are required and what failure modes remain?
answer
- acks=all + idempotence
- RF=3 / min.isr=2
- unclean.leader.election=false
- handle the callback (no fire-and-forget)
- spread replicas across AZs; consumer commits after process
basics
~20 sUse acks=all with enable.idempotence=true on the producer, replication.factor=3 and min.insync.replicas=2 on the topic, and unclean.leader.election.enable=false on the brokers. Handle send failures (don't drop them) and bound delivery.timeout.ms. Remaining risks: simultaneous loss of all ISR replicas and consumer-side processing gaps.
solid answer
~40 sA durable-producer contract layers several settings. **Producer**: `acks=all` (wait for ISR commit), `enable.idempotence=true` (dedup retries, ordered, requires acks=all and max.in.flight ≤ 5), generous `retries`/`delivery.timeout.ms` so transient ISR shrinkage doesn't fail sends, and application code that **handles the send callback/future** rather than fire-and-forget. **Topic**: `replication.factor=3`, `min.insync.replicas=2` so each committed record is on ≥2 brokers and you tolerate one broker failure while still producing. **Broker**: `unclean.leader.election.enable=false` so a non-ISR (out-of-date) replica is never elected leader — electing one would truncate committed records. Even with all this, durability is not absolute: simultaneous loss of all replicas holding a record (e.g. AZ/datacenter outage) loses it, so spread replicas across racks/AZs (`broker.rack`). And end-to-end no-loss also requires consumer-side care — commit offsets only after successful processing.
go deeper
Recognize the standard durable trio: acks=all, replication.factor=3, min.insync.replicas=2.
Add enable.idempotence and unclean.leader.election.enable=false and explain why each is needed.
Reason about delivery.timeout.ms, callback handling, idempotence constraints, and the consistency-over-availability stance.
Articulate the complete contract across producer/topic/broker/consumer, the residual failure modes (correlated AZ loss, fsync, outbox), and rack-aware/geo-replication tradeoffs with explicit RPO.
## Goal 'No data loss' means: once the producer's `send()` reports success, the record is never lost despite broker failures. Achieving it requires aligning producer, topic, and broker settings — `acks=all` alone is insufficient. ## Producer settings - **`acks=all`**: wait until the record is committed across the ISR (see committed-vs-leader question). Mandatory. - **`enable.idempotence=true`** (default since 3.0): the producer gets a producer ID and tags each record with a monotonic sequence number per partition. Brokers dedup retries and detect gaps, giving **exactly-once-into-the-log** and ordering even with retries. It *requires* `acks=all`, `retries>0`, and `max.in.flight.requests.per.connection ≤ 5`; the client enforces these. - **`retries`** (effectively ∞ by default) and **`delivery.timeout.ms`** (default 120000): bound total delivery time. Set high enough that a rolling restart or brief ISR shrink (which yields retriable NotEnoughReplicas) doesn't fail the send. - **Application discipline**: never ignore the `Future`/callback. With async `send()`, a swallowed exception is silent loss. Block on critical sends or handle the callback and propagate failures. ## Topic settings - **`replication.factor=3`**: three copies per partition. - **`min.insync.replicas=2`**: under acks=all, require 2 in-sync copies. Tolerates 1 broker down while still accepting writes; every committed record is on ≥2 brokers. ## Broker settings - **`unclean.leader.election.enable=false`** (default false in modern Kafka): forbids electing a replica that is **not** in the ISR as leader. An unclean election would promote an out-of-date replica whose log is behind the HW, **truncating committed records** — direct data loss. Keeping this false means that if no ISR replica is available the partition goes offline rather than losing data (consistency over availability). ## Rack/AZ awareness - **`broker.rack`** + rack-aware assignment spreads the 3 replicas across racks/availability zones, so a single-AZ outage doesn't take all replicas of a partition. ## Remaining failure modes (honest limits) 1. **All ISR replicas lost simultaneously** (e.g. correlated power/AZ failure of the brokers holding a partition) — the committed record is gone. Mitigate with cross-AZ/rack placement and, for the highest tier, geo-replication (MirrorMaker 2 / Cluster Linking) — but cross-cluster replication is async and has its own RPO. 2. **Producer-side crash before send completes** — records buffered in the client (not yet acknowledged) are lost. Mitigate with synchronous critical sends or an outbox pattern. 3. **Consumer-side gaps** — true end-to-end no-loss requires the consumer to **commit offsets only after processing succeeds** (at-least-once) or use transactions/exactly-once. Auto-commit before processing can drop records logically even if Kafka never lost them. 4. **Disk corruption / fsync timing** — Kafka relies on the OS page cache and flushes lazily; a simultaneous power loss of multiple brokers before flush can lose recently-committed data. Replication across machines is the primary defense, not per-broker fsync. ## The complete recipe (one line) Producer: `acks=all, enable.idempotence=true, delivery.timeout.ms` tuned + handle callbacks. Topic: `RF=3, min.insync.replicas=2`. Broker: `unclean.leader.election.enable=false`, rack-aware. Consumer: commit after process.
- Why does unclean.leader.election.enable=false matter even with acks=all and min.insync.replicas=2?If all ISR replicas become unavailable, an unclean election would promote an out-of-date non-ISR replica as leader, truncating committed records below its log — direct loss. Keeping it false takes the partition offline instead, preserving consistency at the cost of availability.
- You have the full durable config but still lose records. Where do you look beyond the producer?The consumer side and the application. Check whether consumers commit offsets before processing (auto-commit), whether the producer ignores send callbacks/exceptions, and whether replicas are co-located in one AZ so a single outage took all copies.
- What constraints does enable.idempotence=true impose?It requires acks=all, retries>0, and max.in.flight.requests.per.connection ≤ 5. The client validates these and throws a config error otherwise. In return it dedups retries and preserves per-partition ordering.
saying these in an interview costs you the question
- Claiming acks=all alone guarantees no data loss (idempotence, min.isr, RF, and unclean-election settings all matter).
- Forgetting that fire-and-forget async send with a swallowed exception loses data despite acks=all.
- Enabling unclean leader election 'for availability' without acknowledging it sacrifices committed records.
- Ignoring the consumer side — Kafka not losing the record doesn't mean the application processed it.