skip to content

Design an end-to-end no-data-loss producer/topic configuration. Beyond acks=all, what settings are required and what failure modes remain?

level: principalimportance: should knowfreq 40%

answer

  1. acks=all + idempotence
  2. RF=3 / min.isr=2
  3. unclean.leader.election=false
  4. handle the callback (no fire-and-forget)
  5. spread replicas across AZs; consumer commits after process

basics

~20 s

Use acks=all with enable.idempotence=true on the producer, replication.factor=3 and min.insync.replicas=2 on the topic, and unclean.leader.election.enable=false on the brokers. Handle send failures (don't drop them) and bound delivery.timeout.ms. Remaining risks: simultaneous loss of all ISR replicas and consumer-side processing gaps.

solid answer

~40 s

A durable-producer contract layers several settings. **Producer**: `acks=all` (wait for ISR commit), `enable.idempotence=true` (dedup retries, ordered, requires acks=all and max.in.flight ≤ 5), generous `retries`/`delivery.timeout.ms` so transient ISR shrinkage doesn't fail sends, and application code that **handles the send callback/future** rather than fire-and-forget. **Topic**: `replication.factor=3`, `min.insync.replicas=2` so each committed record is on ≥2 brokers and you tolerate one broker failure while still producing. **Broker**: `unclean.leader.election.enable=false` so a non-ISR (out-of-date) replica is never elected leader — electing one would truncate committed records. Even with all this, durability is not absolute: simultaneous loss of all replicas holding a record (e.g. AZ/datacenter outage) loses it, so spread replicas across racks/AZs (`broker.rack`). And end-to-end no-loss also requires consumer-side care — commit offsets only after successful processing.

go deeper

for a junior

Recognize the standard durable trio: acks=all, replication.factor=3, min.insync.replicas=2.

for a middle

Add enable.idempotence and unclean.leader.election.enable=false and explain why each is needed.

for a senior

Reason about delivery.timeout.ms, callback handling, idempotence constraints, and the consistency-over-availability stance.

for a principal

Articulate the complete contract across producer/topic/broker/consumer, the residual failure modes (correlated AZ loss, fsync, outbox), and rack-aware/geo-replication tradeoffs with explicit RPO.

## Goal 'No data loss' means: once the producer's `send()` reports success, the record is never lost despite broker failures. Achieving it requires aligning producer, topic, and broker settings — `acks=all` alone is insufficient. ## Producer settings - **`acks=all`**: wait until the record is committed across the ISR (see committed-vs-leader question). Mandatory. - **`enable.idempotence=true`** (default since 3.0): the producer gets a producer ID and tags each record with a monotonic sequence number per partition. Brokers dedup retries and detect gaps, giving **exactly-once-into-the-log** and ordering even with retries. It *requires* `acks=all`, `retries>0`, and `max.in.flight.requests.per.connection ≤ 5`; the client enforces these. - **`retries`** (effectively ∞ by default) and **`delivery.timeout.ms`** (default 120000): bound total delivery time. Set high enough that a rolling restart or brief ISR shrink (which yields retriable NotEnoughReplicas) doesn't fail the send. - **Application discipline**: never ignore the `Future`/callback. With async `send()`, a swallowed exception is silent loss. Block on critical sends or handle the callback and propagate failures. ## Topic settings - **`replication.factor=3`**: three copies per partition. - **`min.insync.replicas=2`**: under acks=all, require 2 in-sync copies. Tolerates 1 broker down while still accepting writes; every committed record is on ≥2 brokers. ## Broker settings - **`unclean.leader.election.enable=false`** (default false in modern Kafka): forbids electing a replica that is **not** in the ISR as leader. An unclean election would promote an out-of-date replica whose log is behind the HW, **truncating committed records** — direct data loss. Keeping this false means that if no ISR replica is available the partition goes offline rather than losing data (consistency over availability). ## Rack/AZ awareness - **`broker.rack`** + rack-aware assignment spreads the 3 replicas across racks/availability zones, so a single-AZ outage doesn't take all replicas of a partition. ## Remaining failure modes (honest limits) 1. **All ISR replicas lost simultaneously** (e.g. correlated power/AZ failure of the brokers holding a partition) — the committed record is gone. Mitigate with cross-AZ/rack placement and, for the highest tier, geo-replication (MirrorMaker 2 / Cluster Linking) — but cross-cluster replication is async and has its own RPO. 2. **Producer-side crash before send completes** — records buffered in the client (not yet acknowledged) are lost. Mitigate with synchronous critical sends or an outbox pattern. 3. **Consumer-side gaps** — true end-to-end no-loss requires the consumer to **commit offsets only after processing succeeds** (at-least-once) or use transactions/exactly-once. Auto-commit before processing can drop records logically even if Kafka never lost them. 4. **Disk corruption / fsync timing** — Kafka relies on the OS page cache and flushes lazily; a simultaneous power loss of multiple brokers before flush can lose recently-committed data. Replication across machines is the primary defense, not per-broker fsync. ## The complete recipe (one line) Producer: `acks=all, enable.idempotence=true, delivery.timeout.ms` tuned + handle callbacks. Topic: `RF=3, min.insync.replicas=2`. Broker: `unclean.leader.election.enable=false`, rack-aware. Consumer: commit after process.

  • Why does unclean.leader.election.enable=false matter even with acks=all and min.insync.replicas=2?
    If all ISR replicas become unavailable, an unclean election would promote an out-of-date non-ISR replica as leader, truncating committed records below its log — direct loss. Keeping it false takes the partition offline instead, preserving consistency at the cost of availability.
  • You have the full durable config but still lose records. Where do you look beyond the producer?
    The consumer side and the application. Check whether consumers commit offsets before processing (auto-commit), whether the producer ignores send callbacks/exceptions, and whether replicas are co-located in one AZ so a single outage took all copies.
  • What constraints does enable.idempotence=true impose?
    It requires acks=all, retries>0, and max.in.flight.requests.per.connection ≤ 5. The client validates these and throws a config error otherwise. In return it dedups retries and preserves per-partition ordering.

saying these in an interview costs you the question

  • Claiming acks=all alone guarantees no data loss (idempotence, min.isr, RF, and unclean-election settings all matter).
  • Forgetting that fire-and-forget async send with a swallowed exception loses data despite acks=all.
  • Enabling unclean leader election 'for availability' without acknowledging it sacrifices committed records.
  • Ignoring the consumer side — Kafka not losing the record doesn't mean the application processed it.

context