How does the acks setting trade durability against producer throughput and latency?
answer
- 0 = fire-and-forget, can lose, no retries
- 1 = leader only, loss on leader crash
- all/-1 = full ISR, safest, +latency
- acks=all needs min.insync.replicas to be safe
- idempotence forces acks=all
basics
~20 sacks controls how many brokers must confirm a write. acks=0 (no wait) is fastest but can lose data; acks=1 waits for the leader; acks=all waits for all in-sync replicas — most durable but highest latency.
solid answer
~50 sacks sets the durability guarantee per produce request. acks=0: the producer never waits for acknowledgment — highest throughput and lowest latency, but any record can be lost silently (no retries possible since there's no failure signal). acks=1: the leader writes to its log and acks before followers replicate — a leader crash before replication loses recent records. acks=all (a.k.a. acks=-1): the leader waits until all in-sync replicas (the ISR) have the record, bounded by min.insync.replicas, giving the strongest no-loss guarantee at the cost of extra round-trips and latency. The throughput hit of acks=all is often smaller than expected because batching and pipelining (multiple in-flight requests) hide the latency. For exactly-once or no-loss pipelines you must use acks=all with idempotence; for lossy telemetry where speed dominates, acks=0 or acks=1 may be acceptable. Enabling enable.idempotence=true forces acks=all.
go deeper
Know the three values and the basic durability ladder: 0 fast/lossy, 1 leader, all safest.
Explain the acks=1 loss window and how acks=all + min.insync.replicas gives no-loss.
Reason about pipelining hiding acks=all latency and the idempotence/acks coupling.
Set fleet durability policy (replication factor, min.insync.replicas, acks) per data criticality and model the under-replication throughput cliff.
## What acks means `acks` (producer config) tells the leader broker how many acknowledgments the producer requires before considering a produce request successful. It is the core **durability vs latency** dial. ### acks=0 The producer adds the record to a batch, sends it, and **immediately treats it as done** without waiting for any response. - Pros: lowest latency, highest throughput, no ack round-trip. - Cons: **no delivery guarantee** — if the broker is down or the network drops the request, the record is lost silently. Crucially, **retries are meaningless** at acks=0 because the producer never learns of a failure. Offsets returned are not meaningful. - Use: high-volume, loss-tolerant data (metrics, traces) where throughput trumps correctness. ### acks=1 The **leader** writes the record to its local log and acknowledges **before** followers have replicated it. - Pros: one round-trip, good throughput, confirms the leader received it. - Cons: if the leader fails after acking but before a follower replicates, that record is **lost** during leader election. This is the classic 'acks=1 data loss window.' - Use: a middle ground when occasional loss on rare leader failure is tolerable. ### acks=all (acks=-1) The leader waits until **all in-sync replicas (the ISR)** have appended the record before acking. - Durability is governed jointly with the broker/topic config **min.insync.replicas**: the write only succeeds if at least that many replicas are in-sync; otherwise the producer gets NotEnoughReplicas and (with retries) waits or fails. With replication.factor=3 and min.insync.replicas=2, you tolerate one broker loss with no data loss. - Pros: strongest no-loss guarantee; required for idempotent/transactional (exactly-once) producers. - Cons: additional replication round-trip raises per-request latency. ## Why the throughput cost is often modest Latency added by acks=all is not paid serially per record. With batching and `max.in.flight.requests.per.connection > 1`, several requests are in flight at once, so the pipeline hides much of the replication latency — sustained throughput drops far less than the per-request latency increase suggests. ## Interactions to remember - `enable.idempotence=true` (default true in modern clients) **requires acks=all**; setting acks=0 or 1 with idempotence on throws a ConfigException. - Idempotence also constrains `max.in.flight.requests.per.connection` to <= 5 to keep ordering. - `min.insync.replicas` is what actually makes acks=all safe; acks=all alone with min.insync.replicas=1 still allows loss if the lone in-sync replica fails. ## Edge cases - acks=0 + compression still works but a dropped request is unrecoverable. - Under-replicated partitions can stall acks=all writes until ISR is restored — a throughput cliff, not a gradual decline.
- Why does enabling idempotence require acks=all?The idempotent producer must guarantee no duplicate or lost records under retries, which only holds if every committed record is on all in-sync replicas. acks<all could lose an acked record on failover, breaking the exactly-once-per-partition semantics, so the client rejects acks 0/1 with idempotence on.
- Why is the throughput penalty of acks=all often smaller than its latency penalty?Multiple requests are pipelined in flight (max.in.flight.requests.per.connection > 1) and records are batched, so replication latency overlaps across requests. Sustained throughput stays high even though each individual request takes longer to be acked.
saying these in an interview costs you the question
- Saying acks=all waits for ALL replicas (it waits for the in-sync replicas / ISR)
- Claiming acks=0 supports retries (no failure signal, so no retries)
- Thinking acks=all alone guarantees no loss without min.insync.replicas
- Assuming acks=all roughly halves throughput — pipelining hides most of it
- Confusing acks=-1 with a different setting (it equals acks=all)