skip to content

In a multi-region Kafka deployment, why does setting acks=all with replicas in remote regions hurt producer latency, and what knobs trade durability for speed?

level: juniorimportance: must knowfreq 65%

answer

  1. acks=all waits for whole ISR
  2. remote replica = RTT on every write
  3. acks=1 fast but loses data on leader/region loss
  4. pair acks=all with min.insync.replicas>=2
  5. batch/linger amortize, don't remove, RTT

basics

~20 s

acks=all waits for all in-sync replicas to confirm a write. If some replicas live in another region, every write waits for a cross-region round trip, adding latency. You can use acks=1 or fewer in-sync replicas for speed, but you risk losing data if a region fails.

solid answer

~40 s

A producer's `acks` setting controls when a write is acknowledged. `acks=all` (a.k.a. `acks=-1`) waits for every replica in the in-sync replica set (ISR) to persist the record; combined with `min.insync.replicas`, this is Kafka's strong-durability mode. In a multi-region cluster, replicas (and thus the ISR) span regions, so an acked write must complete an inter-region round trip — directly adding that RTT to producer latency and capping throughput per in-flight batch. To trade durability for speed you can: lower `acks` to 1 (leader-only ack, risks data loss if the leader's region dies before followers catch up) or 0 (fire-and-forget); reduce `min.insync.replicas`; or keep all ISR replicas in the local region and rely on async cross-region replication. Each step weakens the guarantee that an acked record survives a region loss.

go deeper

for a junior

Know acks=all is the durable/slow option and that remote replicas add latency.

for a middle

Pair acks=all with min.insync.replicas and reason about which records survive a region loss.

for a senior

Quantify the RTT floor, weigh local-ISR-plus-async-copy patterns, and account for idempotence/EOS requiring acks=all.

for a principal

Set durability SLOs per topic and design region-aware ISR/quorum layouts that meet them within latency budgets.

## What acks means When a producer sends a record, the partition **leader** writes it and waits according to the `acks` config before replying: - `acks=0` — don't wait at all (fire-and-forget); fastest, can silently lose data. - `acks=1` — wait for the leader only; survives a follower failure but not loss of the leader before replication. - `acks=all` / `acks=-1` — wait for every replica currently in the **ISR** (in-sync replica set: replicas caught up within `replica.lag.time.max.ms`). This is the durable mode. ## Why region placement matters `acks=all` is only as durable as `min.insync.replicas` requires: if fewer than that many replicas are in sync, the leader rejects writes (`NotEnoughReplicas`). So you typically want ISR members in more than one region for region-loss durability. But the leader cannot acknowledge until those remote followers fetch and persist the record — that means an inter-region network **round trip** (RTT) sits on the critical path of every acked write. If region RTT is, say, 60 ms, your floor write latency is ~60 ms regardless of disk speed, and per-connection throughput drops because fewer batches are in flight at once. ## Knobs that trade durability for speed 1. **Lower acks** — `acks=1` removes the remote round trip but an acked record can be lost if the leader region fails before async replication copies it. 2. **Lower `min.insync.replicas`** — fewer required confirmations; can allow writes to ack with replicas only in one region. 3. **Local ISR + async cross-region copy** — keep all synchronous replicas local (fast acks) and use MirrorMaker 2 / Cluster Linking to copy to other regions asynchronously. Fast, but a region loss can lose un-replicated records. 4. **Larger `batch.size` / `linger.ms`** — amortize RTT across more records (improves throughput, not single-record latency). ## Edge cases / pitfalls - `acks=all` without `min.insync.replicas>=2` is a trap: if the ISR shrinks to just the leader, writes still ack but you've lost the cross-region guarantee. - Idempotent/transactional producers (`enable.idempotence=true`) effectively require `acks=all`; you can't safely run exactly-once with weak acks. - Throughput vs latency: batching hides RTT for high-volume pipelines but does nothing for a single low-latency request.

  • Why is acks=all with min.insync.replicas=1 considered misconfigured?
    With min.insync.replicas=1 the leader alone satisfies the requirement, so writes ack even when no follower (and no other region) has the data. You get the latency profile of strong durability only when followers are caught up, but no real protection against losing the leader/region — the durability guarantee you think you have isn't enforced.

saying these in an interview costs you the question

  • Saying acks=all guarantees zero latency cost in multi-region (it adds an RTT).
  • Recommending acks=0 or acks=1 for data you can't afford to lose on a region failure.
  • Believing larger batches reduce single-record latency (they reduce per-record overhead/throughput cost, not the RTT floor).
  • Setting acks=all but leaving min.insync.replicas=1 and assuming region-loss durability.

context