How does synchronous cross-AZ (or cross-DC) replication latency affect producer throughput and tail latency in a stretch cluster, and what levers tune that trade-off?
answer
- acks=all waits for remote replica RTT
- batch.size + linger.ms amortize the RTT
- max.in.flight pipelines; idempotence keeps order
- lower minISR → wait for nearer replica
- cross-region sync = avoid, use MM2
basics
~20 sWith acks=all the leader waits for in-sync replicas in other zones before acknowledging, so each write pays a cross-AZ round trip. That raises per-record latency. You hide it with batching (linger.ms, batch.size) and concurrency (in-flight requests), and by keeping zones close.
solid answer
~50 sIn a stretch cluster, acks=all + min.insync.replicas≥2 means the leader cannot acknowledge a produce until at least one replica in another AZ/DC has persisted it — so each commit pays a cross-zone network round trip (single-digit ms intra-region, tens of ms cross-region). This inflates per-request produce latency and, if not absorbed, caps throughput. The levers: batching (linger.ms and batch.size let many records share one round trip, raising throughput per RTT), max.in.flight.requests.per.connection (pipeline multiple batches so the link stays busy, but >1 with retries can reorder unless enable.idempotence=true preserves order), compression (less bytes on the wire), and ensuring leaders are spread evenly so no single AZ shoulders all cross-AZ traffic. You also keep min.insync.replicas as low as durability allows (2, not 3) so you wait for the nearest, not the farthest, replica. For latency-critical paths some teams use acks=1 with the knowledge it trades durability. Cross-region stretch is generally avoided for write-latency reasons in favor of async MM2.
go deeper
Know that acks=all waits for other zones, so writes are slower, and batching helps throughput.
Name linger.ms/batch.size and max.in.flight as the throughput levers and that min ISR affects how far the leader waits.
Explain the RTT/throughput math, idempotence-and-ordering interaction, and intra-region vs cross-region choice.
Design the latency budget, balance leaders, and decide stretch-vs-MM2 boundaries against RPO/RTO and SLA targets.
## Where the latency comes from With **acks=all** and **min.insync.replicas ≥ 2**, the partition leader must wait until the required number of **in-sync replicas** have fsync'd/persisted a record before it returns an ack to the producer. In a stretch cluster those replicas live in **other AZs (or DCs)**, so each commit incurs a **network round-trip** to a remote zone: - **Cross-AZ within a region:** typically ~0.5–2 ms RTT. - **Cross-region:** tens of ms (e.g. 30–80 ms), which is why true cross-region *synchronous* stretch is usually avoided. The leader waits for the **slowest replica it needs** to reach min ISR. Lower min ISR (2) means it waits for the *nearest* qualifying replica; higher min ISR (3) means it waits for *more/farther* ones — directly adding latency. ## Effect on throughput If each record waited synchronously and serially, throughput would be capped at roughly `1 / RTT` requests per connection — disastrous over a 30 ms link. Kafka avoids this with two mechanisms: 1. **Batching** — `batch.size` and `linger.ms`. The producer accumulates records into a batch and sends them in **one request**, so a single cross-AZ round trip commits many records. Throughput per RTT scales with batch size. `linger.ms` adds a small wait to let batches fill. 2. **Pipelining** — `max.in.flight.requests.per.connection`. Multiple batches are in flight to a broker at once, keeping the long-RTT link saturated instead of idle while awaiting acks. The default is 5. ## Ordering caveat With `max.in.flight > 1` **and** retries enabled, a failed-then-retried batch could be reordered behind a later one. **`enable.idempotence=true`** (default in modern Kafka) preserves ordering and exactly-once-per-partition semantics even with up to 5 in-flight requests, so you keep pipelining *and* ordering. ## Other levers - **Compression** (`compression.type=lz4/zstd`) shrinks bytes on the wire — helpful when cross-zone bandwidth or egress cost is the constraint. - **Even leader distribution / preferred leader election** so cross-AZ replication load and client traffic are balanced; a skewed leader placement makes one AZ a hotspot. - **min.insync.replicas as low as durability allows** (2) to wait for the closest replica set, not the farthest. - **acks tuning:** acks=1 removes the cross-AZ wait entirely but sacrifices the durability guarantee — a conscious trade for latency-sensitive, loss-tolerant streams. - **request.timeout.ms / delivery.timeout.ms** may need raising for high-RTT cross-region links so legitimate slow acks aren't treated as failures. ## Tail latency Even with batching, **p99 produce latency** is gated by the cross-zone RTT plus replica fsync and any GC/IO stall on the remote follower. Stretch clusters therefore show a higher latency floor than single-AZ clusters. This is the central reason teams keep stretch deployments **within one region** (low cross-AZ RTT) and use **asynchronous MirrorMaker 2** for cross-region, accepting non-zero RPO instead of paying synchronous cross-region latency on every write.
- Why does increasing max.in.flight.requests.per.connection help throughput over a high-latency link, and what config keeps ordering safe?It pipelines multiple batches so the link stays busy instead of idling for each ack over the long RTT, raising throughput. With retries this could reorder batches, but enable.idempotence=true (default) preserves per-partition ordering even with up to 5 in-flight requests.
- Why is synchronous cross-region stretch usually discouraged, and what is used instead?Cross-region RTT is tens of milliseconds, so acks=all pays that on every commit, inflating produce latency and capping throughput. Teams instead run separate clusters with asynchronous MirrorMaker 2 geo-replication, accepting a small RPO rather than synchronous cross-region write latency.
saying these in an interview costs you the question
- Claiming batching reduces per-record latency (it amortizes the RTT for throughput; it can even add a little latency via linger.ms)
- Saying max.in.flight>1 always reorders (idempotence preserves order)
- Recommending min.insync.replicas=3 to 'go faster' (it makes the leader wait for more/farther replicas)
- Assuming synchronous cross-region stretch is fine (the per-write RTT makes it impractical)