What is the difference between a 'stretch' (stretched) Kafka cluster and a 'replicated' multi-cluster topology across regions, and when would you choose each?
answer
- stretch = one cluster, replicated = many clusters + MM2
- stretch pays RTT per acks=all write
- 2.5 DC tiebreaker for quorum
- replicated = local latency + lag + offset translation
- global ordering only within a single partition leader
basics
~20 sA stretch cluster is one Kafka cluster whose brokers span regions, giving one global view but paying cross-region latency on every write. A replicated topology is separate clusters per region linked by a copy tool (MirrorMaker 2). Stretch favors consistency; replicated favors latency and isolation.
solid answer
~50 sA stretch cluster runs a single logical Kafka cluster with brokers (and KRaft controllers) placed in multiple regions, often 2.5 DCs (two data regions plus a tiebreaker for quorum). Producers with acks=all and replicas in remote regions pay inter-region RTT per write, but you get one offset space, global ordering per partition, and automatic failover with no offset translation. A replicated topology runs independent clusters per region; an async tool like MirrorMaker 2 (or Confluent Cluster Linking) copies topics one way or bidirectionally. Local writes stay fast, regions are failure-isolated, but you inherit replication lag, offset translation, and possible duplicate/out-of-order delivery on failover. Choose stretch when you need strong cross-region durability and a single source of truth and RTT is low (<~50-100ms). Choose replicated for low write latency, region isolation, or when regions are far apart.
go deeper
Know there are two shapes: one big cluster spanning regions vs. separate clusters copied by a tool.
Articulate the latency-vs-consistency tradeoff and name MirrorMaker 2 and acks=all.
Reason about 2.5 DC quorum sizing, ISR/min.insync.replicas, follower fetching, and offset translation costs.
Decide topology per workload from RTT/cost/residency, and design hybrids (stretch in-metro + replicate cross-region).
## The core problem Kafka stores each partition as a replicated log. The leader replica accepts writes; followers in the **ISR** (in-sync replica set) copy them. With `acks=all`, a producer's write is acknowledged only after all in-sync replicas have it. Spreading replicas across regions means an acked write must travel between regions — that is the central tension in multi-region design. ## Stretch (stretched) cluster One single Kafka cluster whose brokers are physically placed in 2+ regions/AZs. Modern Kafka uses **KRaft** (Kafka Raft) controllers for metadata; older deployments used ZooKeeper. Either way the metadata quorum must reach a majority, so a common layout is **2.5 data centers**: two regions hold data brokers and the controller quorum is split 2-2 with a lightweight third site (the '.5') acting as a tiebreaker so a region loss still leaves a controller majority. Properties: - **Single offset space** and single metadata view — consumers and producers see one cluster; no offset translation on failover. - **Global ordering per partition** is preserved because there is exactly one leader per partition. - **Write latency** is dominated by the slowest in-sync replica's region RTT under `acks=all`. - Tools: `min.insync.replicas` and **rack-awareness** (`broker.rack` + the rack-aware replica assignment) ensure replicas land in different regions. **Follower fetching** (`client.rack`, KIP-392) lets consumers read from a local-region follower to cut read latency. - Practical only at low inter-region RTT (same metro / nearby regions, typically <~50-100ms) or you destroy throughput. ## Replicated (connected) topology Independent clusters per region, linked by an async copier: - **MirrorMaker 2 (MM2)**, built on Kafka Connect, copies topics, consumer-group offsets, and ACLs; prefixes remote topics with the source alias by default (`<source>.topic`). - **Confluent Cluster Linking** copies byte-for-byte and preserves offsets without a Connect cluster. Properties: - **Local write latency** — producers hit their nearby cluster. - **Failure isolation** — a region outage doesn't stall writes elsewhere. - Costs: **replication lag** (async, so the remote copy trails), **offset translation** (MM2 uses `RemoteClusterUtils`/the offset-sync topic to map offsets; positions are approximate), and **no global ordering** across clusters. Failover can yield duplicates or gaps. ## Choosing - Need a single source of truth, strong cross-region durability, simple failover, and regions are close → **stretch**. - Need fast local writes, region isolation, far-apart regions, or data-residency boundaries → **replicated**. - Many real systems combine both: stretch within a metro for HA, replicate across distant regions for DR/locality. ## Edge cases - A stretch cluster with replicas only 2-2 across two regions and `min.insync.replicas=2` can stall writes if one region is lost (can't keep 2 ISR) — sizing the quorum and ISR is subtle. - Replicated topologies need idempotent/dedup-tolerant consumers because exactly-once does not span clusters by default.
- Roughly what inter-region latency makes a stretch cluster impractical?Once RTT climbs into the tens-to-hundreds of milliseconds, every acks=all write pays that round trip, collapsing throughput and inflating p99. Stretch is generally viable only for nearby regions / same metro (single-digit to low tens of ms); beyond that, prefer a replicated topology.
- Why does a replicated topology need offset translation but a stretch cluster does not?A stretch cluster has one offset space, so a consumer's committed offset is valid after failover. Separate clusters assign offsets independently for the same logical record, so MirrorMaker 2 maintains an offset-sync topic to approximately map source offsets to target offsets when consumers fail over.
saying these in an interview costs you the question
- Claiming a stretch cluster eliminates cross-region latency (it doesn't — acks=all pays RTT).
- Saying MirrorMaker 2 gives synchronous/strongly-consistent replication (it is async).
- Assuming offsets are identical across replicated clusters without translation.
- Thinking global ordering is preserved across a replicated multi-cluster setup.