What is a 2.5-DC stretch cluster design, and how do observers / asymmetric quorums (e.g. ZooKeeper or KRaft witnesses) help avoid split-brain across two main datacenters?
answer
- 2 full DCs + tiny tie-breaker = 2.5
- even quorum → no majority → split-brain
- third site holds witness/voter, no data
- observers = async, not in ISR/acks (Confluent MRC)
- RF=4, 2/DC, minISR=3
basics
~20 sA 2.5-DC design runs Kafka across two full datacenters plus a tiny third site holding just a tie-breaker (a witness/observer). The third site holds no data but lets the cluster keep a majority quorum if one of the two main DCs fails, avoiding split-brain.
solid answer
~60 sStretching a single Kafka cluster across exactly two datacenters has a quorum problem: the metadata quorum (ZooKeeper ensemble, or the KRaft controller quorum) needs a strict majority to elect a controller, and an even 2-DC split has no majority when the DCs are partitioned — risking split-brain or full unavailability. The 2.5-DC pattern adds a small, cheap third site that hosts only a quorum tie-breaker — an extra ZooKeeper node or a KRaft controller/witness — but no Kafka data brokers. Now the quorum members total an odd number across three locations, so losing either main DC still leaves a majority (the surviving DC + the tie-breaker), allowing controller election and continued operation. Data brokers and synchronous replicas live only in the two main DCs (often RF=4 with 2 per DC and min.insync.replicas=3, or rack-aware across the two). The third site is 'half a DC' — just enough to break ties. Confluent's Multi-Region Clusters also add observers: async followers that don't count toward ISR/acks, used for cheap remote read replicas or fast promotion.
go deeper
Recognize the term: two main datacenters plus a small third tie-breaker site to keep a majority if one DC fails.
Explain why an even two-DC quorum can't elect a controller and how the third site fixes it.
Detail quorum math, RF/min-ISR placement across the two DCs, and the role of observers.
Architect the full design: failure-domain modeling, cross-DC latency budget, KRaft quorum placement, observers vs sync replicas, and stretch vs MM2 trade-offs.
## The two-DC quorum problem A Kafka cluster has a **metadata quorum** that elects the controller and stores cluster state: historically a **ZooKeeper ensemble**, in modern Kafka the **KRaft controller quorum** (a Raft group). Quorum systems need a **strict majority** of members to make progress — 2 of 3, 3 of 5, etc. If you stretch across **exactly two** datacenters and split the quorum evenly (e.g. 2 nodes in DC-A, 2 in DC-B), a network partition between the DCs gives **neither side a majority** → no controller can be elected → the whole cluster stalls. Worse, naive setups risk **split-brain**: both sides believing they are authoritative. Even 3 nodes split 2/1 means losing the DC with 2 nodes kills the quorum. ## The 2.5-DC pattern The fix is a **third site that is 'half a datacenter'** — it hosts only a **tie-breaker / witness**, not data: - One extra **ZooKeeper node** (so 2+2+1 = 5, or 1+1+1 = 3), **or** a **KRaft controller** there. - **No data brokers** and no synchronous replicas at the third site — keeping it cheap and low-bandwidth. Now the quorum is an **odd total spread across three locations**. If either main DC fails or is partitioned away, the **surviving main DC plus the tie-breaker site form a majority**, so a controller is still elected and the cluster keeps running. There is exactly one majority partition, so **no split-brain**. The data plane lives in the **two main DCs**, typically with rack-awareness treating each DC as a rack: e.g. **RF=4 (2 replicas per DC) with min.insync.replicas=3** so every acked write is in both DCs, or RF=2 across the two. The third site does no data replication. ## Observers (Confluent MRC) Confluent **Multi-Region Clusters (MRC)** add **observers** — a different concept: an observer is an **asynchronous follower** that replicates the partition but is **not part of the ISR** and **does not count toward acks or min.insync.replicas**. Uses: - A **remote read replica** for follower-fetching reads in a distant region without slowing synchronous writes. - A standby that can be **promoted** to a synchronous replica during failover, giving faster recovery than rebuilding from scratch. - Per-topic `replica.placement` JSON controls how many sync replicas and observers go in each rack/region. Because observers are async, they keep cross-region write latency off the critical path while still maintaining a remote copy. ## KRaft consideration In KRaft, the **controller quorum** has the same majority math, so the tie-breaker site holds a **controller voter** (or, with KIP-853-style dynamic quorums and witnesses, a lightweight voter). The principle is identical: an odd number of voters across three failure domains. ## Trade-offs to reason about - **Cross-DC synchronous latency:** with replicas in both main DCs and acks=all, every write pays the inter-DC round trip. The two main DCs must be close enough (low-latency link) for acceptable throughput. - **Cost vs availability:** the tie-breaker site is cheap precisely because it carries no data. - **Surviving more than one site:** 2.5-DC survives any one site loss; it does not survive two simultaneous losses. - Don't confuse stretch-cluster (one logical cluster, synchronous) with **MirrorMaker 2** geo-replication (two clusters, asynchronous) — different consistency and RPO/RTO profiles.
- Why can't you just stretch a Kafka cluster across two datacenters with an even ZooKeeper/KRaft quorum?Quorum election needs a strict majority. An even split (e.g. 2/2) leaves neither side with a majority during a DC-to-DC partition, so no controller can be elected and the cluster stalls — or, if misconfigured, both sides act independently (split-brain). The 2.5-DC tie-breaker creates an odd quorum with exactly one majority partition.
- How does an observer differ from a normal follower in Confluent MRC?An observer is an asynchronous replica that copies the partition but is NOT in the ISR and does NOT count toward acks or min.insync.replicas, so it never slows synchronous writes. It serves remote reads and can be promoted to a sync replica during failover. A normal follower is in the ISR and participates in the acks/min-ISR commit path.
- For data durability across the two main DCs, what RF/min-ISR is common and why?Often RF=4 with 2 replicas per DC and min.insync.replicas=3. That forces each acked write into both DCs (you can't satisfy 3 in-sync replicas from a single 2-replica DC), so a whole-DC loss never loses acknowledged data while still leaving enough replicas to stay writable.
saying these in an interview costs you the question
- Calling the third site a full datacenter or putting data brokers there (it only holds a quorum tie-breaker)
- Saying observers participate in ISR/acks (they are asynchronous and excluded)
- Confusing a 2.5-DC stretch cluster with MirrorMaker async geo-replication
- Believing a 2-DC even quorum is safe (it deadlocks or splits on partition)