skip to content

As an architect, when would you choose active-active bidirectional MM2 over alternatives like a stretched cluster or active-passive, and what consistency limits must you communicate to product teams?

level: principalimportance: should knowfreq 35%

answer

  1. active-active = locality + write availability, eventual
  2. stretched = strong consistency, WAN-latency acks
  3. active-passive = simple DR, idle write capacity
  4. contract: RPO>0, dups, no global order, no auto-merge
  5. home-region affinity recovers per-key consistency

basics

~20 s

Choose active-active MM2 when both regions must serve low-latency local writes and tolerate eventual, asynchronously-replicated cross-region data. Avoid it when you need strong global consistency or strict ordering — a stretched cluster gives synchronous consistency (at WAN-latency cost), and active-passive gives simpler failover without write conflicts. The key limit to communicate: it's eventually consistent, at-least-once, with no global ordering or conflict resolution.

solid answer

~50 s

Active-active bidirectional MM2 fits when each region needs **local, low-latency writes** and the business can accept **asynchronous, eventually-consistent** cross-region propagation with app-level conflict handling. It maximizes write availability (either region survives the other's loss) and serves regional traffic fast. The costs: MM2 is async + at-least-once, gives only per-partition ordering, and never resolves conflicts — so product teams must accept duplicates, possible concurrent-update conflicts, and replication lag, designing idempotency keys and a conflict policy (LWW/CRDT/home-region). Alternatives: a **stretched cluster** (single cluster across DCs with rack-aware replicas) gives strong consistency and single ordering but pays WAN latency on every produce ack and is fragile to partitions; **active-passive** (one writer, MM2 to a standby) eliminates write conflicts and is simpler but loses the passive region's write capacity and needs orchestrated failover. Choose active-active for availability + locality; choose stretched for consistency; choose active-passive for simple DR.

go deeper

for a junior

Know active-active means both regions take writes and data syncs eventually, versus active-passive where only one region writes.

for a middle

Contrast active-active, active-passive, and stretched cluster on latency vs consistency at a high level.

for a senior

Justify a choice with RPO/RTO, ordering, and conflict implications; pair active-active with idempotency and a conflict policy.

for a principal

Own the consistency contract communicated to product teams, weigh stretched vs MM2 vs Cluster Linking, design hybrid topologies (home-region affinity, aggregate clusters), and set RPO/RTO/SLOs across the fleet.

## The decision frame Multi-region Kafka is a CAP/PACELC trade. You're picking where to spend latency and what consistency to give up. Three common shapes: ### 1. Active-active bidirectional (MM2) Both clusters independent, both accept writes, MM2 replicates each way with prefixed remote topics. - **Strengths**: each region serves **local low-latency** produces/consumes; **survives a full region loss** for writes (the other region keeps accepting traffic); scales read/write load per region. - **Costs / limits**: **asynchronous** replication (lag, RPO > 0 — in-flight records can be lost if a region dies before mirroring), **at-least-once** (duplicates), **only per-partition ordering** (no global order), and **no conflict resolution**. The app must own idempotency keys and a conflict policy (last-writer-wins, CRDTs, or home-region affinity). ### 2. Stretched (single) cluster across DCs One Kafka cluster whose brokers span DCs/AZs, using rack-aware replica placement so each partition's replicas land in different DCs; `min.insync.replicas` forces cross-DC acks. - **Strengths**: **strong consistency** and a **single ordering** per partition; no MM2, no prefixing, no conflict logic; transparent to apps. - **Costs**: every `acks=all` produce waits for a **cross-DC (WAN) round trip** — high write latency; very sensitive to **inter-DC network partitions** (can lose quorum); typically limited to low-latency links (metro/3-AZ within a region), not truly distant geos. ### 3. Active-passive (one writer + standby) One region is the sole writer; MM2 mirrors to a passive standby that only takes over on failover. - **Strengths**: **no write conflicts** (single writer), simpler reasoning, clean DR target; passive can serve reads. - **Costs**: passive region's **write capacity is idle**; failover must be **orchestrated** (promote standby, redirect producers, reconcile offsets via checkpoints); RPO > 0 since replication is async. ## Choosing - Need **write availability + regional locality**, tolerate eventual consistency → **active-active**. - Need **strong consistency / single global order**, DCs are close (metro/AZ) → **stretched**. - Need **simple DR**, single source of truth, conflicts unacceptable but full active-active not required → **active-passive**. ## What to explicitly communicate to product teams (the consistency contract) 1. **Eventual consistency**: a write in region A is not instantly visible in B; there's replication lag (monitor `replication-latency-ms`, heartbeats). 2. **Non-zero RPO**: if a region is lost, records not yet mirrored are lost — active-active is **not** zero-data-loss. 3. **At-least-once / duplicates**: every consumer of cross-region data must be idempotent (stable idempotency keys + dedup). 4. **No global ordering**: only per-partition order within a single source topic survives; do not build logic assuming a global event order. 5. **No automatic conflict resolution**: concurrent updates to the same key must be resolved by an explicit policy (LWW/CRDT/home-region). 6. **No dual-writes**: each record produced to exactly one region. 7. **Failover semantics**: consumers fail over via translated offsets (MirrorCheckpointConnector / `sync.group.offsets.enabled`), accepting some reprocessing. ## Edge cases / advanced - **Hybrid**: active-active for serving + a per-region aggregate cluster for analytics (extra hop, more lag, but isolates global reads). - **Home-region affinity** turns active-active into 'active-active across the fleet but single-writer per key', recovering strong per-entity consistency while keeping regional locality for most traffic. - **MM2 vs Confluent Cluster Linking / Replicator**: alternatives exist (Cluster Linking preserves offsets byte-for-byte, simplifying failover); the architectural trade-offs above still hold. - **Cost**: bidirectional doubles cross-region bandwidth (everything mirrored both ways); aggregate clusters add more.

  • A product team wants strict global ordering of events across regions on top of active-active. How do you respond?
    Active-active MM2 cannot provide global ordering — only per-partition order within a single source topic survives, and the aggregate interleaves regions. Options: route all events for a given ordering domain to a single home region (single writer per key) to get per-key order, embed logical clocks/sequence numbers and order at the consumer, or if true global order is mandatory, move that workload to a stretched cluster within a metro region. Don't promise global order on async bidirectional replication.
  • Why is active-active not a zero-data-loss (RPO=0) DR solution?
    MM2 replicates asynchronously, so at any moment there are records produced in one region that haven't yet been mirrored to the other. If that region is lost before those records replicate, they're gone — RPO is greater than zero. Zero-RPO requires synchronous replication (e.g., a stretched cluster with cross-DC min.insync.replicas), which costs WAN-latency on every write.
  • How does a stretched cluster differ from active-active in failure behavior during an inter-DC network partition?
    In a stretched cluster a network partition can break the cross-DC replica quorum, making partitions unavailable for acks=all writes (it favors consistency over availability). In active-active each region is an independent cluster, so a partition just stops replication between them; both keep accepting local writes (favoring availability) and reconcile via MM2 once the link heals — at the cost of divergence/conflicts to resolve.

saying these in an interview costs you the question

  • Pitching active-active as zero-data-loss / RPO=0 — it's async, RPO>0
  • Claiming active-active gives strong consistency or global ordering
  • Recommending a stretched cluster across distant geos (WAN latency / partition fragility)
  • Not communicating the at-least-once + conflict-resolution burden to product teams
  • Treating MM2 as automatic failover rather than a tool that provides translated offsets

context