What is an active-passive (primary/standby) disaster recovery topology in Kafka, and how does it differ from active-active?
answer
- Primary serves all, standby is warm spare
- Unidirectional replication via MM2/Replicator/Cluster Linking
- Active-active = bidirectional, both serve
- Simpler: no write conflicts
- Standby idle = wasted capacity, higher RTO
basics
~20 sIn active-passive DR, one Kafka cluster (primary) serves all traffic while a second cluster (standby) only receives a one-way replicated copy. Clients use the standby only after failover. Active-active runs producers/consumers on both clusters at once.
solid answer
~40 sActive-passive DR uses a primary cluster that handles all production traffic and a standby cluster that receives a continuous unidirectional copy of the data via a replication tool (MirrorMaker 2, Confluent Replicator, or Cluster Linking). The standby is idle for clients under normal operation; producers and consumers run only against the primary. If the primary fails, you fail clients over to the standby (now promoted to primary). Active-active, by contrast, runs live producers/consumers on both clusters simultaneously with bidirectional replication, so both serve traffic. Active-passive is simpler to reason about (no conflicting writes, no read-your-own-replicated-write loops) and is the common choice for DR where the second region exists mainly to survive a regional outage, at the cost of idle standby capacity.
go deeper
Know the basic shape: one live cluster, one warm copy, one-way replication, clients fail over on disaster.
Be able to name the replication tools and explain why active-passive avoids the conflicts active-active has.
Discuss the async-lag/RPO trade-off, IdentityReplicationPolicy for clean names, and idle-capacity cost.
Frame the topology choice against business RPO/RTO targets, multi-region cost, and operational testability of failover.
## What problem DR solves Kafka is often the backbone for event streaming. If the entire cluster — or the cloud region hosting it — goes down, you lose the ability to produce and consume. **Disaster Recovery (DR)** topologies replicate data to a second, geographically separate cluster so you can resume operation after a catastrophic loss. ## Active-passive defined In an **active-passive** (also called **primary/standby** or **primary/secondary**) topology: - The **primary** cluster handles 100% of live traffic: all producers write to it and all consumers read from it. - A **standby** cluster continuously receives a **unidirectional** (one-way) copy of the primary's topics. Replication flows primary → standby only. - Under normal operation, no clients talk to the standby. It is effectively a warm spare. - On **failover**, you redirect clients to the standby, which is promoted to be the new primary. Replication is performed by a tool such as **MirrorMaker 2 (MM2)**, **Confluent Replicator**, or **Cluster Linking**. These consume from the source cluster and produce to the target, copying records, and (for MM2) topic configs and consumer-group offsets. ## Active-active contrast In **active-active**, both clusters serve live traffic simultaneously and replication is **bidirectional**. A producer might write to cluster A while another writes to cluster B; each cluster's writes are mirrored to the other. This doubles usable capacity and gives near-zero failover, but it introduces hard problems: avoiding **replication cycles** (a record copied A→B then B→A forever), handling duplicate/ordering semantics across regions, and reconciling concurrent writes. MM2 prevents cycles using the **DefaultReplicationPolicy**, which prefixes mirrored topics with the source cluster alias (e.g. `A.orders` on B), so a topic is never mirrored back onto itself. ## Why pick active-passive - **Simplicity**: one source of truth at a time; no write conflicts, no cross-region ordering puzzles. - **Cost/operational clarity**: the standby is idle, easy to reason about, easy to test failover. - **Trade-off**: standby capacity sits unused, and failover is a deliberate, sometimes manual, action — so RTO is higher than active-active. ## Key edge cases - The standby's topics are typically **remote/prefixed** (e.g. `primary.orders`) under DefaultReplicationPolicy; with **IdentityReplicationPolicy** they keep the same name, which is usually what you want for clean DR cutover. - Replication is asynchronous, so the standby always lags the primary by some amount — this lag is your potential data loss (RPO) on an unplanned failover.
- Why is active-passive often preferred over active-active for pure DR?It avoids write conflicts, replication cycles, and cross-region ordering issues, giving a single source of truth. The standby exists only to survive a regional outage, so simplicity and predictable failover matter more than using the standby's capacity.
- What replicates the data from primary to standby?A replication tool: MirrorMaker 2 (Connect-based), Confluent Replicator, or Cluster Linking. They consume from the source and produce to the target, copying records and often topic configs and consumer-group offsets.
saying these in an interview costs you the question
- Claiming the standby serves read traffic during normal operation (it doesn't in true active-passive)
- Saying active-passive replication is synchronous — Kafka cross-cluster replication is asynchronous
- Confusing intra-cluster broker replication (ISR) with cross-cluster DR replication