skip to content

Active-Passive Disaster Recovery Topologies

Primary-standby topologies: unidirectional replication, RPO and RTO targets, cutover with translated offsets, and failback afterwards. The most common multi-cluster interview scenario.

part ofApache Kafkaoverview, primer and where to startread it →
on this pageshow

questions

6

What is an active-passive (primary/standby) disaster recovery topology in Kafka, and how does it differ from active-active?

level: juniorimportance: must knowfreq 70%

answer

  1. Primary serves all, standby is warm spare
  2. Unidirectional replication via MM2/Replicator/Cluster Linking
  3. Active-active = bidirectional, both serve
  4. Simpler: no write conflicts
  5. Standby idle = wasted capacity, higher RTO

basics

~20 s

In active-passive DR, one Kafka cluster (primary) serves all traffic while a second cluster (standby) only receives a one-way replicated copy. Clients use the standby only after failover. Active-active runs producers/consumers on both clusters at once.

solid answer

~40 s

Active-passive DR uses a primary cluster that handles all production traffic and a standby cluster that receives a continuous unidirectional copy of the data via a replication tool (MirrorMaker 2, Confluent Replicator, or Cluster Linking). The standby is idle for clients under normal operation; producers and consumers run only against the primary. If the primary fails, you fail clients over to the standby (now promoted to primary). Active-active, by contrast, runs live producers/consumers on both clusters simultaneously with bidirectional replication, so both serve traffic. Active-passive is simpler to reason about (no conflicting writes, no read-your-own-replicated-write loops) and is the common choice for DR where the second region exists mainly to survive a regional outage, at the cost of idle standby capacity.

go deeper

for a junior

Know the basic shape: one live cluster, one warm copy, one-way replication, clients fail over on disaster.

for a middle

Be able to name the replication tools and explain why active-passive avoids the conflicts active-active has.

for a senior

Discuss the async-lag/RPO trade-off, IdentityReplicationPolicy for clean names, and idle-capacity cost.

for a principal

Frame the topology choice against business RPO/RTO targets, multi-region cost, and operational testability of failover.

## What problem DR solves Kafka is often the backbone for event streaming. If the entire cluster — or the cloud region hosting it — goes down, you lose the ability to produce and consume. **Disaster Recovery (DR)** topologies replicate data to a second, geographically separate cluster so you can resume operation after a catastrophic loss. ## Active-passive defined In an **active-passive** (also called **primary/standby** or **primary/secondary**) topology: - The **primary** cluster handles 100% of live traffic: all producers write to it and all consumers read from it. - A **standby** cluster continuously receives a **unidirectional** (one-way) copy of the primary's topics. Replication flows primary → standby only. - Under normal operation, no clients talk to the standby. It is effectively a warm spare. - On **failover**, you redirect clients to the standby, which is promoted to be the new primary. Replication is performed by a tool such as **MirrorMaker 2 (MM2)**, **Confluent Replicator**, or **Cluster Linking**. These consume from the source cluster and produce to the target, copying records, and (for MM2) topic configs and consumer-group offsets. ## Active-active contrast In **active-active**, both clusters serve live traffic simultaneously and replication is **bidirectional**. A producer might write to cluster A while another writes to cluster B; each cluster's writes are mirrored to the other. This doubles usable capacity and gives near-zero failover, but it introduces hard problems: avoiding **replication cycles** (a record copied A→B then B→A forever), handling duplicate/ordering semantics across regions, and reconciling concurrent writes. MM2 prevents cycles using the **DefaultReplicationPolicy**, which prefixes mirrored topics with the source cluster alias (e.g. `A.orders` on B), so a topic is never mirrored back onto itself. ## Why pick active-passive - **Simplicity**: one source of truth at a time; no write conflicts, no cross-region ordering puzzles. - **Cost/operational clarity**: the standby is idle, easy to reason about, easy to test failover. - **Trade-off**: standby capacity sits unused, and failover is a deliberate, sometimes manual, action — so RTO is higher than active-active. ## Key edge cases - The standby's topics are typically **remote/prefixed** (e.g. `primary.orders`) under DefaultReplicationPolicy; with **IdentityReplicationPolicy** they keep the same name, which is usually what you want for clean DR cutover. - Replication is asynchronous, so the standby always lags the primary by some amount — this lag is your potential data loss (RPO) on an unplanned failover.

  • Why is active-passive often preferred over active-active for pure DR?
    It avoids write conflicts, replication cycles, and cross-region ordering issues, giving a single source of truth. The standby exists only to survive a regional outage, so simplicity and predictable failover matter more than using the standby's capacity.
  • What replicates the data from primary to standby?
    A replication tool: MirrorMaker 2 (Connect-based), Confluent Replicator, or Cluster Linking. They consume from the source and produce to the target, copying records and often topic configs and consumer-group offsets.

saying these in an interview costs you the question

  • Claiming the standby serves read traffic during normal operation (it doesn't in true active-passive)
  • Saying active-passive replication is synchronous — Kafka cross-cluster replication is asynchronous
  • Confusing intra-cluster broker replication (ISR) with cross-cluster DR replication

context

open as a page

How do RPO and RTO targets shape an active-passive Kafka DR design, and what drives each?

level: middleimportance: must knowfreq 65%

basics

~20 s

RPO (Recovery Point Objective) is the max acceptable data loss, driven by replication lag — async replication means you can lose whatever hasn't reached the standby. RTO (Recovery Time Objective) is the max acceptable downtime, driven by how fast you detect, decide, and cut clients over.

open as a page

On failover to the standby cluster, why are source offsets not valid, and how does MirrorMaker 2 offset translation let consumers resume correctly?

level: seniorimportance: must knowfreq 60%

basics

~20 s

A record's offset on the standby differs from its offset on the primary because replication starts at different points and may compact/skip. MirrorMaker 2's MirrorCheckpointConnector records the primary→standby offset mapping and writes translated consumer-group checkpoints so failed-over consumers resume near where they left off.

open as a page

Contrast planned versus unplanned failover in an active-passive Kafka DR setup. What changes operationally and in achievable RPO?

level: middleimportance: should knowfreq 50%

basics

~20 s

Planned failover (maintenance, drills) lets you stop producers, drain replication so the standby fully catches up, then cut over with RPO near zero. Unplanned failover (primary crashes) gives no chance to drain, so you lose whatever wasn't yet replicated — RPO equals the lag at failure.

open as a page

Walk through how a consumer application actually fails over to the standby cluster using translated checkpoints. What must the client do, and what are the pitfalls?

level: seniorimportance: should knowfreq 45%

basics

~20 s

The consumer reconnects to the standby's bootstrap servers, looks up its group's translated offsets (via RemoteClusterUtils/MirrorClient or pre-synced __consumer_offsets), seeks each partition to those offsets, then resumes. Pitfalls: stale checkpoints causing reprocessing, topic-name prefixes, and running the same group active on both clusters.

open as a page

After recovering the original primary, how do you fail back (resync) to it in an active-passive setup, and why is unplanned-failover fallback the hard case?

level: principalimportance: should knowfreq 40%

basics

~20 s

Fallback means reversing replication so the recovered primary catches up from the now-active standby, then doing a planned failover back. The hard part after an UNPLANNED failover is that the old primary holds writes that never replicated — a divergent tail you must reconcile or discard before reversing, or you get duplicates/conflicts.

open as a page