skip to content

Contrast planned versus unplanned failover in an active-passive Kafka DR setup. What changes operationally and in achievable RPO?

level: middleimportance: should knowfreq 50%

answer

  1. Planned = quiesce + drain to lag 0 → RPO≈0
  2. Unplanned = abrupt, RPO = lag at failure
  3. Unplanned → split-brain risk, fence old primary
  4. Planned failover doubles as a DR drill
  5. Stranded tail on old primary complicates fallback

basics

~20 s

Planned failover (maintenance, drills) lets you stop producers, drain replication so the standby fully catches up, then cut over with RPO near zero. Unplanned failover (primary crashes) gives no chance to drain, so you lose whatever wasn't yet replicated — RPO equals the lag at failure.

solid answer

~50 s

Planned failover is a controlled, deliberate switch: you quiesce producers on the primary, wait for MirrorMaker 2 or Cluster Linking to fully drain the replication backlog (lag → 0), checkpoint final consumer offsets, then repoint clients to the standby and promote it. Because you drained, achievable RPO is near zero and reprocessing is minimal. Unplanned failover follows an abrupt loss (region outage, broker fleet failure): there's no opportunity to drain, so the standby is missing the un-replicated tail and RPO equals the replication lag at the instant of failure. Operationally, unplanned failover also stresses detection and decision-making (avoiding split-brain if the primary partially recovers), may require force-resetting consumer positions from translated checkpoints, and complicates eventual fallback because the old primary now holds writes that never reached the standby. Planned failover is also how you run DR drills to validate RTO.

go deeper

for a junior

Know planned failover is controlled and low-loss; unplanned is abrupt and may lose data.

for a middle

Describe the drain sequence for planned failover and why unplanned RPO equals the lag at failure.

for a senior

Address split-brain fencing, checkpoint staleness on crash, and reprocessing differences.

for a principal

Design detection/decision guards, drill cadence, and fallback reconciliation policy for the stranded tail.

## Two ways the switch happens **Planned failover** is intentional: a region maintenance window, a migration, or a scheduled DR drill. You control timing and sequence. **Unplanned failover** is forced by a real disaster: the primary cluster or its whole region becomes unavailable without warning. ## Planned failover sequence (graceful drain) 1. **Stop/quiesce producers** on the primary so no new records arrive. 2. **Drain replication**: wait until the replication tool reports lag = 0 — every record on the primary is now on the standby. (Monitor MM2 record-lag / Cluster Linking mirror lag.) 3. **Flush consumer checkpoints**: ensure the latest committed group offsets are translated/synced to the standby (`sync.group.offsets.enabled`, or run a final checkpoint). 4. **Promote the standby** and **repoint clients** (bootstrap servers / DNS / config). 5. **Restart producers and consumers** against the new primary. Because you drained, **RPO ≈ 0** and consumers resume with minimal reprocessing. This same sequence is what you exercise in **failover drills** to measure and validate your RTO. ## Unplanned failover sequence The primary is simply gone, so you skip straight to: 1. **Detect** the outage (monitoring/alerting). 2. **Decide** to fail over (ideally with a guard against acting on a transient blip). 3. **Promote** the standby and repoint clients, resuming consumers from the **last translated checkpoints** available (which may lag the true last commit). Consequences: - **RPO = replication lag at failure** — the un-replicated tail is lost (or stranded on the dead primary). - **More reprocessing** because the checkpoints you have are the last ones MM2 managed to emit before the crash. - **Split-brain risk**: if the old primary partially recovers while clients are now on the standby, you can get two diverging "primaries." You must fence the old primary (stop its listeners, revoke client access) before any recovery. ## Fallback implications After an **unplanned** failover, the old primary holds records that never reached the standby. When you eventually fail back (resync), you must decide whether to discard that stranded tail or reconcile it — there's no automatic merge. After a **planned** failover, both clusters were consistent at cutover, so fallback is cleaner. ## Edge cases - A producer burst right before quiescing can extend drain time in a planned failover — budget for it. - For unplanned failover, automation that flips on a single failed health check is dangerous (false positives → unnecessary, lossy cutover); use multi-signal detection and, often, human authorization for the final flip. - Transactional/exactly-once producers complicate both cases because in-flight transactions on the lost primary may be incomplete.

  • Why can planned failover achieve near-zero RPO but unplanned cannot?
    Planned failover stops producers and waits for replication lag to reach zero before cutting over, so nothing is in flight. Unplanned failover happens abruptly with no drain, so any records not yet replicated are lost — RPO equals the lag at the moment of failure.
  • What is split-brain and how do you prevent it during unplanned failover?
    Split-brain is when the old primary recovers and accepts writes while clients have moved to the promoted standby, producing two diverging clusters. Prevent it by fencing the old primary — stopping its listeners or revoking client/network access — before it can resume serving.

saying these in an interview costs you the question

  • Claiming unplanned failover can be lossless like planned failover
  • Forgetting to fence the old primary, allowing split-brain
  • Automating the failover flip on a single health-check signal
  • Assuming fallback after unplanned failover is automatic with no data reconciliation

context