skip to content

After recovering the original primary, how do you fail back (resync) to it in an active-passive setup, and why is unplanned-failover fallback the hard case?

level: principalimportance: should knowfreq 40%

answer

  1. Fallback = reverse replication, then planned cutover back
  2. Planned-failover fallback is clean (consistent at cutover)
  3. Unplanned: divergent tail on old primary
  4. Don't append on stale tail → duplicates/conflicts
  5. Usually: standby = source of truth, wipe + re-replicate primary
  6. Fence old primary, re-create translation, drill first

basics

~20 s

Fallback means reversing replication so the recovered primary catches up from the now-active standby, then doing a planned failover back. The hard part after an UNPLANNED failover is that the old primary holds writes that never replicated — a divergent tail you must reconcile or discard before reversing, or you get duplicates/conflicts.

solid answer

~60 s

Failback (resync) reverses the DR flow: the standby has been the active cluster, so you stand up replication from standby → recovered-primary, let it fully catch up, then perform a controlled planned failover back to the primary (quiesce, drain to lag 0, translate offsets, repoint clients). After a PLANNED failover the two clusters were consistent at cutover, so reversing replication is clean. After an UNPLANNED failover it's hard: the old primary contains a divergent tail — records it accepted that never reached the standby before it died. If you naively replicate standby → primary on top of that, you get duplicates, offset conflicts, and possibly broken compaction or transactional state. You must decide a reconciliation policy: usually treat the standby as the new source of truth and rebuild the old primary from scratch (wipe its topics, re-replicate fully), discarding the stranded tail, or run an application-level merge if those lost records matter. You also re-establish offset translation in the new direction and validate with a drill before going live.

go deeper

for a junior

Know fallback means going back to the original primary after it recovers.

for a middle

Explain reversing replication and that unplanned failover leaves data on the old primary that complicates this.

for a senior

Detail the divergent-tail problem and the wipe-and-re-replicate reconciliation, plus offset re-translation.

for a principal

Own the reconciliation policy (discard vs merge), split-brain fencing, and drill-validated, business-signed-off fallback runbook.

## What fallback (failback / resync) means After a failover, the **standby** is now serving as the active cluster. **Fallback** is returning to the original primary once it's healthy: you reverse the replication direction so the recovered primary is brought up to date from the current active cluster, then perform a controlled cutover back. It's effectively "a planned failover, in the opposite direction." ## Clean case: fallback after a planned failover Because the original failover was **planned** — producers quiesced, replication drained to lag 0 — the two clusters were **consistent** at the moment of cutover. The old primary holds exactly the data the standby had. So: 1. Configure replication **standby → old-primary** (a new MM2 flow or reverse Cluster Link). 2. Let it copy everything produced on the standby since cutover; wait for lag 0. 3. Re-establish **offset translation** in the new direction (checkpoints standby → old-primary). 4. Do a planned failover back: quiesce, drain, translate offsets, repoint clients. No divergence, so this is mechanically just another planned cutover. ## Hard case: fallback after an unplanned failover When the primary died abruptly, it had **un-replicated records** — a **divergent tail** of writes it accepted that never reached the standby (the RPO loss). When the old primary comes back, its logs still contain those records, and they conflict with the standby's now-authoritative timeline: - The same offsets/partitions exist on both clusters with **different content** beyond the divergence point. - Naively replicating standby → old-primary would **append on top of** the stale tail, producing **duplicates and offset/content conflicts**, and can corrupt **log-compacted** topics or **transactional/EOS** state. ### Reconciliation strategies 1. **Rebuild from scratch (most common)**: declare the standby the **single source of truth**, **wipe** the old primary's topics (or the whole cluster), and **re-replicate fully** from the standby. The stranded tail is **discarded** — acceptable when that data was already counted as RPO loss. 2. **Application-level merge**: if the lost records are business-critical (e.g. payments), extract the divergent tail from the recovered primary and **replay it through the application** so the active cluster can absorb it idempotently. This is bespoke and risky; only do it when the data truly matters. 3. **Forensic discard with audit**: keep the old primary's tail as an offline forensic copy, but don't merge — used where regulators require evidence but the system can't reconcile automatically. ## Operational guardrails - **Fence the old primary** the entire time it's recovering so it never accepts client traffic and never becomes a second active cluster (split-brain). - **Re-create internal MM2 topics** for the new direction (offset-syncs, checkpoints) so translation works the right way. - **Validate with a drill**: do the fallback in a maintenance window and verify consumer positions, lag, and data integrity before repointing production clients. - **Document the decision** (discard vs merge) — it's a business call about the RPO-lost data, not purely technical. ## Why this is a principal-level concern Fallback ties together replication direction, offset translation, divergence reconciliation, split-brain prevention, and a business decision about lost data. The technically tempting "just reverse replication" silently corrupts data after an unplanned failover. The disciplined answer is: treat the surviving active cluster as source of truth, rebuild the recovered one, and only do application-level merge when the lost tail is worth the bespoke risk.

  • Why is naively reversing replication onto a recovered primary after an unplanned failover dangerous?
    The old primary holds a divergent tail of un-replicated writes. Replicating the standby's now-authoritative data on top of that produces duplicate records, offset/content conflicts, and can corrupt log-compacted or transactional state. You must reconcile (usually wipe and re-replicate) before reversing.
  • What is the most common reconciliation choice and when would you deviate?
    Most commonly you declare the standby the single source of truth, wipe the old primary's topics, and re-replicate fully — discarding the stranded tail (already counted as RPO loss). You deviate to an application-level merge only when those lost records are business-critical, like payments, and the bespoke replay risk is justified.
  • How do you prevent split-brain during the recovery and fallback window?
    Fence the recovered old primary so it never serves clients — stop its listeners or revoke network/client access — until it has been resynced and you deliberately cut back to it in a controlled planned failover.

saying these in an interview costs you the question

  • Saying you can just reverse MM2/replication onto the recovered primary with no reconciliation after an unplanned failover
  • Ignoring the divergent tail of un-replicated records on the old primary
  • Allowing the old primary to serve traffic during recovery (split-brain)
  • Assuming fallback is always automatic and lossless
  • Forgetting to re-establish offset translation in the reversed direction

context