skip to content

During a multi-year architecture transition, a team keeps a legacy monolith and a new microservices platform running in parallel for over a year, with dual-write synchronization keeping both in sync. What risks does this long-lived transition state introduce, and how would you mitigate them?

level: principalimportance: should knowfreq 45%

answer

  1. dual-write drift is the classic failure
  2. doubled ops/cognitive load
  3. temporary becomes permanent without forcing function
  4. CDC/single-source-of-truth beats true dual writes
  5. strangler fig shrinks dual-run window

basics

~20 s

Running two systems in parallel for a long time means double the cost, double the bugs, and data can quietly drift out of sync between them — you mitigate by keeping the overlap as short as possible and constantly verifying the two sides agree.

solid answer

~50 s

Long-lived dual-run/dual-write transition states carry compounding risks: data drift, where the two systems silently disagree because one write path fails or retries differently — the classic dual-write inconsistency problem; doubled operational and cognitive load, two systems to monitor, patch, and staff; 'temporary' becoming permanent because funding or attention moves on before cutover completes; and scope/feature freeze pressure, where neither system gets meaningfully improved because both are 'going away soon.' Mitigations include treating the transition state as its own architecture with SLAs and an explicit owner; using a single source of truth with change-data-capture or event sourcing instead of true dual writes wherever possible, to eliminate the write-path race entirely; automated reconciliation jobs that detect and alert on drift; a hard, calendared decommission date with an executive sponsor; and feature-flagged, incremental traffic shifting, such as a strangler-fig cutover, so the dual-run window itself is measured in weeks per slice, not a monolithic year-long freeze.

go deeper

for a junior

Can identify that running two systems at once is more work than one; not expected to reason about dual-write consistency failure modes or reconciliation strategy.

for a middle

Understands data can drift between the two systems and knows monitoring/reconciliation is needed; may not yet default to single-source-of-truth/CDC designs over true dual writes.

for a senior

Actively designs the synchronization approach, choosing CDC vs dual write, builds reconciliation and alerting, and pushes to shrink the transition window via incremental cutover rather than accepting a year-long freeze.

for a principal

Sets organizational policy that no transition state ships without a calendared decommission date and named owner, arbitrates when competing priorities threaten to let a transition state become permanent, and can articulate the total cost of an indefinitely extended transition to secure funding for finishing it.

## What a dual-run state is A **dual-run transition state** keeps two systems — typically a legacy system being retired and its replacement — operating side by side for an extended period, with some mechanism keeping their data consistent so that both can serve reads, and sometimes writes, correctly. Two mechanisms, with very different risk: - **Application-level dual writes** — the riskiest version of this. The calling code, or two separate client paths, writes the same logical change to both systems as two separate operations, one after another, with no distributed transaction spanning them. Because there is no atomicity across the two writes, any partial failure — the first write succeeds and the second times out, retries, or fails outright — leaves the two systems in different states with nothing forcing them back into agreement. - **A single source of truth with replication** — a safer mechanism that achieves the same goal differently. Designate one system as the single source of truth for a given entity, and derive the other side via **change-data-capture (CDC)** or event-driven replication reading off the source's transaction log, so there is exactly one write path and the second system is an eventually-consistent, derived read replica rather than an independently-written peer. ## Why organizations accept it Organizations accept this messier intermediate state because a hard, all-at-once cutover from an old system to a new one is usually too risky to attempt directly on a system carrying live production load — customers, revenue, or regulatory data can't tolerate an unplanned outage or corruption if the cutover goes wrong, so a period where both systems are live and can be compared and gradually shifted between is the practical way to build confidence and allow safe rollback. The dual-run period exists specifically to buy the organization the ability to reverse course cheaply if the new system misbehaves, and to migrate traffic incrementally rather than betting the whole system on a single cutover event. ## The trade-off The trade-off is that every additional week of dual-run buys risk reduction and reversibility at the cost of **doubled operational surface area**: - two systems to monitor, patch, capacity-plan, and staff on-call for, and - two data stores that both need to keep working correctly even though one is explicitly being phased out. Teams also face a **scope-freeze tension** — neither system gets meaningful new investment, because the legacy side is 'going away soon' so bugs get patched but features don't, and the new side is still mid-migration so it can't yet take on unrelated feature work either — which can leave the business unable to respond to changing requirements on either front for the duration of the transition. ## Failure modes 1. The signature failure mode of true dual writes is **silent data drift**: some fraction of records diverge between the two systems because of partial write failures, differing validation logic, or race conditions between concurrent writers, and without an active reconciliation process, nobody notices until a downstream consumer — finance, a customer complaint, an audit — surfaces the mismatch, often long after the drift began. 2. A second failure mode is **the transition state becoming permanent**: without a calendared decommission date and an accountable owner, the 'temporary' bridging state loses priority to newer funded initiatives, and the organization ends up paying the double-operational-cost tax indefinitely — years past the original estimate is a common real-world outcome, with dedicated engineers effectively permanently assigned just to keep the legacy side alive and in sync. 3. A third is **reconciliation debt**: teams intend to build drift-detection tooling 'later' but never prioritize it over feature work, so the sync mechanism runs unmonitored and any drift compounds silently for the entire life of the transition. ## A worked example, and the mitigation playbook Consider a payments team migrating account balances from a legacy monolith to a new microservices ledger, using application-level dual writes for six months before finance discovers roughly 2% of accounts have quietly diverged, traced to write failures on the second leg of the dual write during periods of elevated latency, with no reconciliation job in place to catch it. The concrete fixes illustrate the general mitigation playbook for long transition states: 1. **Replace true dual writes** with a single-source-of-truth-plus-CDC design so there is exactly one authoritative write path and no write-path race to fail. 2. **Add automated reconciliation jobs** that continuously diff the two systems and alert on divergence. 3. **Treat the transition state itself as a first-class architecture** with its own monitoring, SLAs, and on-call ownership rather than an informal hack. 4. **Set a hard, calendared decommission date** for the legacy side with a named executive sponsor accountable for hitting it. 5. **Where possible, use an incremental strangler-fig cutover**, migrating one capability or customer segment at a time behind a routing facade, so any single slice's dual-run window is measured in weeks, not the entire system sitting in a fragile, long-lived intermediate state for a year or more.

  • Why is dual-write synchronization inherently riskier than a single-source-of-truth-plus-replica approach?
    In true dual writes, the client must successfully write to two systems as if it were one atomic operation, but there's no distributed transaction across them — a partial failure leaves the systems permanently inconsistent unless there's compensating reconciliation. A single-source-of-truth design with CDC or event-driven replication has exactly one write path and treats the second system as a derived, eventually-consistent view, which removes the race entirely.
  • What organizational signal tells you a transition state is at risk of becoming permanent?
    When the decommission of the legacy side has no calendared date, no named executive sponsor, and its removal isn't on any team's roadmap or OKRs — 'temporary' states without an active forcing function and an owner accountable for closing them out reliably drift because there's always a more urgent, funded priority competing for the same engineers.
  • How does the strangler fig pattern reduce the risk profile of a long transition compared to a flat 'run both for a year' approach?
    Instead of running the entire system in dual-run for the whole migration, strangler fig moves one capability or traffic slice at a time behind a facade or proxy, so each slice's dual-run window is short and independently verifiable and reversible, rather than the whole system sitting in a fragile, long-lived intermediate state for a year or more.

Like living out of two houses during a slow move — every week you forget to move something to the new place, mail keeps arriving at the old address, you're paying two rents, and the longer the move drags on the more likely you just never finish moving.

saying these in an interview costs you the question

  • Treats a long dual-run period as a purely operational/DevOps concern with no data-consistency risk
  • Has no answer for what happens when a dual write partially fails
  • No decommission date or owner for the legacy side
  • Assumes 'temporary' systems don't need monitoring/SLAs because they're going away soon
  • Doesn't distinguish true dual writes from single-source-of-truth-plus-replication as different risk profiles

context