How does a parallel run (shadow traffic / dark launch) work when cutting over from a monolith's old code path to a newly extracted service, and what specifically does it validate that a staging-environment test cannot?
answer
- shadow/dark traffic, dual execution
- old system stays authoritative during comparison
- diff/compare outputs, don't affect users
- real prod data catches edge cases staging can't
- GitHub Scientist library
basics
~20 sYou send real production traffic to both the old and new systems at once, compare their outputs, but only use the old system's answer — so you catch bugs on real data without risking real users.
solid answer
~50 sA parallel run duplicates real production requests to both the legacy code path and the newly extracted service simultaneously; the legacy result is still what's actually returned to the user (or is the system of record), while the new service's result is captured and compared for divergence, without its output affecting production behavior. This validates against the full, messy variety of real production data and traffic patterns — edge cases, weird historical records, concurrency timing, load characteristics — that a staging environment's synthetic or sampled test data almost never reproduces at the same scale or diversity. Typically the team runs comparison at increasing scale (e.g., 1% then 10% then 100% of traffic shadowed), logs and alerts on mismatches, and only flips the traffic-routing switch to make the new service authoritative once divergence is at or near zero over a sustained period. GitHub's Scientist library is a well-known tool built specifically for this pattern.
go deeper
Should grasp that you can 'test' the new code on real traffic without it affecting real users, by comparing rather than serving its output.
Should describe running both systems on the same request and comparing outputs, and know this catches real-data edge cases staging misses.
Should address the write-side-effect problem, define a concrete rollout/threshold strategy, and distinguish parallel run from canary release.
Should design the observability/tooling investment (diffing pipeline, alerting, sampling strategy) as a reusable capability for the whole migration program, and set acceptable divergence thresholds per domain.
## What a parallel run is A parallel run — also called **shadow traffic** or a **dark launch** — duplicates real production requests to both the existing (legacy) code path and a newly built replacement simultaneously, but only the legacy result is actually returned to the caller or persisted as the system of record; the new path's output is captured and compared against the legacy output for divergence, with no effect on what the user experiences. It is the standard technique for validating a newly extracted service's correctness before making it authoritative, and it sits logically between a purely offline test and a live traffic cutover: the request data is completely real, but the blast radius of a bug in the new path is, ideally, zero. ## What staging fundamentally cannot reproduce The reason this validates something staging environments fundamentally cannot is the nature of production data and load. Staging environments typically run against synthetic, sampled, or stale data, and rarely reproduce the full diversity of real records: - years-old customer accounts with unusual field combinations; - edge-case input from real users; - concurrent requests racing against each other at real traffic volumes and timing patterns. Bugs that only manifest against that diversity — a null field that's always populated in test fixtures but occasionally empty in production, a rounding difference that only shows up past a certain currency amount, a race condition that only occurs under real concurrent load — simply never surface in staging no matter how thorough the test suite, but they will surface, and get caught, in a parallel run against genuine production traffic. ## Ramping, and deciding when to cut over Operationally, teams typically ramp the shadowed percentage of traffic gradually — say 1%, then 10%, then 100% — logging every mismatch between the legacy and new outputs and alerting when the divergence rate exceeds a threshold. GitHub's open-source **Scientist** library is a well-known tool purpose-built for exactly this pattern: 1. wrap a legacy code path and a new "candidate" path; 2. run both; 3. publish the comparison; 4. and always return the legacy result while the team studies the diffs. The decision to finally cut over — make the new service authoritative — is generally made against a divergence-rate threshold sustained over a period long enough to cover the traffic patterns that matter, rather than a fixed calendar duration; a two-day shadow window, for instance, will completely miss any bug that only manifests during a monthly billing run or a quarterly batch process, so the shadow period needs to be chosen with those periodic scenarios explicitly in mind. ## The side-effect constraint The most important constraint on this technique is what to do about the new path's side effects, particularly for operations that aren't naturally read-only. If the shadowed candidate implementation is allowed to actually execute its side effects — genuinely charging a card, sending an email, writing a real database row — then shadowing doubles those side effects for every single request, which for something like a payment or refund operation is a serious production incident in its own right, not a safe test. The standard mitigation is either: - to shadow only up to the point just before the side effect (compare the computed result rather than actually executing it); - or to route the candidate's side effects to a sandboxed or non-production-affecting downstream system specifically so the comparison can happen without real-world consequences. ## Parallel run against canary release It's also worth distinguishing a parallel run from a canary release, since interviewers often probe for the difference. The contrast: | Parallel run | Canary release | |---|---| | carries essentially zero user-facing risk (the old system's answer is always what's used) | actually routes a small slice of real traffic's user-visible response through the new system, so that traffic slice bears real risk if the new system misbehaves | | at the cost of extra compute (running both paths) and the inability to safely validate every kind of side effect | so it validates genuine end-to-end behavior including real side effects | In practice, the two are often used sequentially within a single cutover: parallel-run first, on as much traffic as compute allows, to build statistical confidence with zero user risk, then a canary release to prove the new system under genuinely live conditions, before finally routing all traffic through the new path and retiring the legacy code.
- What has to be true about the operation being shadowed for a parallel run to be safe, especially for write operations?Ideally the shadowed path should be read-only or its side effects sandboxed/no-op'd, because if the new service's write path also has real side effects it doubles those effects for every request. For genuinely stateful writes, teams often shadow only up to the point of the write (compare the computed result, not the executed side effect) or run against a mirrored/non-production-affecting downstream.
- How would you decide when a parallel run has run 'long enough' to trust the cutover?Rather than a fixed calendar duration, teams typically define a divergence-rate threshold sustained over a period that covers relevant traffic cycles — daily peak, weekly batch jobs, monthly billing runs — since bugs tied to rare periodic processes won't show up in a two-day shadow window. They also weight by the severity of mismatches found, since a few cosmetic diffs matter less than any diff in a monetary field.
- What's the difference between a parallel run and a canary release in this context?A parallel run executes both old and new systems for the same request and compares outputs while only the old system's result is user-visible, so it carries zero user-facing risk but doubles compute cost and can't shadow every kind of side effect. A canary release actually routes a small percentage of real traffic's user-visible response through the new system, so it validates true end-to-end production behavior but does carry real risk to that traffic slice — the two techniques are often used sequentially, parallel-run first, canary second.
Like a trainee air traffic controller who sits next to a certified controller and independently calls out the same instructions on a private channel — nobody acts on the trainee's calls yet, but every mismatch between the trainee's call and the real instruction gets reviewed, and only once the trainee's calls match reliably do they get put in charge.
saying these in an interview costs you the question
- Doesn't mention that the old system's result remains authoritative/user-facing during the parallel run
- Suggests shadowing write operations with real side effects without addressing double-execution risk
- Confuses parallel run with a simple A/B test that exposes users to the new system's actual output
- Assumes staging environment testing is equivalent to a production parallel run
- No mention of comparing/diffing outputs — just running both systems with no validation step