What is traffic shadowing, and what does diffing the shadowed responses tell you?
answer
- A copy of real requests
- The response never reaches the caller
- The old version is the oracle
- Writes and outbound calls must be discarded
- Mask volatile fields before comparing
basics
~20 sTraffic shadowing duplicates live requests to a new version running alongside the current one and discards its responses. Comparing the two responses on real inputs shows where the new version behaves differently, before any user sees it.
solid answer
~50 sShadowing - also called mirroring - sends a copy of real requests to a candidate version while the live version keeps serving users; the candidate's responses go to a comparator, never to the caller. Replay is the offline sibling: recorded requests fed to the candidate later. The diff is the oracle, and its oracle is the current version, so what you learn is *where behaviour changed*, not what is correct: pre-existing bugs are reproduced identically and never surface. Two things dominate whether it works. The candidate must be side-effect free - no writes to shared records, no outbound payments, notifications or filings - or shadowing does real damage. And responses must be normalised before comparison, because timestamps, generated identifiers, collection ordering and money-rounding formats otherwise swamp the diff with noise. It is strongest for rewrites of a calculation-heavy component against real input diversity.
code
pseudocode · 13 linesnormalize(response):
body = response.body
body.generated_at = "<masked>"
body.preview_id = "<masked>"
body.deductions = sort(body.deductions, by="code")
return to_fixed_scale(body, places=2)
on mirrored(request):
live = call(current_version, request)
candidate = call(shadow_version, request, side_effects="discard")
if normalize(live) != normalize(candidate):
record_diff(signature(normalize(live), normalize(candidate)), request.employee_hash)go deeper
Know the shape of it: real requests are copied to a second version whose answers are thrown away, and the two answers are compared. Nothing the copy returns can reach a user.
Explain the mechanics an interviewer will push on: suppressing writes and outbound calls, normalising timestamps, generated identifiers and collection order before comparison, sampling to control extra load, and grouping diffs by signature rather than counting them.
Argue the limits out loud - the old version is the oracle, so shared bugs stay invisible; masking rules are blind spots; a replay corpus carries production data and its obligations; and the write path remains unverified until a real rollout.
Judge when the technique is worth its cost at all: it suits like-for-like replacement of a deterministic, input-rich component, and is a poor investment for nondeterministic, personalised or write-dominated endpoints where the same effort buys more as a staged rollout.
### The technique Shadowing sits between a pre-release test and a real release. A copy of production traffic - all of it, or a sampled fraction - is sent to a candidate version deployed alongside the current one. The current version keeps serving users normally; the candidate's responses are collected and thrown away, or handed to a comparator. Nothing the candidate returns reaches a user, so a wrong answer costs nothing directly. **Replay** is the same idea shifted in time: requests are recorded to a corpus and fed to the candidate later, which trades freshness for repeatability and lets you re-run the same corpus after each fix. The value is input diversity. A regression pack contains the cases somebody thought of. A week of production traffic contains the cases nobody thought of: the employee with three concurrent contracts, the pay period containing a leap day, the record with a null cost centre that has been there since a migration in 2019. ### The diff is the oracle - and that shapes what you can learn An oracle is whatever decides that observed behaviour is wrong. Here the oracle is the current version's response. That has a sharp consequence: **shadowing detects change, not correctness**. If the current version has quietly mis-rounded a deduction for two years, the candidate that reproduces the bug produces zero diffs and looks perfect. Conversely a diff is not automatically a defect - if the change was intended, every affected request diffs, and the work is to confirm the population of diffs matches the population you meant to change. So the artefact you want is not a pass/fail but a **classified diff report**: how many requests differed, grouped by the shape of the difference, with an exemplar for each group. ### Side effects: the failure that ends careers The candidate receives real requests. If it shares the primary data store, it will write. If it can reach the outbound integrations, it will send. For a payroll engine that means duplicate ledger entries, duplicate notifications to employees, or - in the worst case - a second disbursement instruction. Before a single request is mirrored you need one of: a candidate wired to an isolated copy of the data, a request context flag that makes every write and every outbound call a no-op sink, or restriction to genuinely read-only endpoints. Read-only endpoints are the natural first target for exactly this reason, and a calculation preview is read-only. A second, quieter side effect is load. Shadowing doubles the request volume hitting shared downstream dependencies unless the candidate is isolated from them too. Sampling a fraction is the usual mitigation, at the cost of covering fewer rare inputs. ### Normalisation: where most first attempts drown Raw response comparison produces a diff rate near 100 percent for reasons that are all uninteresting: - **Timestamps** and other clock-derived values - the candidate is milliseconds behind. - **Generated identifiers** - a new preview identifier per call. - **Ordering** of collections that carry no defined order. - **Formatting** of money and decimals, where one version emits a trailing zero and the other does not. - **Time-of-read skew** - the underlying record changed between the two calls. The comparator therefore canonicalises first: blank out volatile fields, sort unordered collections by a stable key, normalise numeric representation to a fixed scale, and only then compare. Every masking rule is a small loss of coverage, so each one deserves to be written down and revisited - masking a field is exactly how a real regression hides. ### Privacy and the corpus Mirrored and recorded traffic is production data. A replay corpus of payroll requests contains names, salaries and tax identifiers, so it inherits production's access controls, retention limits and redaction obligations. Treating a recorded corpus as an ordinary build asset is a common and serious mistake. ### A worked example A team rewrites the gross-to-net calculator behind a payroll engine and shadows the preview endpoint for one week of a three-week release train. Of 92,600 previews, 287 - about 0.31 percent - differ after normalisation. Grouped by shape, 284 share one signature: the net amount is one day of proration low. Every one of them belongs to an employee whose hire date falls exactly on the first day of the pay period. That is an off-by-one boundary in the proration window, and it is precisely the case the handwritten regression pack did not contain, because nobody writing test data thought to hire someone on the boundary. The remaining three diffs are ordering noise from an unsorted deductions list and are fixed by a normalisation rule. ### When it earns its keep, and when it does not It is strongest for like-for-like replacement of a component with rich input space and deterministic output: a calculation engine, a parser, a pricing or entitlement service, a query layer. It is weak or unusable when the response is legitimately nondeterministic, when the endpoint is inherently write-heavy, when responses are personalised in ways that shift constantly, or when the candidate is not a like-for-like replacement at all - a deliberately new behaviour diffs everywhere and tells you nothing. It also proves nothing about the write path, since you disabled it; the write path still needs a staged rollout with real users to be verified.
- Your first shadow run reports a diff on 96% of requests - what do you look at before filing any defect?Normalisation. A diff rate that high is almost always volatile fields: timestamps, generated identifiers, unordered collections, or differing numeric formatting. Mask or canonicalise those, re-run, and group the surviving diffs by signature. Only then is a diff worth investigating. Keep the list of masking rules explicit and reviewed, because every mask is a blind spot where a genuine regression can hide.
- Shadowing produced zero diffs across a week of traffic. What can you still not claim?That the candidate is correct. The oracle was the current version, so any behaviour both versions get wrong is invisible. You also proved nothing about the write path, outbound side effects or performance under real serving load, since writes were discarded and the candidate served no user. And you only covered the input shapes that week's traffic happened to contain - a quarter-end or year-end shape may never have appeared.
- How does shadowing differ from a staged rollout?In shadowing the candidate answers nobody: its responses are discarded, so an error is harmless and the old version is the comparison oracle. In a staged rollout the new version genuinely serves a slice of users, so its writes are real, its errors are real, and the signal is telemetry rather than a response diff. They are complementary - shadow to de-risk the read path, then ramp to verify the write path and the user-visible outcome.
saying these in an interview costs you the question
- Points the shadow at the primary data store
- Compares raw responses without masking volatile fields
- Treats zero diffs as proof the candidate is correct
- Files every diff as a defect without grouping them
- Stores a replay corpus as an ordinary build asset
- Ignores the doubled load on shared dependencies