In an LLM rollout, what can shadow traffic measure and what can it never measure?
answer
- real inputs, no real users
- mirror the request, discard the answer
- operational truth without exposure
- same request answered twice, diff them
- side effects are the trap
basics
~20 sShadow traffic mirrors real requests to a candidate whose output is scored but never shown, so it measures behaviour on the true request distribution at real latency and cost with zero user risk. It can never measure user response, because no user ever sees the output.
solid answer
~50 sShadow evaluation mirrors a slice of production requests — say 5% of live grocery search queries — into the candidate system in parallel with the live one. The candidate's output is logged and scored; the shopper only ever sees the production result. What this buys is **online input distribution with offline-style safety**. You see the real request mix, including the seasonal and long-tail queries a frozen set misses, and you can measure real latency, error rate, cost per request and malformed-output rate under production concurrency. You can also diff candidate against incumbent on identical inputs, which is a far tighter comparison than an A/B split gives. What it cannot give is **user response**. Nobody clicks, buys, retries or abandons on a shadow output, so any metric downstream of the user is unavailable. Shadow evaluation therefore closes the input-distribution gap between harness and production, but not the outcome gap — a candidate that looks flawless in shadow can still lose an A/B test.
go deeper
Know that shadow traffic means real requests are copied to a new system whose answers are logged but never shown to anyone, so nothing can go wrong for users.
Explain what that buys and costs: real input distribution, real latency, real cost and rare failure rates at volume, but no user response at all, and doubled inference on the mirrored slice.
Show judgment about when shadow earns its cost — structural changes rather than prompt tweaks — and how to handle stateful systems whose tool calls cannot simply be run twice.
Own where shadow sits in a release policy: which classes of change must pass a shadow stage before exposure, what mirroring rate the inference budget supports, and how shadow findings feed back into the offline suite.
## What shadow evaluation is Shadow evaluation — also called dark launching or mirroring — routes a copy of real production requests to a candidate system that runs alongside the live one. The candidate's response is captured and scored; it is discarded from the user's point of view. The user is served entirely by the incumbent and cannot tell anything else ran. The mirrored share is a knob. Mirroring everything doubles inference cost, so teams typically sample: mirror 5% of live requests, or mirror everything for a subset of segments you care about. Mirroring is usually asynchronous so the shadow call cannot add latency to or fail the user's request — a shadow path that can break production defeats its own purpose. ## The gap it closes The central weakness of offline evaluation is that inputs are a sample somebody chose. Shadow evaluation removes that weakness completely: the inputs *are* production, in the proportions production actually has them, including the seasonal spike, the malformed client payload, the query in a language nobody put in the eval set, and the long tail that individually looks rare but collectively is most of your volume. That unlocks several measurements a harness cannot make honestly: - **Operational behaviour under real load** — p50/p95 latency, timeout rate, error rate and throughput at production concurrency rather than in a sequential batch run. - **Real cost per request** on the true input-length distribution, which is what your bill will look like. - **Failure rates that are only visible at volume** — schema violations, refusals, truncation, tool-call errors that occur in one request in ten thousand and therefore never appear in a thousand-case eval set. - **Paired comparison** — because both systems saw the *same* request, you can diff outputs one to one. Where they agree, nothing changed; where they disagree, you have a small, high-value review queue rather than a whole corpus to grade. That last property is the underrated one. A/B tests compare distributions across different users; shadow compares two answers to the same question. For detecting *changed behaviour* — as opposed to measuring *value* — the paired design is far more sensitive per unit of traffic. ## The gap it cannot close No user ever sees a shadow output, so every metric that lives downstream of the user is unavailable: purchases, task completion, escalation, retry, session length, revenue. Shadow tells you what the candidate *says*; it cannot tell you what anyone *does* about it. Second, systems with side effects cannot be shadowed naively. If the candidate is an agent that calls tools, mirroring it means it will really send the email, really issue the refund, really write to the database — twice. Shadowing anything stateful requires stubbing or sandboxing the write path, and the moment you stub it, the shadow run stops being fully faithful to production. Third, shadowing cannot capture feedback effects. In a search or recommendation system, what the model returns changes what users click, which changes the training and ranking data, which changes future results. A shadow system sits outside that loop by construction. Fourth, scoring is still your problem. Shadow gives you real inputs; it does not give you labels. You still need programmatic checks, a judge model, or human review to say whether a shadow output was good — and at production volume you are sampling those, not scoring everything. ## Where it sits in the release path The natural sequence is offline gate, then shadow, then canary, then split test. Offline catches known regressions cheaply. Shadow proves the candidate behaves sanely on the real request distribution and is operationally viable — right latency, right cost, no unexpected refusal or malformed-output rate — all without a single user exposed. Only then does a canary put real users behind it, and only then does a randomized split answer whether it is actually better. Skipping shadow is common and usually fine for a prompt tweak. It earns its cost when the change is structural — a new model, a new retrieval stack, a rewritten pipeline — where the questions "does this survive real traffic" and "what does it cost at our actual input lengths" are exactly the ones a frozen set cannot answer and an A/B test answers too expensively. ## A worked shape For a grocery search-query rewriter: mirror 5% of live queries into the candidate rewriter, run both rewritten and incumbent query through the same retrieval, log both result sets. Score automatically for empty-result rate, latency delta and out-of-catalogue terms; send the disagreeing pairs to a judge and a human sample. If the candidate produces empty results twice as often on seasonal terms, you learn that before any shopper sees it. If it looks clean, you still do not know it sells more — that requires a split test where shoppers can actually add things to the cart. ## What interviewers listen for The expected answer names the exact tradeoff: shadow buys the production input distribution and operational truth at zero user risk, and pays by learning nothing about user response. Bonus credit for raising side effects on stateful systems, the cost of doubled inference, and the paired-comparison advantage over a split test.
- How do you shadow an agent that calls real tools?You cannot mirror it naively — it would really send the email or issue the refund a second time. Either restrict shadowing to read-only tools, or route write calls to a sandbox or recorded stub. Both weaken fidelity: the stubbed path no longer reflects real tool latency and real failure modes, so you shadow the reasoning and tool selection and validate the write path elsewhere.
- What fraction of traffic should you mirror, and what drives that number?Enough to cover the segments and rare failure modes you care about, bounded by inference cost, since every mirrored request is paid for twice with no user benefit. Teams commonly mirror a single-digit percentage for volume metrics and mirror a targeted segment at a much higher rate when that segment is the concern.
- Why is comparing candidate and incumbent on shadow traffic more sensitive than an A/B split of the same size?Because it is a paired comparison. Both systems answered the identical request, so between-user variance — which dominates A/B noise — cancels out, and you can focus review on the pairs that disagree. It is more sensitive for detecting behavioural change, but it still cannot measure value, since no user acted on either answer.
- After a clean shadow run, what evidence do you still lack before ramping?Everything downstream of the user: whether people accept the output, buy, complete the task, or escalate. Shadow proves the candidate is operationally viable and behaviourally sane on real inputs. Whether it is better requires exposure — a canary to bound risk, then a randomized split to measure the effect.
saying these in an interview costs you the question
- Believes shadow traffic can measure conversion or task completion
- Mirrors an agent with write tools without stubbing side effects
- Ignores that mirroring doubles inference cost on the mirrored slice
- Treats a clean shadow run as sufficient evidence to ship at 100%
- Assumes shadow outputs come with labels and need no scoring