A candidate ticket router scores live tickets in shadow mode - why can that window not show whether it routes better than the incumbent?
answer
- inputs are real, consequences are not
- output recorded, then thrown away
- nobody works the candidate's queue
- outcomes exist only where both agreed
- parity check, not a quality verdict
basics
~20 sShadow mode compares mechanics, not outcomes. The incumbent's queue choice is what agents actually work, so no ticket is ever placed where the candidate said, and resolution time and reroute rate are never produced for the disagreements.
solid answer
~40 sShadow mode mirrors live tickets to the candidate router, records what it said, and throws the output away: agents still work the queue the incumbent picked. That makes the window a **parity check on mechanics** - scoring latency, error and timeout rates under real traffic, the shape of the candidate's score distribution, the rate at which it disagrees with the incumbent, and how much the queue mix would move. It produces no outcome for the disagreements, because the candidate's queue never received those tickets: nobody worked them, so resolution time, reroute rate and handling time do not exist for them. Outcomes do exist on the agreed tickets, but that is exactly the subset where the two routers are interchangeable, so it measures nothing about the change.
go deeper
Recall the shape: a copy of the live input goes to the candidate, its answer is written to a log and discarded, and the incumbent still decides everything. Nothing the candidate says reaches a customer or an agent.
Explain which signals survive and which do not, and why. Latency, error rate, null rates, score distributions and disagreement counts are real; anything that needs someone to act on the candidate's decision is unavailable.
Show that you would present a shadow result as a deployability verdict, not a quality one, and that you would adjudicate a sample of the disagreements with a human rather than argue from the aggregate agreement number.
The tradeoff to speak to is how much evidence is worth buying before any exposure at all. Shadow is the cheapest window and the least informative one, and a team that keeps extending it is usually avoiding the decision it actually has to make.
## What a shadow window is Shadow scoring (dark-launch scoring, mirrored scoring) sends a copy of real production input to a candidate model version, runs it, records what it said, and then **discards the output**. On the support desk here, every incoming ticket still goes to the queue the incumbent router chose. Agents see nothing different, customers see nothing different, and the candidate's chosen queue exists only as a row in a log. That design buys one thing and gives up another. It buys **real inputs**: the candidate is scored on live ticket text, live customer history and the same online feature values the serving path fetched, at the same hour and the same load. It gives up **real consequences**: no ticket is ever placed where the candidate said to place it. ## What the window can measure - **Operational parity** - the p50/p95/p99 of the candidate's scoring call, its error and timeout rate, its memory and accelerator footprint, on production traffic rather than a benchmark harness. - **Input compatibility** - missing features, null rates, schema mismatches, category values the candidate never saw during training. These surface in the first hour of a shadow window and do not surface in an offline evaluation on a stored dataset. - **Score-distribution parity** - the candidate's output score distribution against the incumbent's over the same tickets. A two-sample Kolmogorov-Smirnov test, or the Jensen-Shannon divergence between the two histograms, turns 'looks different' into a number. - **Disagreement rate** - the share of tickets on which the two routers choose different queues, broken down by channel, language, ticket type and hour. - **Blast radius** - how far the queue mix would move if the candidate owned all the traffic. A shift from 12% to 30% of tickets into a specialist queue is a staffing problem worth knowing about before exposure, not after. ## Why no outcome comes out The chain that produces a resolution time is: a ticket lands in a queue, an agent working that queue picks it up, the agent resolves it or reroutes it. Shadow breaks that chain at the first link, and it breaks it for exactly the tickets that carry the information. Split the window in two: 1. **Tickets where the routers agree.** An outcome exists, because the queue the candidate named is the queue the ticket actually went to. But the two routers are interchangeable on this subset, so the outcome describes the ticket, not the change. 2. **Tickets where they disagree.** No outcome exists in the candidate's world at all. The ticket went to the incumbent's queue; the candidate's queue never received it; no agent ever handled it under the candidate's decision. | Signal | Available in a shadow window | Why | |---|---|---| | Candidate scoring latency and error rate | Yes | The candidate really ran, on real input | | Null or missing feature rate on live tickets | Yes | The candidate really read the online feature values | | Disagreement rate and its segment breakdown | Yes | Both decisions are recorded per ticket | | Resolution time on a disagreed ticket | No | The ticket was never placed in the candidate's queue | | Reroute rate caused by the candidate | No | No agent ever received a ticket on the candidate's decision | | Backlog and handling time in a queue that would grow | No | Capacity effects only exist when the decision is real | There is a second-order loss too. If the candidate concentrates tickets into one specialist queue, that queue's real backlog and its agents' handling time change - and a shadow run cannot simulate that, because the tickets never arrived there. The candidate's decisions also never enter the feedback loop that produces the next training set, so nothing it does in shadow can teach it anything. ## The trap The trap is treating a clean shadow window as evidence of quality. 'Latency is fine and it agrees with the incumbent on 82% of tickets' says the candidate is **deployable**, not that it is **better**. Quality requires somebody to act on the decision, and acting on the decision is the step after this one. ## Reading a shadow window honestly - Treat it as a go/no-go on safety and compatibility, and say so out loud when you present it. - Treat the disagreements as a review queue: sample them and have an experienced agent judge which queue was right. That small adjudicated sample is the closest thing to an outcome the window can produce, and it is worth far more than the aggregate agreement number. - Write down the expected queue-mix shift and hand it to whoever staffs the queues before anything is exposed. - Record which periods were shadowed. A window that ran only on a quiet weekday tells you nothing about peak-hour behaviour.
- The candidate won an offline head-to-head against the incumbent. What does a clean shadow window add before the candidate is exposed to any real traffic?It adds evidence the offline comparison could not produce: the candidate's behaviour on live feature values rather than a stored dataset, its latency and error rate under real load, the volume and segment shape of its disagreements, and the queue-mix shift it would cause. It adds nothing about the outcome, so an offline winner that passes shadow still has to be exposed to real traffic before anyone knows whether it routes better.
- Agreement between the two routers is 99%. Is that a good sign or a bad one?It is ambiguous, and it is a reason to inspect rather than celebrate. Near-total agreement means the candidate is safe to deploy and probably also means it will change almost nothing, so the effort may not be worth it. It can also mean a wiring fault: the candidate may be reading a cached incumbent decision, or falling back to a default that happens to match. Check the score distributions, not just the decisions.
- Can you measure training-serving skew from a shadow window?Partly. The window shows the feature values the candidate actually received in production, so you can compare their distribution against the distribution in the training snapshot and catch a transform that behaves differently on the serving path. What it cannot show is whether that difference costs you anything in outcome, because no outcome exists on the disagreements.
A flight simulator records every input a trainee makes while the instructor's hands stay on the real controls. It can prove the trainee reacted in time; it cannot tell you where the aircraft would have landed.
saying these in an interview costs you the question
- Claiming shadow mode proves the candidate resolves tickets faster
- Believing the candidate's queue receives the mirrored tickets
- Treating a high agreement rate as a quality result
- Assuming shadow reveals how a queue's backlog would change
- Thinking shadow output must be shown to a small user group