skip to content

Mirroring every ticket to the shadow router doubles your inference bill - how do you sample the shadow traffic without spoiling the comparison?

level: seniorimportance: nice to knowfreq 28%

answer

  1. duplicated inference ends windows early
  2. hash a stable key, not random
  3. widening the sample stays a superset
  4. rare segments need their own rate
  5. concurrency does not extrapolate

basics

~20 s

Sample deterministically on a hash of the ticket id so a re-run mirrors the same tickets, stratify so rare ticket types survive, record each stratum's rate for reweighting, and keep the rule independent of the routing decision.

solid answer

~40 s

Mirroring at 100% doubles scoring spend for the whole window, which is usually what cuts the window short. Sample instead, but sample in a way you can defend. Use a **deterministic hash of the ticket id** rather than a fresh random draw, so a second window covers the same tickets and the two are comparable. **Stratify**: a uniform 20% sample of a ticket type that is 0.5% of volume yields about 12 mirrored tickets an hour, so over-sample that stratum and record its rate. Never key the sample on anything correlated with the routing decision - mirroring only the tickets the incumbent was unsure about inflates the disagreement rate by construction. And accept the limit: a 5% sample answers questions about decisions, never about whether the candidate's latency holds at full traffic.

go deeper

for a junior

The idea to take away is that a shadow window costs a second full inference per mirrored ticket, so teams mirror a fraction of traffic rather than all of it.

for a middle

Explain deterministic sampling on a hash of a stable id and what it buys: repeatable windows, a superset when the rate is raised, and stability when a ticket is retried.

for a senior

Show that you would stratify for rare segments, carry the per-stratum weight on every record, keep the rule independent of the decision, and answer latency and capacity questions with a separate full-rate burst.

for a principal

The tradeoff to own is spend against evidence: how much duplicated inference a pre-exposure check is worth, and when further sampling stops changing any decision the team is going to make.

## Where the money goes A shadow window runs the candidate on top of everything already running, so the marginal cost is a second full inference per mirrored ticket plus whatever extra feature reads the mirror performs. If the candidate is heavier than the incumbent - a larger model, a richer feature set - the multiplier is worse than two. Teams almost never stop a shadow window because they have learned enough; they stop it because the duplicated bill got noticed. Sampling is what lets the window run long enough to cover a real business cycle. Work it through on this desk. At 200 tickets a minute, mirroring everything means 200 extra scoring calls a minute. At 20% it is 40 a minute, and with an 18% disagreement rate that still yields about 7 disagreements a minute, roughly 430 an hour - vastly more than anyone will adjudicate. The decision-level questions are cheap to answer; it is the coverage and the tail questions that need care. ## Deterministic, not random Draw the sample from a hash of a stable key - the ticket id - and compare it against the mirroring percentage, rather than rolling a fresh random number per ticket. Three things follow: 1. **Reproducibility.** A second window at the same percentage mirrors the same tickets, so two runs compare like with like instead of two unrelated samples. 2. **Monotone widening.** Raising the percentage from 5% to 20% is a superset: every ticket already mirrored stays mirrored, so the earlier data remains part of the larger sample. 3. **Stability across retries.** The same ticket scored twice is either mirrored both times or neither, so a retry does not silently change the sample's composition. ## Stratify, and carry the weights Uniform sampling starves exactly the segments a routing change is most likely to break. On this desk a ticket type at 0.5% of volume arrives about once a minute; at 20% mirroring that is 12 an hour, and after a day you still have a handful. The fix is to stratify - mirror 100% of the rare types and a small fraction of the dominant one - and to **record the sampling rate on every paired record**. Without the weight, every aggregate computed later is wrong in a way nobody notices, because the sample over-represents the segments you deliberately over-sampled. | Stratum | Share of volume | Mirror rate | Mirrored per hour | |---|---|---|---| | Common billing questions | 70% | 10% | 840 | | Technical faults | 25% | 20% | 600 | | Chargebacks and disputes | 4.5% | 60% | 324 | | Regulated-market escalations | 0.5% | 100% | 60 | (Rates are per hour at 200 tickets a minute, so 12,000 tickets an hour in total.) ## The rule must not know the answer The sampling condition has to be independent of anything correlated with the routing decision. Mirroring only the tickets where the incumbent's top score is low sounds efficient and produces a disagreement rate that is meaningless, because it selects the cases where any two routers are most likely to differ. If you want a focused look at the uncertain band, take it as a **named extra stratum with its own weight**, alongside an unbiased base sample - not instead of one. ## What a sample cannot answer A 5% mirror tells you almost nothing about behaviour at 100%, because the questions it cannot reach are concurrency questions and concurrency does not extrapolate from a thin sample: - **Tail latency** - the candidate's p99 under a twentieth of the traffic is not its p99 at full rate, since the queueing that produces a tail barely exists. - **Memory and pool pressure** - footprint, connection reuse and worker contention only appear when the calls overlap. - **Online feature-store load** - the read volume the mirror adds scales with the sample, so a thin sample hides the load a full mirror would create. - **Cache behaviour** - a prediction or embedding cache warms at a rate the sample never reproduces, so its hit rate is understated. Answer those separately, with a short full-rate burst during a known-quiet period or with a load test against a replayed stream, and keep the cheap ongoing sample for the decision-level questions. ## When to close the window Stop on coverage, not on a round number of days. The window should have spanned a full business cycle - weekday and weekend, the peak hour, at least one batch import or campaign spike - and the adjudicated disagreements should have stopped producing new failure classes. If every new hour of review shows the same three patterns, further duplicated spend is buying nothing the window can give you.

  • Why is a fresh random draw per ticket worse than a hash of the ticket id?
    Because it makes the sample unrepeatable and unstable. Two windows at the same rate cover different tickets, so their disagreement rates are not directly comparable; raising the rate produces a sample that is not a superset of the old one; and a retried ticket can be mirrored on one attempt and not the next, quietly changing the composition. A hash gives all three properties for free.
  • The shadow sample is drawn only from tickets arriving during business hours. What does that cost you?
    It hides the traffic that behaves least like the average. Overnight tickets skew toward other regions and languages, automated submissions and batch imports, which is exactly where a new router's feature coverage is thinnest. The latency picture is also flattered, because the quiet hours never exercise the contention the candidate meets at peak. Treat the window as covering business hours only and say so in the result.

saying these in an interview costs you the question

  • Assuming p99 latency at 5% mirroring predicts it at full traffic
  • Sampling only the tickets the incumbent scored with low confidence
  • Drawing a fresh random number per ticket instead of hashing a key
  • Reporting aggregates from a stratified sample without the weights
  • Running a uniform sample and concluding rare ticket types are fine