skip to content

questions

21

In an offline gate for a candidate model, why must the shipping score come from a held-out data split, not the training rows?

level: juniorimportance: must knowfreq 72%

answer

  1. whose rows are these, exactly?
  2. recall is not prediction
  3. capacity lifts the training score
  4. held-out rows stand in for tomorrow
  5. both model versions, identical rows

basics

~20 s

A model scored on the rows it was fitted on is being asked to recall, not predict, so that score improves with capacity and understates the errors it will make on the next order. Only rows the fit never saw estimate that.

solid answer

~40 s

The offline gate asks one question: on orders this model has never seen, does the candidate substitution ranker propose replacements shoppers accept more often than the model in production does? A score taken on the training rows cannot answer it, because a flexible model raises that number by fitting particulars of individual rows - this store, this item, this hour - that will not recur in the same combination. So the gate holds back a slice of logged orders and scores the candidate model version and the incumbent model version on *exactly those rows*, through the same feature-computation path. Two cautions: a held-out **data** split is not a **traffic** split - no shopper is served by it - and a slice stops being held out once you have tuned against it repeatedly.

go deeper

for a junior

Recall the one line: a model is graded on rows it has not seen, because a score on its own training rows measures recall rather than skill. Say the held-out slice is historical data, not live traffic.

for a middle

Explain the mechanics: capacity lifts the training score, the validation slice decays as you tune against it, and the gate is a like-for-like comparison of two model versions on identical rows and identical feature values.

for a senior

Show where the honest split gets compromised in a real pipeline - reused slices, leakage from post-order features, a baseline score copied from an old run - and say what you would change in the gate to close each one.

for a principal

Frame the gate as the cheapest filter in a chain of instruments and be explicit about what it is licensed to decide. The tradeoff is how much evidence you demand offline before spending live traffic on the question.

## What the offline gate is estimating An offline gate is the check a **candidate model version** passes before anyone lets it near live requests. For a substitution ranker in an online grocery service - the model that proposes a replacement when a picker finds an ordered item out of stock - the gate asks exactly one thing: *on orders this model has never seen, does it propose substitutes shoppers accept more often than the incumbent model version does?* The quantity being estimated is behaviour on data drawn from the same process but kept out of the fit. That is a prediction about the future made from the past, and it only holds if the rows used to make it were genuinely beyond the fit's reach. ## Why the training score climbs whatever you do A model scored on its own training rows is being asked to recall, not to predict. Given enough capacity it can push that number up by fitting properties of individual rows - this store, this item, this hour of this weekday - that will not recur in the same combination. Three consequences follow: - The training score **improves with capacity over a wide range**, so it cannot choose between two candidate model versions: the bigger one nearly always wins on it. - It hides the quantity you care about. The gap between the training score and the unseen-data score *is* the risk, and the training score cannot show a gap to itself. - It rewards the behaviour production punishes - memorising the substitute that worked last winter instead of learning the property that made it work. ## The shape of the split Most gates hold back more than one slice, because the slices are licensed to decide different things. | slice | who touches it | what it may decide | |---|---|---| | training rows | the fit | the model's parameters | | validation rows | the training loop, repeatedly | hyperparameters, early stopping, feature choices | | final scoring window | the gate, once | whether the candidate model version clears the shipping bar | | live traffic | nobody, until the gate passes | the launch itself | The distinction a design round is listening for: **a held-out data split is not a traffic split.** The held-out rows are historical orders replayed offline and no shopper is served by them. A traffic split assigns real users to a model version and is a different instrument, reached only after this one passes. ## The comparison, not the number An absolute held-out score is close to meaningless - "acceptance 0.61" tells nobody whether to ship. The gate is a **comparison**, and it is fair only when both sides see the same thing: - the same rows, from the same scoring window; - feature values produced by the same computation path, so a difference in scores is a difference in models rather than a difference in feature code; - the incumbent model version **re-scored in the same run**, not a number copied from its own gate months ago, because the data and the feature definitions have moved since. A candidate model that wins on rows the incumbent never saw, or with features the incumbent never got, has not won anything. ## Where a held-out score quietly stops being honest 1. **Reuse.** Every decision taken after looking at a slice leaks a little of it into the model. A slice that has picked the learning rate forty times is a validation slice; treat the survivor's score as optimistic and keep a final scoring window that was read once. 2. **Leakage through features.** A feature computed from events *after* the order - the eventual refund, the day's closing inventory count - makes any split look excellent and cannot be reproduced at pick time. 3. **The wrong kind of split.** Shuffling rows at random hides a change in demand regime; scoring forward in time does not. 4. **Population drift.** The held-out window is still the past. It cannot contain a store that opened last week or a supplier that changed last month. ## What clearing the gate is worth Passing means the candidate model version is not obviously worse and is worth the cost of a live comparison. It is a veto with teeth, not evidence of a win. The mechanisms that make an offline number optimistic - a mismatch between the offline and serving feature paths, feature values that are older at serving time than in the replay, labels that exist only where the incumbent offered something, and a gate metric that is not the business outcome - are all still ahead of you. Presenting a held-out number as the launch decision is the classic junior answer; the correct framing is that it is the cheapest filter in the chain and the least conclusive one.

  • Why score the incumbent model version again in the same gate run instead of reusing its recorded score?
    Because the incumbent's old number came from different rows, possibly different feature definitions, and a different demand period. Re-scoring it on this window through today's feature path makes the difference attributable to the models. Otherwise the comparison measures how the world moved, not which ranker is better.
  • If a slice has been used for early stopping all quarter, can it still gate the launch?
    No. Each look at it transfers a little of it into the model, so the survivor's score on that slice is optimistic. Keep it as the validation slice and reserve a separate scoring window - ideally a later time period - that the gate reads once.
  • What makes a feature leak a label rather than predict it?
    Timing. If the feature's value depends on events that occurred after the moment the prediction would have been made - the refund, the final stock count, the shopper's later basket - it is unavailable at pick time and inflates every split it appears in. The test is whether the serving path could compute it from data present at the request.

saying these in an interview costs you the question

  • Quotes the training-set acceptance metric as the gate result.
  • Assumes a higher training score means a better production model.
  • Tunes against the held-out slice for weeks and still calls it held out.
  • Treats a held-out data split as a share of live traffic.
  • Compares the candidate model's held-out score with the incumbent's months-old number.
  • Builds a feature from events that happen after the order is placed.
open as a page

A candidate ticket router scores live tickets in shadow mode - why can that window not show whether it routes better than the incumbent?

level: juniorimportance: must knowfreq 64%

basics

~20 s

Shadow mode compares mechanics, not outcomes. The incumbent's queue choice is what agents actually work, so no ticket is ever placed where the candidate said, and resolution time and reroute rate are never produced for the disagreements.

open as a page

In a sponsored-listing strip, how does team-draft interleaving compare two ranking policies within a single user's list?

level: middleimportance: must knowfreq 58%

basics

~20 s

Team-draft interleaving fills the strip by alternating drafts: a coin flip picks who starts, then each policy takes its highest-ranked listing not already placed. Every slot carries a hidden team tag, and a click credits the team that drafted that slot.

open as a page

How do you wire the mirrored call to a shadow ticket-routing model so it cannot add latency or failures to the live request?

level: middleimportance: must knowfreq 55%

basics

~20 s

Keep the candidate off the critical path. Answer the live request from the incumbent first, hand a copy to a bounded background worker with its own pool and timeout, and let shadow failures only increment a counter.

open as a page

Why can a candidate send-time model ramped to 10% of users deliver two notifications to one user in a day?

level: middleimportance: must knowfreq 62%

basics

~20 s

Because the ramp assignment was not sticky. The same user fell on the candidate side in one planning run and on the incumbent side in another, so both sides queued a send for that day. Deterministic per-user assignment with one claimed send key prevents it.

open as a page

Your permanent sponsored-ranking holdout shows +3.1% while the year's shipped launches summed to +5.6% - why the gap?

level: seniorimportance: must knowfreq 52%

basics

~20 s

Short launch readouts are not additive. Each was measured while the change was still novel, only the launches that read positive were kept, and launches touching the same slots cannot both claim their full effect. The holdout measures what actually survived together.

open as a page

Why is interleaving invalid for a sponsored-strip change that adds a fourth slot and widens the candidate pool?

level: seniorimportance: must knowfreq 46%

basics

~20 s

Interleaving compares two orderings of one shared pool inside one unchanged container. A fourth slot means no single merged strip represents both policies, and a widened pool lets the candidate win on listings the incumbent never had rather than on better ordering.

open as a page

A substitution ranker cleared the offline gate by a wide margin but barely moved live acceptance. How would you find the cause?

level: seniorimportance: must knowfreq 80%

basics

~20 s

Test four mechanisms in cost order: diff the feature vectors the serving and offline paths build for the same requests, re-run the replay with values aged to production's read-time age, check label coverage, then check refunds.

open as a page

A candidate send-time model is rolled back to the previous artifact, yet send hours stay wrong and nothing errors — what else must the rollback restore?

level: seniorimportance: must knowfreq 74%

basics

~20 s

The feature definitions the candidate introduced, and the queued work planned under it. Reverting only the artifact leaves the previous model reading features whose meaning changed, which produces plausible wrong hours with no error, and leaves already-queued sends carrying the candidate's decisions.

open as a page

While a candidate send-time model serves 10% of users, why does the previous model version stay loaded and serving too?

level: juniorimportance: should knowfreq 50%

basics

~20 s

A percentage ramp splits live traffic between two model versions instead of replacing one with the other. The previous version keeps serving the other 90%, and keeping it loaded is what makes reverting a routing change rather than a redeploy.

open as a page

What must a candidate model's shipping bar fix in advance, besides a number the offline gate metric has to clear?

level: middleimportance: should knowfreq 46%

basics

~20 s

A shipping bar is a decision rule written before the run: which offline gate metric on which scoring window, which slices may not regress, a margin bigger than rebuild-to-rebuild variation, and the cost and latency envelope the candidate model must also fit.

open as a page

A substitution ranker's offline gate shuffles a year of logged orders at random. What does a time-ordered split fix?

level: middleimportance: should knowfreq 58%

basics

~20 s

A random shuffle scores the candidate model on orders from the same days it trained on, so a holiday demand shift sits in both halves. A time-ordered split trains before a cut and scores after it, the way production always runs.

open as a page

In a candidate send-time model's traffic ramp, what sets how long one step must bake before the next increase?

level: middleimportance: should knowfreq 46%

basics

~20 s

The latency of the guardrail signals the step is supposed to be halted on. A step must span at least one full cycle of the action being changed plus the users' reaction lag, or the halt it promises can never fire.

open as a page

Interleaving on the sponsored strip picks the candidate ranker in two days - why doesn't that settle the launch?

level: seniorimportance: should knowfreq 42%

basics

~20 s

An interleaving run returns a preference between two orderings of one list, measured in credited clicks. It carries no session-level effect size, no guardrail outside the strip, and no effect that needs weeks to appear, so the candidate still walks the ramp.

open as a page

In a shadow window the candidate router disagrees with the incumbent on 18% of tickets - what must each paired record hold to make those diffs diagnosable later?

level: seniorimportance: should knowfreq 51%

basics

~20 s

Enough to reproduce and attribute the diff: both decisions and both raw scores, both model versions and both feature-definition versions, the feature values used or a snapshot reference, per-call latency, and the segment fields you will group by.

open as a page

Your shadow ticket router reuses the live serving code, so it also writes prediction logs, updates the prediction cache and emits routing events - what must you isolate first?

level: seniorimportance: should knowfreq 47%

basics

~20 s

Every write the scoring path performs. The shadow run must not populate the shared prediction cache, emit routing events, spend shared quota or reuse idempotency keys, and its prediction rows must be tagged so training never reads them.

open as a page

The opt-out guardrail trips at the 25% step of a model ramp — what does halting the ramp leave running that a kill switch does not?

level: seniorimportance: should knowfreq 55%

basics

~20 s

Halting freezes the percentage where it is, so the 25% already assigned keep being served by the candidate model version. It stops the ramp growing; it does not remove exposure. A kill switch routes everyone back to the incumbent immediately.

open as a page

Your permanent holdout keeps 1% of users out of every sponsored-ranking launch - would you keep, shrink or retire it?

level: principalimportance: should knowfreq 31%

basics

~20 s

There is no default answer: the holdout is the one instrument that prices a year of launches together, and it costs a permanently worse experience for 1% of users plus a frozen serving path maintained forever. Decide by naming the decision it feeds.

open as a page

Why is a permanent sponsored-ranking holdout not a clean counterfactual once the marketplace adapts around it?

level: seniorimportance: nice to knowfreq 28%

basics

~20 s

Only the serving policy is held out. Sellers, listings, bids and the retrieval index are shared with the launched population, so part of every launch reaches held-out users - and if the holdout's policy is frozen instead, its own staleness is counted as launch effect.

open as a page

A candidate substitution model is scored by replaying logged orders. What can that replay not tell you about substitutes the incumbent never offered?

level: seniorimportance: nice to knowfreq 34%

basics

~20 s

Nothing at all: a logged order records whether the shopper took the substitute the incumbent proposed, so an item that was never offered has no outcome. The replay can only grade the candidate model where it agrees with the incumbent.

open as a page

Mirroring every ticket to the shadow router doubles your inference bill - how do you sample the shadow traffic without spoiling the comparison?

level: seniorimportance: nice to knowfreq 28%

basics

~20 s

Sample deterministically on a hash of the ticket id so a re-run mirrors the same tickets, stratify so rare ticket types survive, record each stratum's rate for reweighting, and keep the rule independent of the routing decision.

open as a page