skip to content

What must a candidate model's shipping bar fix in advance, besides a number the offline gate metric has to clear?

level: middleimportance: should knowfreq 46%

answer

  1. a rule, not a number
  2. against what, on which rows?
  3. measure the rebuild spread first
  4. slices, cost and latency clauses
  5. written before the scores arrive

basics

~20 s

A shipping bar is a decision rule written before the run: which offline gate metric on which scoring window, which slices may not regress, a margin bigger than rebuild-to-rebuild variation, and the cost and latency envelope the candidate model must also fit.

solid answer

~40 s

"Beats the incumbent" is not a bar. Fix five things before the candidate model is scored: (1) the **offline gate metric** and the scoring window it is computed on; (2) the comparison - the incumbent model version re-scored on the same rows through the same feature path; (3) a **margin** that exceeds how much the same configuration moves between two rebuilds, so the gate does not ship noise; (4) the slices that may not regress - region, category, high-value baskets - because an aggregate win hides a regional loss; (5) the non-quality envelope: serving cost per thousand predictions, the latency the substitution call must fit, and any new feature dependency the candidate introduces. Writing it afterwards lets the window and the slices be chosen to fit the result.

code

pseudocode · 19 lines
pseudocode
gate(cand_model, incumbent_model, window, bar):
    rows = replay_orders(window)                  // logged orders, features as of each order
    cand = score(cand_model, rows)
    base = score(incumbent_model, rows)           // same rows, same feature path, same run

    if cand.cost_per_1k > bar.cost_ceiling:
        return REJECT("cost envelope")
    if cand.p99_ms > bar.latency_ceiling:
        return REJECT("latency envelope")

    for slice in bar.protected_slices:            // region, category, high-value baskets
        if cand.acceptance[slice] < base.acceptance[slice] - bar.slice_tolerance:
            return REJECT("slice regression: " + slice)

    margin = cand.acceptance.all - base.acceptance.all
    if margin < bar.min_margin:                   // min_margin > rebuild-to-rebuild spread
        return REJECT("inside rebuild noise")

    return PASS                                   // licenses a live comparison, not a launch

go deeper

for a junior

Know that the bar is agreed before the model is scored, and that it covers more than accuracy: the model also has to fit the cost and the time budget the serving path allows.

for a middle

Be able to list the clauses and justify each - metric and window, the re-scored baseline, a margin above rebuild noise, protected slices, the envelope - and explain why writing them afterwards voids the test.

for a senior

Show how you would measure the rebuild spread in your own pipeline, choose the protected slices from where the business actually loses money, and record the bar as an artifact beside the run.

for a principal

The open call is how strict to be: a high bar burns candidate versions that would have won live, a low one spends live comparison capacity on noise. Tie the bar's strictness to the cost of the next instrument.

## The bar is a decision rule, not a number A **shipping bar** is what the offline gate compares against. Teams usually say it as a number - "+1 point of substitute acceptance" - but a number alone is not decidable: it does not say on which rows, against what, how much is enough to be real, or what else the candidate model version must not break. The useful form is a short rule, written down before the scoring run, that a reviewer can apply without discretion. ## The clauses that make it decidable | clause | what it states | what it prevents | |---|---|---| | metric and window | which offline gate metric, on which scoring window | swapping to the metric that happened to move | | comparison | the incumbent model version re-scored on the same rows, same feature path | a win against a stale or differently-fed baseline | | margin | how far above the incumbent the candidate must land | shipping a difference smaller than rebuild noise | | slices | segments that may not regress, and by how much | an aggregate win that hides a regional loss | | envelope | cost per prediction, latency, new feature dependencies | a quality win the serving path cannot afford | | authority | what passing licenses | mistaking the gate for a launch decision | ## Why the margin must clear rebuild variation Retrain the same configuration twice - new random seed, one more day of data, a different shuffle order - and the offline gate metric moves. That movement is the floor of what the gate can resolve. If two rebuilds of the *incumbent* configuration differ by 0.3 points of acceptance, a candidate that wins by 0.2 has shown nothing, and a bar set at "any improvement" will ship a coin flip roughly half the time. So the margin is measured, not guessed: rebuild the incumbent configuration a few times, look at the spread, and set the bar outside it. This is a property of the training pipeline, and it changes when the pipeline changes - re-measure it when feature definitions or the sampling of training examples move. ## The slices are part of the bar Aggregate acceptance is a weighted average dominated by the largest stores and the commonest categories. A candidate model version can gain a point overall while losing three points on chilled goods or in one region whose stores substitute differently. If the launch will be global, the bar has to say which slices are protected and what regression is tolerated in each - otherwise the gate passes a model that is a net win and a local disaster. Keep the slice list short and stable; a slice list assembled after the scores are in is not a test. ## The envelope clauses Quality is not the only thing a candidate model has to clear: - **Cost**: a candidate that needs a heavier scoring pass raises the per-prediction bill on every out-of-stock event, which for a large grocery service is a standing cost, not a one-off. - **Latency**: the substitution decision happens while a picker waits. A ranker that cannot answer inside its slice of the request budget is not shippable at any accuracy. - **New dependencies**: if the candidate needs a feature the online feature store does not serve today, the gate has to know, because the launch now carries a feature-pipeline change as well as a model artifact change. ## Declaring it before the run The reason to write the bar first is mechanical: after the numbers are in, every free choice becomes a way to reach the answer you want. Which window, which slices, which metric, whether to include the week of the supply incident - each is defensible in isolation and, chosen after the fact, together they will pass almost anything. A bar recorded before the scoring run turns those choices into a test. Record it as a small artifact next to the run - metric, window, margin, slices, envelope - so the decision is reproducible months later. ## What passing actually licenses Clearing the bar earns the candidate model version the *next* instrument, not the traffic. The offline gate is the cheapest filter and the least conclusive: it is computed on replayed history, with labels the incumbent's own behaviour shaped, on feature values assembled in batch rather than read under a deadline. A team that treats a cleared bar as a launch decision has skipped everything the gate cannot see. The honest sentence is: *the candidate is not obviously worse, so it is worth spending live comparison on.*

  • How do you set the margin rather than guess it?
    Rebuild the incumbent configuration several times with a different seed and one more day of data, and look at how far the offline gate metric moves between rebuilds. That spread is the floor of what the gate can resolve; put the margin outside it, and re-measure whenever the feature definitions or example sampling change.
  • The candidate model gains a point overall but loses two on chilled goods. Does it pass?
    Only if the bar said so in advance. A protected-slice clause exists precisely for this: either chilled goods was on the list, in which case it fails and goes back, or it was not, in which case you ship and add the slice - you do not decide now, with the number in front of you.
  • Should the offline shipping bar include a business metric like refunds avoided?
    If it can be computed on replayed orders, include it as a protected slice rather than the headline: refunds and basket value are the outcome you care about, but their offline form is weak because it is measured only where the incumbent offered something. Keep the launch metric explicit in the bar so nobody confuses acceptance with value.

saying these in an interview costs you the question

  • States the bar as 'better than the current model' and stops.
  • Chooses the scoring window after seeing the candidate's score.
  • Sets a margin smaller than the pipeline's rebuild-to-rebuild spread.
  • Checks aggregate acceptance only, with no protected slices.
  • Leaves serving cost and latency out of the gate entirely.
  • Treats a cleared bar as the decision to launch.