skip to content

A candidate substitution model is scored by replaying logged orders. What can that replay not tell you about substitutes the incumbent never offered?

level: seniorimportance: nice to knowfreq 34%

answer

  1. who created these labels?
  2. outcomes exist only for what was shown
  3. unlabelled pick: miss or drop
  4. drop rows equals optimistic
  5. report coverage beside the lift

basics

~20 s

Nothing at all: a logged order records whether the shopper took the substitute the incumbent proposed, so an item that was never offered has no outcome. The replay can only grade the candidate model where it agrees with the incumbent.

solid answer

~50 s

Labels in this system are created by the incumbent's own choices. The order log says the shopper accepted or rejected *the substitute that was shown at pick time*; every other item in the catalogue has no outcome attached. So when the candidate model's top pick differs from what was offered, the replay has no label and must do something arbitrary with the row - count it as a miss, which punishes exactly the candidate that found a better substitute, or drop it, which quietly restricts scoring to the rows where the two models agree. Both are wrong in a known direction. The practical fix is to report **coverage** - the share of the candidate's picks that have a logged outcome - beside the gate score, and to log the offered set and a small exploration slice at serving time so future replays have labels outside the incumbent's habit.

code

pseudocode · 16 lines
pseudocode
scored = 0; hits = 0; unlabelled = 0

for order in logged_orders(window):
    offered  = order.offered_substitute        // what the incumbent showed at pick time
    outcome  = order.accepted                  // label exists ONLY for `offered`
    features = feature_values_as_of(order.timestamp)
    pick     = top1(cand_model.rank(order.oos_item, features))

    if pick == offered:
        scored = scored + 1
        hits   = hits + outcome                // a real shopper decision
    else:
        unlabelled = unlabelled + 1            // nobody ever saw `pick`

coverage   = scored / (scored + unlabelled)    // report this beside the score
acceptance = hits / scored                     // measured only where the models agreed

go deeper

for a junior

Grasp the core fact: a log can only record what a shopper reacted to. If the model would have offered a different item, there is no record of what would have happened.

for a middle

Explain what the replay does with an unlabelled pick and which way each choice bends the score - counting misses is pessimistic, dropping rows is optimistic and changes the population being scored.

for a senior

Show the operational consequence: the gate quietly prefers candidates that agree with the incumbent. Report coverage beside the metric and change what the serving path logs so the next gate has labels outside the current policy.

for a principal

The call is how much current experience to spend on future evaluability - a standing exploration slice costs shoppers something today and is the only thing that makes tomorrow's gate able to see past the incumbent's habits.

## What a logged order actually contains When a picker finds an ordered item out of stock, the serving path proposes a substitute; the shopper accepts it, rejects it, or the order goes short. The log records that event: the out-of-stock item, **the substitute that was offered**, the outcome, and the context. What it does not record - because it never happened - is what the shopper would have done with a different substitute. That single sentence is the whole difficulty. The labels available to the offline gate were manufactured by the policy currently in production. The incumbent model version chose what to show; the shopper could only react to that. ## The two ways to handle an unlabelled pick, and their directions When the candidate model version's top pick is an item the incumbent never offered for that order, the replay has to decide something. | handling | what it assumes | direction of the error | |---|---|---| | count it as a miss | any unoffered item would have been rejected | pessimistic - punishes a candidate whose value is finding substitutes outside the incumbent's habit | | drop the row | the dropped orders resemble the kept ones | optimistic - scores the candidate only where it agreed with the incumbent, and silently changes which orders were evaluated | | count it as a hit | any unoffered item would have been accepted | unusable - there is no argument for it | Neither of the first two is neutral, and the choice is rarely written down. A gate that drops rows is the common default because it produces the nicer number, and it is the one that most often surprises a team at launch. ## Why the gate drifts toward agreement with the incumbent Follow the consequence. A candidate that reorders items the incumbent already surfaced is fully scoreable. A candidate that proposes genuinely different substitutes has most of its picks unlabelled, so it is scored on a shrinking, unrepresentative slice of its own behaviour. Over several gate cycles this produces a systematic preference for candidates that stay close to the current policy - the gate becomes a conservatism filter, and the improvements it keeps passing are the small ones. This is also why an offline lift can be large and the live effect small. If the measured lift came almost entirely from rows where both models picked the same item, the number describes a narrow disagreement, not the candidate's real behaviour in production. ## Report coverage beside the score The minimum discipline costs nothing: alongside the offline gate metric, report **coverage** - the fraction of scored orders where the candidate's top pick had a logged outcome. A lift of two points at 80% coverage and a lift of two points at 25% coverage are different claims, and the second one should not clear a shipping bar. Slice coverage the same way as the metric, because it is usually lowest exactly where the incumbent's habits are strongest. ## What to log at serving time so the next replay is better The long-term fix is upstream: the offline gate can only be as honest as what the serving path recorded. - **Log the decision, not just the outcome.** Store the full candidate set considered, the item offered, its position if several were shown, and the incumbent's scores - so a future replay can tell "never offered" apart from "offered and refused". - **Log the inputs as read.** Persist the feature values the serving path actually used, with their timestamps, so the replay can reproduce the decision instead of recomputing it from batch-complete data. - **Reserve a small exploration slice.** If a small, deliberately varied share of substitution offers departs from the incumbent's top pick, the log eventually contains outcomes outside the current policy's habit, and future gates have labels where they need them. This is a standing cost paid for future evaluability, and it has to be sized so the shopper experience stays acceptable. - **Log the inventory state at pick time**, since whether an item was even available bounds what any model could have offered. ## Saying it in a design round The checkable answer is one line - *the replay has no outcome for an item nobody saw* - and the useful follow-through is three: name the direction each handling biases the score, report coverage next to the number, and change what the serving path logs so the next gate is not stuck with the same blind spot.

  • Which handling of unlabelled picks is safer as a default?
    Counting them as misses, because its error is pessimistic and visible: a candidate that departs from the incumbent is under-credited, and you will know it because coverage is low. Dropping the rows silently narrows the evaluated population and produces the flattering number, which is the one that surprises the team at launch.
  • What does coverage tell you that the gate metric does not?
    How much of the candidate model's actual behaviour was graded. Coverage is the share of scored orders where the candidate's top pick had a logged outcome; a lift measured at 25% coverage describes a narrow band of agreement with the incumbent, not the model you would ship.
  • Why log the full candidate set the serving path considered, not just the offered item?
    So a later replay can distinguish 'never offered' from 'offered and refused', and can tell whether an item was available at all at pick time. Without it, every disagreement between models looks identical to the gate, and the coverage number cannot be computed.

It is a menu test where the only dishes that ever got reviews are the ones the previous chef chose to put on the menu. The new chef's best dish has no reviews - not a bad score, no score at all.

saying these in an interview costs you the question

  • Treats a rejected substitute and a never-offered item as the same evidence.
  • Drops unlabelled rows and reports the remaining lift without qualification.
  • Assumes any substitute the incumbent skipped would have been rejected.
  • Reports a gate score without saying what share of picks had a label.
  • Logs only the accepted outcome and not the decision that produced it.