skip to content

A substitution ranker cleared the offline gate by a wide margin but barely moved live acceptance. How would you find the cause?

level: seniorimportance: must knowfreq 80%

answer

  1. what could the gate not see?
  2. same feature vector, both paths?
  3. production reads older values than the replay
  4. was the lift measured on agreement only?
  5. acceptance is not refunds avoided

basics

~20 s

Test four mechanisms in cost order: diff the feature vectors the serving and offline paths build for the same requests, re-run the replay with values aged to production's read-time age, check label coverage, then check refunds.

solid answer

~50 s

Four mechanisms explain most such gaps, and each has a test. **Train-serve skew**: re-score a sample of live requests through both the offline and the serving feature path and diff field by field - a different default, unit or trimming rule shows up immediately. **Feature staleness**: the replay reads values assembled in batch, while the serving path reads an online feature store refreshed on a lag, so re-run the replay with every value aged by the real read-time age and see whether the lift survives. **Exposure in the labels**: check what fraction of the candidate's picks had a logged outcome; a lift measured mostly on agreement with the incumbent will not reproduce. **Metric mismatch**: acceptance is not the business outcome - confirm whether accepted substitutions actually reduce refunds and preserve basket value. A fifth possibility is simply that the scoring window and the live weeks are different populations.

go deeper

for a junior

Know that an offline win is a prediction about production, not a result from it, and that the features a model gets when it is serving may not be the ones it was graded on.

for a middle

Name the four mechanisms - skew, staleness, exposure in the labels, metric mismatch - and describe the test for each rather than reasoning about which is most likely.

for a senior

Run the diagnosis in cost order, get the direction of the staleness claim right, and say what each finding changes about the gate itself so the same surprise does not recur.

for a principal

The standing question is how much the gate should be made to resemble production - aged features, exploration logging, business outcomes in the bar - against the cost of every one of those, and what you stop believing offline numbers for.

## Start by admitting what the gate could not see The offline gate is computed on replayed history, from labels the incumbent's own behaviour produced, on feature values assembled in batch rather than read under a request deadline. Every one of those three is a place where the offline world and the live world can differ. A big offline lift that lands near zero is not surprising; it is the default outcome when nobody checked the three. ## The four mechanisms and the test for each | mechanism | what happened | the cheapest test | |---|---|---| | train-serve skew | the serving path computes at least one feature differently from the offline path | re-score a sample of live requests through both paths and diff field by field | | feature staleness | the online feature store's values are older at read time than the replay's were | re-run the replay with every value aged by the measured read-time age | | exposure in the labels | the lift was measured mostly where the candidate agreed with the incumbent | compute the share of the candidate's picks that had a logged outcome | | metric mismatch | acceptance moved, the outcome the launch was for did not | check refunds and basket value on the same accepted substitutions | **Skew** is first because it is the cheapest and the most common. The offline path computes a feature over a complete day of events with no deadline; the serving path computes it in milliseconds, from whatever has arrived, with its own default for a missing value. A single field - a unit, a rounding rule, a null filled with zero on one side and with a category mean on the other - is enough to feed the model a vector it never saw in training. The test is a paired diff, not an argument: take live requests, produce the feature vector both ways, and compare. **Staleness** is the direction people get backwards. The replay is the *fresher* world: it reads values recomputed after the fact, complete and correct as of each order. Production reads an online feature store refreshed on an interval, so an inventory-derived feature can be up to one refresh interval old when the picker is standing in the aisle - which is exactly when the substitute recommended by the candidate may itself be out of stock. Measure the actual age of values at read time, then re-run the replay with values artificially aged by that amount. If the lift collapses, freshness was the gate's blind spot, and the fix is to age the replay permanently so the gate stops overstating. **Exposure** is the label problem: outcomes exist only for substitutes the incumbent actually offered, so a lift can be measured almost entirely on the rows where the two models chose the same item. A coverage number settles it. **Metric mismatch** is the one that survives all the technical fixes. Acceptance of a proposed substitute is a proxy: a model can raise it by proposing obviously-similar, cheaper items that shoppers wave through and then return, leaving refunds flat or worse and basket value down. If live acceptance had moved and refunds had not, that is the diagnosis; if acceptance itself barely moved, this is not your explanation and the first three are. ## An order that costs least and rules out most 1. **Confirm the readings are comparable at all** - same slices, same definition of acceptance, same population of orders. A surprising number of vanished lifts are two different denominators. 2. **Paired feature diff** on a sample of live requests. This is hours of work and eliminates the most likely cause. 3. **Aged replay** using the measured read-time age of the online feature store's values. 4. **Coverage** of the candidate's picks in the gate's scoring window. 5. **Outcome check**: for the accepted substitutions, did refunds fall and basket value hold? ## What each finding changes about the gate - Skew: fix the serving path or the offline definition so one computation is shared, then re-gate. Nothing else is trustworthy until this is closed. - Staleness: age the replay's features by the production read-time age as standing practice, so the shipping bar is set against the world the model actually sees. - Exposure: report coverage beside the gate metric and begin logging a small exploration slice so future gates have labels outside the incumbent's habit. - Metric mismatch: change what the bar protects - keep acceptance as the gate metric if it is the only thing measurable offline, but state the business outcome explicitly so nobody reads one as the other again. ## The sentence an interviewer is listening for *The offline number and the live number are measuring different things, and the job is to name which difference.* A candidate who reaches for "we need a bigger model" or "the bar was too low" has skipped the diagnosis; the four mechanisms above are checkable in a day and each one changes the gate afterwards.

  • Which way does feature staleness bias the offline gate, and why do people get it backwards?
    The replay is fresher, so the gate overstates. Offline features are recomputed after the fact from complete data, while the serving path reads an online feature store refreshed on an interval and gets values up to one interval old. People assume the replay uses 'old data' because the orders are historical - but the orders are old, the feature values are not.
  • How do you test train-serve skew without waiting for a live experiment?
    Take a sample of real requests, produce the feature vector through the serving path and through the offline path for the same timestamps, and diff field by field. Report the fraction of requests with any mismatch and the worst offending fields; a unit, a null default or a trimming rule usually accounts for most of it.
  • Live acceptance rose but refunds and basket value are flat. What have you learned?
    That the gate metric is not the outcome the launch was for. The candidate is persuading shoppers to take substitutes without reducing the refunds or protecting the basket value that justified the project. The bar has to name the business outcome, and the launch decision needs a reading on it rather than on acceptance alone.
  • Could the gap simply be that the two periods differ?
    Yes, and it is worth ruling out early: the scoring window is historical and the live weeks are not, so a supply disruption, a promotion or a seasonal shift can move acceptance on its own. Compare the incumbent's own acceptance in both periods; if that moved too, the difference is the world, not the candidate.

saying these in an interview costs you the question

  • Concludes the model is undertrained and asks for a bigger one.
  • Raises the shipping bar instead of finding the divergence.
  • Assumes the offline replay used staler features than production does.
  • Reads acceptance of a substitute as refunds avoided.
  • Compares two readings computed over different populations or denominators.
  • Never checks whether the serving path builds the same feature vector.