skip to content

Why do offline eval wins often fail to reproduce online, as when a grocery search-query rewriter gains 9 points offline but loses 0.4% cart conversion?

level: seniorimportance: must knowfreq 58%

answer

  1. two measurements, different validity
  2. the sample moved under you
  3. proxy metric is an assumption
  4. harness has no user in it
  5. rule out latency before behaviour

basics

~20 s

Three causes dominate: the frozen set no longer matches live traffic, the offline metric proxies something users do not reward, and the harness has no user in it to react. A large offline gain on the wrong population or the wrong proxy buys nothing online.

solid answer

~50 s

Start by separating the three failure classes rather than guessing. **Distribution shift** — the frozen set was sampled from an older traffic mix. A grocery rewriter's eval set built from last quarter's queries carries no seasonal produce terms, and today's traffic is full of them; the 9-point gain is real on a population you no longer serve. **Metric mismatch** — the offline score rewards output properties (well-formed, more specific, more relevant by a rubric) that do not map to the outcome the business is paid in. Rewritten queries can be objectively better and still surface a narrower result set that shoppers do not add to cart. **Missing context and user response** — the harness sees a bare query; production has a cart, a store, a substitution history and a live user who reacts to unfamiliar results and rephrases. The diagnosis is empirical: re-score the offline set on a fresh traffic sample to test shift, then slice the online result to find which segment carries the loss.

go deeper

for a junior

Know that a better score in testing does not guarantee a better product, and that the frozen examples used for scoring may not look like what real users send today.

for a middle

Explain the concrete mechanisms — an eval set drawn from an older traffic mix, a metric that measures output properties rather than user outcomes, and eval records missing the context production supplies.

for a senior

Demonstrate a diagnostic order: verify the served system equals the evaluated one, check latency and errors, re-score on fresh inputs, then slice the online result by segment before questioning the metric itself.

for a principal

Own the policy question: what evidence is required before a rollout, who arbitrates when harness and production disagree, and how the offline suite is kept honest as a predictor rather than allowed to drift into a number the team optimizes for its own sake.

## The gap is the normal case, not an anomaly Every experienced team has shipped a change that won convincingly on the harness and lost in production. It is worth internalizing that this is the *expected* behaviour of two measurements with different validity, not evidence that someone did evaluation wrong. Offline evaluation is a controlled experiment on a chosen sample with a proxy outcome; online evaluation is an uncontrolled measurement on the real population with the real outcome. The gap between them has a small number of recurring causes, and a senior engineer names them and tests for each rather than re-running the harness harder. ## Cause one: distribution shift The frozen set was drawn at some past moment. Traffic moves — seasonally, with marketing pushes, with a new client app, with a new customer segment. A grocery search-query rewriter whose eval set came from last quarter's logs contains no summer produce vocabulary; live traffic in July is dense with it. The rewriter's 9-point gain may be genuine on staple-item queries and neutral-to-harmful on the seasonal terms that now dominate volume. Weighted by real traffic, the aggregate gain shrinks or inverts. The test is cheap: take a fresh sample of live inputs, score both variants on it, and compare against the frozen-set result. A large discrepancy is shift; a small one sends you to the other causes. ## Cause two: metric mismatch Offline scoring measures properties of the output because that is all it can see. Those properties are chosen as *proxies* for value, and the proxy relationship is an assumption. Rewritten queries that are more specific score higher on relevance rubrics and may retrieve a narrower, cleaner result set — while shoppers, who browse rather than target, were converting off the messier broad set. Similarly, a support answer that is more faithful and better hedged can score higher and drive more escalations because it now refuses more. The symptom that identifies this cause: the offline metric moves, the online *quality* signal is flat or fine, and the business metric moves the other way. That combination means the proxy, not the population, is the problem. The fix is to change what offline measures, usually by adding cases and rubric criteria that encode the behaviour the business actually rewards. ## Cause three: missing context Harness inputs are almost always thinner than production requests. In production the rewriter may run with the shopper's cart, store, dietary preferences, substitution history and prior session queries in context; the frozen record often preserves only the query string. A change that helps in the thin case can be redundant or actively harmful in the rich one — for example re-specifying constraints that personalization already applied. The reverse also occurs: a change that needs context to help looks flat offline and wins online. The test is to reconstruct full production context for a subset of eval cases and re-score. If the gain disappears with realistic context, you have found it. ## Cause four: user response A harness has no user. Real users adapt: they learn what phrasing works, retry when results look unfamiliar, and abandon when a familiar interface behaves in a new way. Even a strictly better system can lose in the short run because habituated users are disrupted — an effect that decays over days and that a one-week read can misattribute as a permanent loss. Novelty runs the other way: a change can win early because it is new and regress to baseline once the novelty passes. Both mean a single short read can lie in either direction, which is why serious teams check whether the effect is stable over the experiment window rather than only whether it is significant. ## Cause five: the boring ones Before reaching for behavioural explanations, rule out the mundane. The online variant may not be the offline variant — different system prompt assembly, different truncation, different retrieval settings, a different model version in the serving path. Latency is a frequent silent killer: a rewriter that adds a model call before search adds hundreds of milliseconds to every query, and in a conversion funnel latency alone can cost more than quality gains. Errors and timeouts under real concurrency do not exist in a harness that runs sequentially with generous retries. ## The diagnostic order A usable sequence: (1) confirm the two variants are the same system — pin versions and diff the assembled prompts; (2) check latency and error rate in the treatment arm; (3) re-score offline on a fresh traffic sample to test shift; (4) slice the online result by segment, query type and device to locate where the loss lives; (5) re-score with full production context; (6) only then question the metric itself. ## What to do with the answer Each cause has a different repair. Shift is repaired by refreshing and reweighting the offline set to live traffic. Mismatch is repaired by changing the offline metric or rubric, not by ignoring the online result. Missing context is repaired by richer eval records. User response is repaired by running longer and reading the trend. In all cases the online result stands as the decision — the offline suite is what gets corrected, never the other way round. ## What interviewers listen for Weak answers say "the eval set was bad" and stop. Strong answers enumerate distinct mechanisms, propose a cheap discriminating test for each, and are explicit that when the two disagree the production measurement wins and the harness is the artifact that must be fixed.

  • How would you distinguish distribution shift from metric mismatch without running another experiment?
    Re-score both variants offline on a fresh sample of live inputs. If the gain shrinks or inverts on fresh data, it is distribution shift — the frozen set no longer represents traffic. If the gain persists on fresh inputs while the online business metric still moved the wrong way, the offline metric is not proxying value, and the rubric or scoring criteria need to change.
  • The treatment arm looks worse in week one and even in week two. When is that still not a real quality loss?
    When it is disruption of habituated users rather than worse output. Existing users have learned the old behaviour and a changed interface costs them time before it pays. The tell is a trend: a disruption effect narrows week over week and concentrates in returning users, while a genuine quality loss is flat over time and present in new users too.
  • Your offline set is refreshed and the metric is validated, yet online still disagrees. What is left?
    Usually the systems are not identical. Diff the actual assembled prompt, retrieval parameters, truncation and model version between harness and serving path, then check treatment-arm latency and error rate under real concurrency. Added latency in a conversion funnel routinely outweighs quality gains, and it never appears in a sequential harness run.
  • Should a large offline gain with a small online loss ever be shipped anyway?
    Only with an explicit reason that the online read is untrustworthy — the experiment was underpowered, contaminated by a concurrent change, or measured across an unrepresentative window. Absent that, shipping means preferring a proxy over the outcome it was supposed to proxy. The better move is to extend the experiment or slice it until you know where the loss lives.

saying these in an interview costs you the question

  • Blames the online experiment and ships on the offline number anyway
  • Assumes any offline-online gap means the eval set is simply too small
  • Never checks whether the served variant matches the evaluated one
  • Ignores added latency as a cause of conversion loss
  • Treats offline metrics as if they were business outcomes

context