skip to content

Offline recall rose 12% but launch moved enrolments 0% - how should offline results gate launches?

level: principalimportance: should knowfreq 42%

answer

  1. audit the protocol before the theory
  2. proxy metric versus causal effect
  3. one held-out label is not all right answers
  4. zero measured is an interval, not a null
  5. record offline and online deltas over time

basics

~20 s

Treat an offline top-k win as a screen, not a decision. It scores re-ranking of behaviour that already happened, not behaviour change. Calibrate the bar against your own recorded history of offline versus online deltas rather than a threshold someone invented.

solid answer

~50 s

The two numbers answer different questions. Offline recall asks whether the model would have placed the one course a learner actually enrolled in near the top of a list. The launch asks whether showing a different list causes an enrolment that would not otherwise have happened. A model can win the first and change nothing in the second: the held-out course is only one acceptable answer, it is frequently something the learner would have found anyway, and the shelf may account for a small share of how anyone reaches a course. So offline gates entry to the experiment queue, not the launch. Before accepting the divergence as real I re-check the protocol itself, because a leaky split or a sampled candidate set manufactures gains of exactly this size. Then I keep a ledger of every shipped pair of offline and online deltas and set the offline bar from that history.

go deeper

for a junior

Be ready to say that an offline improvement is evidence about a proxy, and that the launch measures a different thing, so the two can disagree without either being a bug.

for a middle

Explain the concrete mechanisms: the held-out item is one acceptable answer among many, the shelf drives only part of the outcome, and a leaky split or sampled candidate set can produce a gain of this size on its own.

for a senior

Show the audit you would run first, then the analysis of the online result as an interval rather than a zero, and name which protocol decision you would change before the next attempt.

for a principal

Own the gating policy itself: align the offline metric with the launch metric, keep a ledger of offline and online deltas, set the bar from that history, and defend keeping a screen at all when experiment capacity is scarce.

## First: is the divergence real, or did the protocol invent it? Before theorising about user behaviour, audit the measurement. A 12 percent relative offline gain is comfortably inside the range that protocol defects produce on their own: - a split whose held-out events sit at scattered dates, so training carried later activity and the model was rewarded for knowing the future; - a small sampled candidate set, which compresses the tail toward the top and can favour whichever model happens to crowd the head; - an offline exclusion rule that differs from serving, for instance scoring courses the live shelf would never display; - a cutoff larger than the number of slots on the page, so retrieval headroom improved while the visible list did not. Re-score both models under the strictest protocol you can run. If the gap collapses, the story is over and the lesson is about evaluation, not about learners. ## Second: why a genuine offline gain can still move nothing **One label is not the set of right answers.** The held-out enrolment is a single acceptable outcome among many. A model that ranks a *better* course first is scored as wrong, and a model that ranks the *recorded* course first may have earned nothing, because that course is often one the learner was going to find through search or a direct link regardless. **Some targets are unwinnable.** If the held-out course is one the platform had already shown the learner and they declined, then predicted it, no ranking policy can convert it. Offline the item scores as a hit; online it is a course the learner has already rejected. **Attribution and dilution.** If the recommendation shelf drives a modest share of enrolments, even a genuine improvement in shelf quality is a small fraction of the top-line number, and the launch may only be able to bound the effect rather than detect it. Zero measured is not zero; it is an interval whose width you should quote. **Substitution rather than creation.** Better ranking can bring forward an enrolment the learner would have made anyway, moving the timing without changing the total. ## What to change about the gating process 1. **Align the offline metric with the online decision metric.** If the launch is judged on enrolments, the offline target event is an enrolment, at the cutoff the surface shows, on the learners the surface reaches. Every step of distance between the two is a place for the correlation to leak out. 2. **Make offline a screen with a known false-positive rate.** Its job is to decide who gets scarce online traffic. That means a bar high enough to be worth an experiment slot, not a bar that pretends to settle the question. 3. **Keep a launch ledger.** Record, for every shipped change, the offline delta and the online delta with its interval. After a dozen entries you have the empirical relationship for your product, and you can set the offline threshold from data instead of intuition. You may discover it is weak - which is itself the most valuable thing this exercise produces. 4. **Report offline results with uncertainty.** A point estimate over a few thousand held-out learners invites over-reading; a spread across learners makes a 12 percent relative gain look appropriately less certain. 5. **Version the protocol.** One written definition of split rule, candidate set, cutoff, exclusions and target. Numbers from different versions are never compared. Without this, the baseline drifts and the ledger is worthless. 6. **Allow mechanistic exceptions in both directions.** A change with a flat offline number may still be worth traffic when there is a clear causal argument - fresher items, better coverage of the new catalog, lower latency - and a spectacular offline number with no such story deserves more scepticism, not less. ## The over-correction to avoid The wrong conclusion is that offline evaluation is worthless and everything should go to traffic. Online capacity is finite and expensive; if nothing is filtered, the experiment queue becomes the bottleneck and the throughput of the whole team drops. The goal is a calibrated screen whose failure rate you know, not the abolition of screening. ## What this question is really testing Whether the candidate defends the offline number, blames the experiment, or does the mature thing: audit the protocol first, accept that offline and online measure different quantities second, and then change the *process* so the organisation learns the strength of its own proxy instead of relitigating each launch by argument.

  • What would convince you the offline metric, rather than the model, produced the 12 percent?
    Re-score both models under a fixed-date split with full-catalog candidates and the serving exclusion rule. If the margin shrinks or flips, the protocol produced it. A second check is the composition of the wins: if the newly captured targets are courses already shown and declined earlier, the model is fitting the logging record rather than learner intent.
  • Is a flat online result the same as no effect?
    No. It bounds the effect at roughly what your traffic and duration could resolve, and for a shelf that contributes a modest share of enrolments that band can be wide. Report the interval, say whether the effect size you cared about still sits inside it, and state what additional exposure or duration would be needed to exclude it.
  • How do you keep the offline gate honest over several years?
    Version the protocol like code, with one document naming split rule, candidate set, cutoff, exclusions and target event, and forbid comparisons across versions. Re-derive the offline threshold from the launch ledger every few quarters, because the relationship drifts as the surface, the catalog and the logging policy change.

saying these in an interview costs you the question

  • Concludes offline evaluation is useless and skips it
  • Blames the experiment and ships on the offline number
  • Assumes a 12 percent offline gain implies a 12 percent online gain
  • Never audits whether the offline protocol itself was leaky
  • Treats a flat online result as a proven zero effect
  • Treats the single held-out item as the only correct answer

context