skip to content

When is replaying a candidate ranker on last month's impression log a trustworthy estimate?

level: principalimportance: nice to knowfreq 28%

answer

  1. reweight by a ratio of policy probabilities
  2. a deterministic log answers nothing
  3. the interesting candidate has the least support
  4. count effective sessions, not logged rows
  5. a filter, not a verdict

basics

~20 s

Only when the log recorded the probability with which each item was shown, the candidate ranker mostly promotes items the old ranker did sometimes show, and enough logged sessions survive importance weighting to give a usable interval.

solid answer

~50 s

Off-policy replay reweights logged outcomes by the ratio of how likely the candidate ranker was to show what was shown to how likely the logging ranker was, so it estimates the candidate's reward on traffic it never served. Three conditions have to hold. The logging policy must have been stochastic with its propensities recorded — a fully deterministic ranker leaves nothing to reweight. There must be overlap: if the candidate's top listing was never surfaced, no reweighting can tell you what would have happened, and in practice the overlap is thin exactly where the candidate is most different. And enough sessions must survive the weighting, which you check with the effective sample size, not the raw log size. Report the interval, the coverage of the candidate's top items, and the clipping you applied. Treat a replay win as a gate for an online test, never a substitute for one.

go deeper

for a junior

Know the basic caution: a log only records what the old ranker actually showed, so you cannot learn from it how users would react to something they were never shown.

for a middle

Explain the reweighting itself — outcomes are scaled by the ratio of the candidate's probability of the logged action to the logging policy's — and why a deterministic logging ranker makes that ratio useless.

for a senior

Demonstrate the diagnostics you run before believing a number: effective sample size, the weight distribution, coverage of the candidate's top items, and how clipping shifts the estimate and in which direction.

for a principal

Own the position that offline evaluability is bought in advance through logging design and a standing randomised slice, and set the rule for when a replay result is allowed to gate a launch versus merely shortlist candidates.

## What replay is trying to do You have a month of impression logs generated by the production ranker, and a candidate ranker you would like to judge without shipping it. **Off-policy replay**, also called counterfactual or off-policy evaluation, asks: what average outcome would the candidate have produced on this same traffic? The machinery is importance sampling. For each logged session you know what the logging policy did and with what probability; you can compute what the candidate would have done and with what probability; you reweight the observed outcome by the ratio of the two. Sessions the candidate would have handled the way production actually did get up-weighted; sessions where it would have done something else get down-weighted toward zero. Note the difference from correcting position bias in labels: there the denominator is the examination propensity of a *slot*, here it is the probability the logging policy chose that *list*. Both are inverse propensity corrections; they answer different questions. ## Condition 1: logged propensities Importance weighting needs the denominator. A deterministic ranker assigns probability one to the list it produced and zero to every other, so the ratio is zero almost everywhere and no candidate that differs at all can be evaluated. The fix is architectural and has to be decided before the log exists: keep a small amount of deliberate randomness in what is served, and write the propensity of each impression into the log next to the outcome. Reconstructing propensities after the fact from a model checkpoint is fragile — features drift, models get rolled back, ties break differently. ## Condition 2: overlap, and why it collapses The support condition says the candidate may only put probability on actions the logging policy could have taken. A candidate listing ranker whose top result for a query is a listing the old ranker never surfaced is asking about a counterfactual the data cannot answer, and importance weighting cannot invent it — it will silently return an estimate dominated by whichever few sessions happened to overlap. This is the practical killer, and it bites hardest in exactly the situation you care about. A candidate that behaves like production is easy to evaluate and not worth shipping. A candidate that reorders aggressively is worth evaluating and nearly impossible to evaluate. On a large catalog the overlap between the candidate's top-1 and the logged top-k is often a small single-digit percentage. ## Condition 3: enough effective data After weighting, the estimate is effectively built from far fewer sessions than the log contains. The standard diagnostic is the **effective sample size**: the squared sum of the weights divided by the sum of their squares. A log of fifty million impressions can carry an effective sample size in the low thousands, and reporting a tight-looking mean off it is how teams talk themselves into bad launches. Look at the weight distribution too — if the top ten weights carry most of the mass, the number is an anecdote. ## Controls, and what they cost - **Clipping the importance weights** at some maximum bounds the variance and biases the estimate toward the logging policy's observed outcome. That direction is worth stating out loud: a clipped replay tends to under-credit a candidate that is genuinely different, which makes a clipped win more believable than a clipped loss. - **Doubly robust estimation** combines importance weighting with a learned model of the outcome, so the estimate stays consistent if either the propensity model or the outcome model is right. It reduces variance materially and is the sensible default, but it inherits the outcome model's extrapolation exactly where the log is thin. - **Restricting the estimand.** Reporting the candidate's performance on the overlapping slice, together with the size of that slice, is more honest than an extrapolated number for the whole population. If the estimate covers 8 percent of queries, say so; the remaining 92 percent is an open question, not a neutral one. ## Condition 4: the world stood still Last month's log carries last month's catalog, seasonality, promotions and user mix. A ranker evaluated on a holiday log and shipped in February is being judged on a different problem. And the log was generated by a ranker that had been shaping user behaviour for months; the candidate would, over time, shift what users see and click, which no single-shot replay can capture. ## How to use it well Treat replay as a **filter, not a verdict**. It is very good at what it is cheap at: killing candidates decisively, catching pathologies, ranking a shortlist of variants, and doing all that in hours rather than weeks. It is weak at certifying a winner, because the candidates most likely to win are the ones with the least support. The standing organisational lesson is that offline evaluability is a property you have to pay for in advance. A team that logs propensities and keeps a modest randomised slice can answer counterfactual questions for years; a team with a fully deterministic serving stack has a log that can only ever describe what already happened.

  • What do you do when the candidate's top listing was almost never logged?
    Say so rather than extrapolate. Report the estimate on the overlapping slice with its coverage, and treat the rest as unmeasured. If that region matters, buy support for it: log propensities and keep a small randomised slice that surfaces items the current ranker suppresses, so next quarter's replay can answer the question. Otherwise the honest conclusion is that only a live test settles it.
  • Why does clipping importance weights make a replay win more credible than a replay loss?
    Clipping pulls the estimate toward the outcome the logging policy actually produced, so it systematically under-credits a candidate that behaves differently. A candidate that still comes out ahead did so despite a conservative correction. A candidate that loses may simply have been penalised by the clipping, so a clipped loss is weaker evidence and deserves a look at the unclipped estimate and the weight distribution.
  • What would you require in the logging stack before promising the business offline ranker evaluation?
    Every impression stamped with the slot, the surface, the model version, and the propensity in force at serve time; a small persistent randomised slice so support does not go to zero; and retention long enough to cover a full seasonal cycle. Without those, replay is not merely noisy, it is undefined, and no amount of modelling recovers the missing denominator.

Judging a restaurant's new menu from last month's receipts. You learn a lot about dishes that were on both menus, and nothing at all about the ones nobody was ever offered.

saying these in an interview costs you the question

  • Runs a replay on logs that never recorded serving propensities
  • Reports a replay mean with no interval or coverage figure
  • Ignores that the log only contains what the old ranker chose to show
  • Trusts a headline number computed from a handful of heavy weights
  • Treats an offline replay win as a decision to ship

context