skip to content

Why can interleaving two search rankings detect a winner with less traffic than an A/B test?

level: seniorimportance: should knowfreq 40%

answer

  1. Both systems in one result page
  2. Each user is their own control
  3. The variance you cancel is between people
  4. Turn-taking with randomised first pick
  5. Tells you which is preferred, not by how much revenue

basics

~20 s

Interleaving blends both rankings into one result list per query, so every user compares both systems at once. That paired within-user design removes between-user variance, which is what makes an A/B test need large samples.

solid answer

~50 s

In an A/B test, one population sees ranking A and a different population sees ranking B; the comparison must overcome the enormous variance between users, whose activity levels differ by orders of magnitude. **Interleaving** merges the two rankings into a single list — in team-draft interleaving the two rankers take turns picking their next unused document, with the pick order randomised, and the system records which side contributed each position. Clicks are then attributed to the contributing ranker and the winner is whichever side collects more credit. Because each impression exposes both systems, the comparison is paired and the between-user variance cancels, so it typically resolves a preference with far less traffic. The trade-offs are real: interleaving yields a *relative preference* between two rankings, not an absolute engagement or revenue number, it only compares rankings rather than whole experiences, and it needs care with deduplication, position bias, and result-level tracking. Teams use it as a fast filter and confirm winners with an A/B test.

go deeper

for a junior

Know that interleaving mixes two rankings into one result page rather than splitting users into two groups.

for a middle

Explain the merge procedure and click attribution, and say why a within-impression paired comparison needs less traffic than a between-user split.

for a senior

Demonstrate the trade-offs: relative preference only, ranking-only scope, attribution bookkeeping, and where interleaving sits between the offline suite and the launch A/B test.

for a principal

Own the three-layer evaluation programme and the policy of which decisions each layer is allowed to make, including the periodic holdout that checks accumulated interleaving wins against real business metrics.

## The problem with between-user tests A standard online controlled experiment splits users into two buckets and compares aggregate metrics. The statistical difficulty is that users are wildly heterogeneous: a few power users issue hundreds of queries a week while most issue a handful, and engagement metrics are heavy-tailed. The variance you must overcome is dominated by *which users landed in which bucket*, not by the ranking difference you care about. That is why detecting a small relevance improvement can require weeks of traffic. ## What interleaving does instead Interleaving turns the between-user comparison into a within-impression one. For a single query, both rankers produce their list. The two lists are merged into one list of the usual length, keeping track of which ranker contributed each document. The user sees one ordinary result page. Clicks are credited to the contributing ranker, and across many impressions the ranker collecting more credit wins. Every user now contributes evidence about both systems, so user-level differences cancel out in the paired comparison. What remains is the signal you wanted. ## Team-draft interleaving The most common construction borrows the playground team-picking procedure. Repeatedly: randomly decide which ranker picks first this round, then each ranker in turn appends its highest-ranked document not already in the merged list. Randomising who picks first each round prevents a systematic positional advantage for either side, and skipping already-chosen documents handles overlap — documents both rankers rank highly appear once and, since neither side gets exclusive credit for them, they contribute nothing to the comparison. This is a feature: the comparison focuses on where the two systems actually disagree. An earlier scheme, balanced interleaving, merges by walking both lists and appending the next unseen document from each; it is simpler but has a known bias in certain overlap configurations, which is why team-draft became the default. ## What interleaving cannot do **It gives a preference, not a level.** Interleaving says B is preferred to A. It cannot tell you what revenue per session, sessions per user, or overall click-through rate will be after launch, because no user ever experienced B alone. Business metrics still need an A/B test. **It compares rankings only.** If the change also alters snippets, the UI, latency, or the number of results, the merged list does not represent either experience faithfully, and the technique is a poor fit. **It needs faithful attribution.** You must log which ranker contributed each shown document, deduplicate correctly, and handle results that both sides return. Bugs in that bookkeeping produce a confidently wrong verdict. **It still lives with click bias.** Randomised pick order neutralises the systematic positional advantage between the two sides, but clicks remain a noisy proxy; a change that wins clicks by promoting clickbait wins interleaving too. **Neither ranking is shown as designed.** The user sees a blend, so any effect that depends on the coherence of a whole result page — diversity, deduplication across the page, a curated top result — is distorted by the merge. ## How it fits the evaluation programme A mature search team runs three layers: 1. **Offline metrics** on a judgment set: cheap, reproducible, runs on every commit, filters out obviously bad candidates and guards against regressions. 2. **Interleaving** on live traffic: fast, sensitive, resolves "is B's ranking better than A's?" in days rather than weeks, used to triage a queue of candidate changes. 3. **A/B tests** on the survivors: slower and less sensitive per unit of traffic, but the only source of absolute business metrics, long-horizon effects such as retention, and guardrail measurements like latency and error rate. Each layer answers a question the layer above cannot. Skipping interleaving is common and survivable; skipping the A/B test before a significant launch is how teams ship a ranking that wins clicks and loses money. ## Practical cautions Run interleaving long enough to cover weekday and weekend behaviour. Watch for novelty effects in the first days. Guard against a ranker that games attribution — for instance one that returns near-duplicates of the other side's documents. And keep a holdout: even a well-run interleaving programme benefits from a periodic long-run A/B check that the accumulated wins actually moved the metrics that matter. ## What an interviewer wants The within-impression paired design and why it cancels user variance, at least a sketch of team-draft construction, the honest limitation that only a relative preference comes out, and the layered programme that ends in an A/B test for launch decisions.

  • In team-draft interleaving, why is the choice of which ranker picks first randomised each round?
    Picking first means placing a document higher, and higher positions collect more clicks regardless of quality. If one ranker always picked first it would accumulate a systematic positional advantage and win by construction. Randomising the order each round makes that advantage symmetric in expectation, so the click difference reflects the documents rather than the merge procedure.
  • What happens to documents that both rankings return in their top results?
    They appear once in the merged list and, because neither side can claim exclusive credit for them, they contribute nothing to the comparison. The verdict therefore rests entirely on where the two rankers disagree. That concentrates the statistical power usefully, but it also means interleaving is uninformative when the two rankings are nearly identical.
  • Why would you still run an A/B test after interleaving picks a winner?
    Interleaving yields only a relative preference between two rankings on a blended page. It cannot measure absolute revenue, sessions per user, retention, or guardrails such as latency and error rate, and no user ever experienced the new ranking as designed. An A/B test exposes the real end-to-end experience and produces the numbers a launch decision needs.

An A/B test is two restaurants judged by two different sets of diners; interleaving puts both kitchens' dishes on one plate and watches which the same diner eats first.

saying these in an interview costs you the question

  • Claiming interleaving gives absolute engagement or revenue numbers
  • Thinking interleaving means showing ranking A to half the users
  • Concatenating one list after the other and calling it interleaving
  • Assuming interleaving eliminates all click bias
  • Skipping the A/B test because interleaving already picked a winner

context