skip to content

Interleaving on the sponsored strip picks the candidate ranker in two days - why doesn't that settle the launch?

level: seniorimportance: should knowfreq 42%

answer

  1. ordinal preference, not an effect size
  2. one list, so nothing off-list
  3. the window is one impression
  4. nobody lived in the candidate's world
  5. screen before the traffic split

basics

~20 s

An interleaving run returns a preference between two orderings of one list, measured in credited clicks. It carries no session-level effect size, no guardrail outside the strip, and no effect that needs weeks to appear, so the candidate still walks the ramp.

solid answer

~50 s

Interleaving answers one question well: inside this strip, right now, which ordering do people reach for. It is a paired verdict on a click-shaped proxy, reported as a share of user-level wins. What it does not contain is an effect size on the launch metric - spend per session, marketplace outcomes - and it has no view of anything outside the merged list, so abandonment, organic engagement and seller-side exposure are invisible to it. It also sees only immediate behaviour: novelty, habit formation and anything that emerges once the ranking owns all traffic are outside its window by construction. So a win promotes the candidate from offline head-to-head to a real traffic split, where the launch metric and the guardrails are readable, with a long-term holdout behind the whole programme. Interleaving is a screen, not a verdict.

go deeper

for a junior

Interleaving tells you which of two orderings people clicked more inside one strip. It does not tell you how much money, time or sessions changed, because nobody in the test saw the candidate's world on its own.

for a middle

Be able to say what the output is - a share of user-level wins on credited clicks - and name two things it cannot contain, such as a session-level effect size and any guardrail outside the list.

for a senior

Position it in the chain: offline head-to-head, interleaving as a fast screen, traffic split for the launch metric and guardrails, long-term holdout for the cumulative view. Explain what each stage adds.

for a principal

The judgment call is where to spend evidence. A cheap paired screen ahead of an expensive split changes how many candidates you can afford to try, which changes how ambitious the team can be.

## What the verdict actually says A team-draft interleaving run over a sponsored strip returns something like: **users preferred the candidate's contributions in 54% of decisive impressions**. Read that sentence literally, because every word of it is a limit. - **Preferred** - measured by a credited click, a proxy for value, not value itself. A click that ends in an immediate bounce counts the same as one that ends in a purchase, unless you define the credited event more strictly. - **The candidate's contributions** - the comparison is between listings inside one strip, not between two versions of the page. - **Decisive impressions** - impressions where the two policies actually differed. Agreed slots carry no information. - The result is a **share of wins**, which cannot be converted into a percentage change in any business quantity, because nothing in the design ever showed a user the candidate's world in full. ## Four things the run does not contain 1. **An effect size on the launch metric.** Nobody in the test lived in the candidate's world, so there is no candidate-world spend per session to compare against an incumbent-world one. The verdict is ordinal - this ordering beats that one - not a magnitude the business can plan against. 2. **Guardrails outside the strip.** Session length, page abandonment, engagement with organic results, complaint rate, latency of the ranking call: all of these belong to the page and the session, and both arms shared the page. 3. **Anything that takes time.** A user reacting to an unfamiliar ordering in its first hour is not the same user four weeks later. Interleaving's window is the impression. 4. **Anything that only appears at full traffic.** When the ranking owns all traffic, its own clicks become tomorrow's training data and its exposure decisions reshape which sellers get seen. In an interleaved run both policies are simultaneously shaping the same small stream, so those loops are muffled. ## The sequence a candidate ranker actually walks | stage | question it answers | what it cannot answer | |---|---|---| | offline head-to-head on logged data | is the candidate plausibly better on recorded outcomes | anything about live behaviour | | interleaving on the strip | do people reach for this ordering over that one | how much anything moves, and anything off-list | | traffic split at a share of users | what the launch metric and the guardrails do in the candidate's world | what a year of such launches sums to | | long-term holdout | what the cumulative change looks like after novelty decays | which individual launch caused what | Each instrument is cheaper and blinder than the next. Interleaving's place in that chain is as a **fast screen between the offline winner and the traffic split**: it can kill a candidate that offline evaluation liked, before you spend the traffic and the calendar time a split costs, and it can rank several candidate orderings against each other quickly. ## Reading a preference win honestly The failure to avoid is treating the ordinal verdict as a licence. Two phrasings of the same result: - *The candidate wins interleaving, so it is worth the cost of a real traffic split.* This is what the run supports. - *The candidate wins interleaving, so ship it.* This is not, and the gap between the two sentences is where most of the value of a rollout process lives. There is also a direction worth stating plainly: a candidate can win a preference comparison and still lose on the launch metric. Suppose the candidate promotes listings that draw more clicks but convert less, or crowds one high-click seller into the strip so repeat exposure rises. Clicks move up, credited wins move up, the launch metric does not follow. That divergence is exactly why the launch metric is read in a design where somebody actually lived in the candidate's world. ## What the win does buy It is a real result and worth having. A preference verdict arrives in days rather than weeks, it needs no separate treatment population, and it is paired - the same person judged both orderings, so the differences between heavy and light clickers do not have to average out across groups. Used as a screen it removes weak candidates early and lets the expensive instruments be spent on the survivors. That is the honest framing in a design round: **name the instrument, name its question, and name the instrument that answers the next one.**

  • Can you make interleaving report a revenue difference by crediting spend instead of clicks?
    You can credit any per-listing outcome, including order value, and it sharpens the preference signal. What it still cannot give you is a session-level or page-level effect, because both teams shared one page: spend that moved from the strip to an organic result, or a session the user abandoned, is not attributable to a slot.
  • Interleaving prefers the candidate but the traffic split reads flat on the launch metric. Which do you believe?
    Both, and the divergence is information. Interleaving says people reach for the candidate's listings; the split says the session outcome did not move. Common explanations are a click-versus-conversion gap, or a gain inside the strip that came out of organic results. Investigate the mechanism before overriding either reading.
  • Why run interleaving at all if the change is going to a traffic split anyway?
    Cost and order. A preference verdict comes back in days and can eliminate a candidate the offline head-to-head liked, so the slower and more expensive instrument is spent only on survivors. It also ranks several candidate orderings against one another quickly, which a split does badly.

saying these in an interview costs you the question

  • Reads a share of interleaving wins as a revenue lift
  • Believes a preference win removes the need for a traffic split
  • Thinks interleaving can catch page abandonment or session length
  • Assumes the preferred ordering must move the launch metric
  • Treats a two-day run as evidence about behaviour after a month