skip to content

questions

6

In a sponsored-listing strip, how does team-draft interleaving compare two ranking policies within a single user's list?

level: middleimportance: must knowfreq 58%

answer

  1. one list, two authors
  2. coin flip chooses who drafts first
  3. teams alternate, skipping duplicates
  4. hidden team tag on every slot
  5. click credits the drafting team

basics

~20 s

Team-draft interleaving fills the strip by alternating drafts: a coin flip picks who starts, then each policy takes its highest-ranked listing not already placed. Every slot carries a hidden team tag, and a click credits the team that drafted that slot.

solid answer

~40 s

Both policies rank the same candidate pool. Per request a fair coin decides which team drafts first, then the two alternate: on its turn a team places its own highest-ranked listing that is not already in the strip, so duplicates are placed once and attributed to whoever took them first. The merged strip renders in the normal layout and the user sees an ordinary strip. Each filled slot keeps a hidden team tag in the impression log, and a click scores a point for the team that contributed that listing. Per user you total the credited clicks and record a win, a loss or a tie, then count wins across users. The comparison is within-user and paired: the same person, the same page, the same moment, two orderings.

code

pseudocode · 20 lines
pseudocode
function interleave(rankingA, rankingB, slots):
    merged = []
    teamOf = {}
    aDraftsFirst = coin_flip()

    while length(merged) < slots:
        aTurn = (length(merged) is even) == aDraftsFirst
        source = rankingA if aTurn else rankingB
        team   = "A" if aTurn else "B"

        pick = first listing in source that is not in merged
        if pick is none:
            break
        append pick to merged
        teamOf[pick] = team

    return merged, teamOf

function credit(clickedListing, teamOf):
    return teamOf[clickedListing]

go deeper

for a junior

Hold on to the core image: one strip, built by two ranking policies taking turns, shown to one person. Clicks are credited to whichever policy contributed the listing that was clicked.

for a middle

Be able to run the protocol aloud: same candidate pool, coin flip for first pick, alternate, skip listings already placed, tag every slot, credit the click to the tag, aggregate to one outcome per user.

for a senior

Show that you know attribution is a logging contract. Say where the team tag is written, how it reaches the click event, and what a client-side reorder or a lazy-loaded slot does to credit.

for a principal

Frame it as choosing an instrument. Interleaving buys a paired within-user comparison on a click-shaped question and pays for it in scope - it can only judge orderings of one pool inside one unchanged layout.

## Two rankings, one strip A marketplace renders a short strip of sponsored listings above the organic results - say three slots. An incumbent ranking policy fills them today; a candidate policy would fill them differently. A between-user traffic split answers this by showing some users the incumbent's strip and other users the candidate's strip and comparing the two groups of people. **Team-draft interleaving answers it differently: it builds one strip out of both policies and shows it to everyone in the test, so the comparison happens inside a single user's list.** ## The draft, step by step 1. Ask both policies for a ranked list over **the same candidate pool** - the same retrieved listings, eligible for the same slots. 2. Flip a fair coin per request to decide which team drafts the first slot. 3. Alternate. On its turn a team takes its own highest-ranked listing that is **not already in the merged strip**. 4. Write a hidden team tag beside each filled slot in the impression log. 5. Render the merged strip in the unchanged layout. Nothing on screen marks which team produced which slot. The coin flip is not decoration. With a fixed draft order one team would own slot 1 on every request, and because the top slot collects most of the attention, that team would win whatever it ranked. Randomising the first pick balances each team's share of the top slot in expectation across requests, so the position advantage does not sit with one policy. ## Crediting a click - A click on a slot credits **the team that drafted that listing**, not the team that drafted first and not both teams. - A listing both policies rank highly is placed once, by whoever reaches it first; it is attributed to exactly one team. - Impressions with no click contribute nothing to either team - the verdict is built only from credited clicks. - The per-user total is compared, and the user contributes one outcome: a win for the team with more credited clicks, or a tie. - Wins are counted across users over the window, and the policy with the larger share of wins is preferred. ## Why the unit of analysis matters | comparison | what differs between the things compared | who sees the candidate | |---|---|---| | between-user traffic split | the entire strip a person is shown | only the assigned fraction | | within-user interleaving | the ordering inside one shared strip | every user in the test | Because each user is exposed to both policies at once, the user's own appetite for sponsored listings - heavy clicker, never clicks - is held constant on both sides of the comparison instead of having to average out across two groups of people. That is the paired structure interleaving buys, and it is the whole reason the technique exists. ## What the tags have to survive The protocol lives or dies on attribution, which is a logging problem more than a ranking problem: - The team tag must be written at **impression** time, per slot, and must travel to the click log through whatever id the click carries. - A re-render, a lazy-loaded slot or a client-side reorder must not renumber slots, or credit lands on the wrong team. - If the two policies return heavily overlapping lists, most slots are agreed and only the differing ones carry information; near-identical policies produce few decisive impressions. - Deduplication has to be by the listing's identity, not by its position, or the same listing can be drafted twice. ## The shape of the readout The output is a **preference**: users clicked the candidate's contributions more often than the incumbent's, inside the same strip. It is expressed as a share of user-level wins, not as a percentage change in revenue or in sessions. That is a deliberate narrowing - the instrument is asking one question, which of these two orderings people reach for, and it answers that question with the layout, the slot count and the candidate pool all held fixed. Every other question about the change - what it does to spend, to abandonment, to how the strip is perceived after a month - is asked with a different instrument. A variant called balanced interleaving merges by walking both lists with a shared cursor rather than by drafting, and it is known to favour one list in some configurations; the team-draft form is the one usually taught because its attribution rule is simple to state and simple to log.

  • What happens when both ranking policies put the same listing first?
    Whoever drafts first takes it, and it is tagged to that team alone; the other team's turn moves to its next unplaced listing. The listing is never shown twice. Heavy agreement between the two policies is normal and it shrinks the number of slots that actually differ, so a run between near-identical policies carries little information per impression.
  • A user sees the interleaved strip on twelve pages in one week. How do you turn that into one verdict?
    Credit each click to its drafting team as it happens, sum the credited clicks per user over the whole window, then emit one outcome for that user - win, loss or tie. Counting raw clicks instead would let a single heavy user dominate the run, because the thing being compared is people's preferences, not sessions.
  • Does the user need to be assigned to the interleaving test consistently?
    Yes for the aggregation to mean anything: the same user must keep seeing interleaved strips for the window, or their credited clicks are split across a treated and an untreated life. The draft order still flips per request - that randomisation is inside the test, not the assignment to it.

saying these in an interview costs you the question

  • Thinks each user sees one policy's full strip, not a merged one
  • Credits every click in the strip to the candidate policy
  • Believes the same team always drafts the first slot
  • Expects interleaving to report a percentage lift in revenue
  • Places a listing twice when both policies rank it highly
  • Sums raw clicks across users instead of user-level outcomes
open as a page

Your permanent sponsored-ranking holdout shows +3.1% while the year's shipped launches summed to +5.6% - why the gap?

level: seniorimportance: must knowfreq 52%

basics

~20 s

Short launch readouts are not additive. Each was measured while the change was still novel, only the launches that read positive were kept, and launches touching the same slots cannot both claim their full effect. The holdout measures what actually survived together.

open as a page

Why is interleaving invalid for a sponsored-strip change that adds a fourth slot and widens the candidate pool?

level: seniorimportance: must knowfreq 46%

basics

~20 s

Interleaving compares two orderings of one shared pool inside one unchanged container. A fourth slot means no single merged strip represents both policies, and a widened pool lets the candidate win on listings the incumbent never had rather than on better ordering.

open as a page

Interleaving on the sponsored strip picks the candidate ranker in two days - why doesn't that settle the launch?

level: seniorimportance: should knowfreq 42%

basics

~20 s

An interleaving run returns a preference between two orderings of one list, measured in credited clicks. It carries no session-level effect size, no guardrail outside the strip, and no effect that needs weeks to appear, so the candidate still walks the ramp.

open as a page

Your permanent holdout keeps 1% of users out of every sponsored-ranking launch - would you keep, shrink or retire it?

level: principalimportance: should knowfreq 31%

basics

~20 s

There is no default answer: the holdout is the one instrument that prices a year of launches together, and it costs a permanently worse experience for 1% of users plus a frozen serving path maintained forever. Decide by naming the decision it feeds.

open as a page

Why is a permanent sponsored-ranking holdout not a clean counterfactual once the marketplace adapts around it?

level: seniorimportance: nice to knowfreq 28%

basics

~20 s

Only the serving policy is held out. Sellers, listings, bids and the retrieval index are shared with the launched population, so part of every launch reaches held-out users - and if the holdout's policy is frozen instead, its own staleness is counted as launch effect.

open as a page