skip to content

Why is a permanent sponsored-ranking holdout not a clean counterfactual once the marketplace adapts around it?

level: seniorimportance: nice to knowfreq 28%

answer

  1. only the ranking decision is held out
  2. sellers and inventory are shared
  3. adapted supply narrows the gap
  4. a frozen arm also goes stale
  5. write down what is actually frozen

basics

~20 s

Only the serving policy is held out. Sellers, listings, bids and the retrieval index are shared with the launched population, so part of every launch reaches held-out users - and if the holdout's policy is frozen instead, its own staleness is counted as launch effect.

solid answer

~50 s

A holdout promises a world without this year's launches, but it can only hold out what the ranking policy decides. Everything upstream is common: sellers reprice and rewrite listings in response to what the other 99% see, the retrieval index is shared, and page-level changes ship to everyone - all of which narrows the measured gap, so the holdout understates. The second problem depends on the design. If the holdout keeps retraining on pooled logs dominated by launched traffic, its baseline improves for free and the gap narrows further. If instead you freeze the holdout's policy version, that arm also stops absorbing routine retraining, so ordinary staleness is added to the launch effect and the gap widens. Neither design is clean; the honest move is to record which components are actually frozen and which bias you have accepted.

code

json · 21 lines
json
{
  "arm": "long-term-holdout",
  "share_of_users": 0.01,
  "design": "freeze-the-serving-policy",
  "frozen": {
    "ranking_policy_version": "sponsored-rank-v41",
    "feature_pipeline_version": "fp-v17",
    "serving_config_version": "cfg-v09"
  },
  "not_held_out": [
    "seller listings, prices and bid strategies",
    "the contents of the retrieval index",
    "page layout and eligibility rules shipped to everyone",
    "anything a frozen policy cannot re-decide"
  ],
  "confounded_with": [
    "staleness: the frozen policy also stops absorbing routine refreshes",
    "patch drift: the frozen path is maintained but never improved"
  ],
  "reads_as": "launch effects reachable by the ranking policy, plus staleness"
}

go deeper

for a junior

The held-out users get an older ranking, but they shop in the same marketplace, over the same listings, from sellers who have already adapted to the new ranking. That shared world is what makes the comparison imperfect.

for a middle

Be able to list what is genuinely held out - the ranking policy and its immediate configuration - against what is shared: inventory, seller behaviour, the retrieval index and page-level changes.

for a senior

Name both designs and their opposite biases, state which one you would build, and produce the artefact that records what is frozen so the reading can be interpreted months later by someone else.

for a principal

Decide what the organisation is willing to spend on a measurement it knows is biased, and set the policy for what may ship into the arm - every exception trades a little measurement validity for user safety or correctness.

## What a holdout can and cannot hold out The intended reading of a permanent holdout is a counterfactual: **what would this marketplace look like if none of this year's ranking launches had happened?** The mechanism that implements it is narrow - a small share of users are routed to a different ranking decision. Everything that is not a ranking decision is shared with the rest of the world, and in a marketplace that is a great deal. ## Leak one: the supply side is common Sellers respond to whatever ranking the 99% experience: - They rewrite titles and attributes toward what wins slots under the new policy. - They shift bids and budgets toward the categories the new policy favours. - Poorly performing listings are withdrawn; new ones are created in their place. - The retrieval index that both arms draw from therefore contains post-launch inventory. When a held-out user is served by the old policy over adapted inventory, some of the launch's benefit arrives anyway. **Direction: the measured gap narrows, so the holdout understates the programme.** Page-level and policy changes that ship to everyone - layout, eligibility rules, disclosure requirements - do the same thing. ## Leak two, or its opposite, depending on the design There are two ways to build the arm and they fail differently. | design | what the holdout runs | its bias | |---|---|---| | hold out the launches, keep retraining | last year's policy shape, refreshed on pooled logs | the pooled corpus is dominated by launched traffic, so the baseline improves for free and the gap narrows | | freeze the policy version outright | a pinned policy, feature pipeline and serving config | the arm stops absorbing routine refreshes, so staleness is counted as launch effect and the gap widens | Both are defensible and neither is neutral. The frozen design is the more common choice because it is easier to reason about, but you have to say out loud that its reading is **launch effects plus staleness**, and that the staleness component grows every month the freeze lasts. ## The operational cost of the frozen path A frozen arm is not free to keep alive: 1. The pinned policy version, its feature pipeline version and its serving config must keep running while everything around them moves. 2. Security and infrastructure patches still have to land on that path, and every patch is a small chance of changing its behaviour. 3. Upstream schema changes will eventually break a pipeline nobody is developing, and the arm fails quietly rather than loudly. 4. Somebody must own the rule about what may be shipped into the holdout - safety, legal and correctness fixes normally must be, and each exception is a documented dent in the counterfactual. A holdout that has silently drifted or been patched into something nobody can describe is worse than no holdout, because its number still gets quoted. ## Reading it honestly The two biases point in opposite directions, so you cannot simply say the holdout is conservative or optimistic without saying which design you built. The practical discipline is: - Record, as a durable artefact, exactly what is frozen and what is not. - Report the holdout reading with that artefact attached, so a reader knows whether staleness is included. - Re-derive the staleness component occasionally - for instance by refreshing the frozen policy once and observing how much of the gap closes without any launch being involved. - Accept that the holdout measures **launch effects reachable by the ranking policy**, not all the effects of the year's work, and do not let the number be quoted as if it measured everything. ## Why this is still worth doing None of this makes the instrument useless. A contaminated counterfactual with known direction is far more informative than the alternative, which is summing short-run readouts and hoping. The point of understanding the leaks is to state the reading as a range with a direction - a floor on the programme's effect from the shared supply side, an upper adjustment for staleness if the policy is frozen - rather than as a single confident number that the next planning cycle will treat as exact.

  • Which way does each bias push the measured gap?
    Shared supply and shared page-level changes carry part of the launch into the held-out arm, which narrows the gap and understates the programme. A frozen policy that also stops absorbing routine refreshes widens the gap, because ordinary staleness is added to the launch effect. Say which design you built before quoting the number.
  • How would you separate staleness from launch effect in a frozen arm?
    Refresh the frozen policy once on current data without applying any of the shipped launch changes, and observe how much of the gap closes. What closes is staleness; what remains is closer to the launch effect. Do it rarely and record it, since each refresh resets the arm's history.
  • Should security and correctness fixes ship into the holdout?
    Yes - a holdout is a measurement device, not a licence to leave users on a broken or unsafe path. The requirement is that every exception is recorded against the arm's definition, so anyone reading the number knows which changes it no longer excludes.

saying these in an interview costs you the question

  • Treats the holdout reading as an exact counterfactual
  • Thinks freezing the policy removes every source of contamination
  • Ignores that sellers adapt listings to the launched ranking
  • Believes a frozen arm needs no maintenance because it never changes
  • Assumes both biases push the measured gap the same way
  • Withholds safety and correctness fixes to protect the measurement