Your permanent holdout keeps 1% of users out of every sponsored-ranking launch - would you keep, shrink or retire it?
answer
- name the decision it feeds
- someone lives on the old experience
- a frozen path must still be patched
- rotation costs the cumulative reading
- reverse tests attribute, holdouts total
basics
~20 sThere is no default answer: the holdout is the one instrument that prices a year of launches together, and it costs a permanently worse experience for 1% of users plus a frozen serving path maintained forever. Decide by naming the decision it feeds.
solid answer
~50 sStart from what consumes the number. If annual planning applies a discount factor derived from the holdout, and slow-moving harms such as fatigue or seller concentration are real risks in this marketplace, the arm is earning its keep and should stay. The costs are concrete: 1% of users live on an ageing experience indefinitely, some sellers get systematically less exposure in that slice, and a pinned policy, feature pipeline and serving config must be kept running and patched by a team that is not developing them. Between keep and retire sit two middles - shrink the share, or replace the permanent arm with periodic reverse tests that turn one shipped change back off for a few weeks. Shrinking keeps the multi-year reading; reverse tests give per-launch answers but never the programme total. State the tradeoff, then pick.
go deeper
Keeping a permanent holdout means some real users never receive any improvement. That is the price of being able to say what a year of changes was actually worth.
Be able to state both sides concretely: the arm gives a cumulative reading after novelty decays, and it costs an ageing experience for a fixed group plus a frozen serving path that still needs patching.
Show the operational reality - who owns the frozen path, what may ship into it, how exceptions are recorded - and know that reverse tests attribute a single launch while only the holdout totals the programme.
Commit: name the consumer and the cadence, set the share from the cost you accept, choose permanent or rotating knowing rotation destroys the multi-year reading, and say what would make you retire the arm.
## What the holdout buys Only two instruments can see an effect after novelty has decayed: a permanent holdout and a reverse test. The holdout is the only one that prices **the whole programme at once**, and that single number does work nothing else does: - It supplies a **discount factor** for short-run launch readouts, which is what makes next year's plan honest. - It catches **slow-moving harms** - fatigue with a denser strip, concentration of exposure among a few sellers, an erosion of trust - that no two-week window can show. - It is an **organisational commitment device**: a team cannot accumulate a year of unverifiable wins when a standing arm reconciles them. ## What it costs | cost | who pays it | how it grows | |---|---|---| | a permanently older experience | the 1% of users in the arm | compounds every year the programme ships | | less exposure for some sellers | the supply side, in that slice | grows as the ranking diverges from the launched one | | a pinned policy, feature pipeline and serving config | the platform team | every upstream change is a chance the frozen path breaks | | an increasingly hard reading | whoever quotes the number | contamination and staleness both accumulate | The first and third are the ones that decide most arguments. The user cost is genuine and it is not shared evenly - the same people bear it for years. The maintenance cost is the one that quietly kills holdouts: a path nobody develops and everybody must patch eventually breaks in a way that is discovered when the number looks strange. ## The decision, stated as a decision A good answer commits to four things rather than listing considerations: 1. **Name the consumer.** Which decision, on which calendar, uses this number? If nobody can name one, the arm is a habit and should be retired. 2. **Name the lifetime and the refresh rule.** A permanent arm and a rotating one are different instruments: rotation returns users to the current experience and spreads the cost, but it destroys the multi-year cumulative reading, because after the first rotation nobody in the arm has missed every launch. 3. **Size it by cost, not by appetite.** The share is a statement about how many users you are willing to keep on an older experience, and it should be set at the smallest value that still supports the reading you committed to in point one. 4. **Write down what may ship into it.** Safety, legal and correctness fixes must reach held-out users; each such exception is a recorded dent in the counterfactual. ## The middles worth proposing - **Shrink.** If the arm exists for an annual discount factor and a fatigue watch, a smaller share may still serve, and the user cost falls proportionally. This preserves the cumulative reading, which rotation does not. - **Reverse tests instead.** Turn one shipped change back off for a slice for a few weeks. This gives a decayed reading for that specific change - which the holdout can never attribute - at the cost of never producing a programme total, and of a re-exposure effect when a familiar experience is withdrawn. - **Both, at different cadences.** A small permanent arm for the annual number, plus reverse tests when a particular launch's claim needs interrogating. This is usually the strongest answer and it is explicit about which instrument answers which question. ## Where the judgment actually lies The question is not statistical, it is organisational. It asks how much a business is willing to pay, in permanently degraded experience for a known set of users and in carrying a frozen path, to be able to answer **did a year of work actually help**. A team that ships a few cautious ranking changes a year and reads each one at a longer horizon may honestly not need a standing arm. A team shipping continuously into a surface where the effects compound and the harms are slow almost certainly does. The answer that fails a design round is the one that treats the holdout as free discipline, because someone is paying for it every day and, if they are never named, the arm eventually gets deleted by whoever is asked to fix it.
- What does rotating the holdout population each quarter cost you?The multi-year reading. Once users rotate in and out, nobody in the arm has missed every launch, so the arm can only price the launches since the last rotation. What rotation buys is fairness - the user cost is spread rather than borne by the same people for years - and a fresher, less contaminated comparison.
- A reverse test on one launch reads much smaller than that launch's original readout. What do you conclude?That the original two-week number contained novelty, overlap with later changes, or both. A reverse test is read after the change has become normal, so it is the closer estimate of standing value. Note the mirror effect: withdrawing a familiar experience can itself provoke a reaction, so read it over a few weeks, not days.
- How do you stop a frozen holdout path from rotting?Treat it as a supported production path: it runs the same deployment and monitoring as everything else, it is covered by the same patching, and one owner is named for it. Record every change that lands on it, because each one is a dent in what the arm is supposed to represent.
- The holdout has not been read in a year. Keep it?No. An arm nobody consumes is pure cost - a worse experience for real users and a frozen path someone maintains - with no decision attached. Either attach it to a named decision on a stated cadence, or retire it and use reverse tests when a specific claim needs checking.
A retail chain keeps one store on the old layout for years so it can tell whether a decade of remodels helped. The knowledge is real, and that store's staff and customers pay for it the whole time.
saying these in an interview costs you the question
- Keeps the holdout because it is best practice, naming no consumer
- Thinks rotating the holdout keeps the multi-year reading intact
- Treats the permanently degraded 1% as a rounding error
- Forgets that the frozen serving path costs engineering effort
- Believes reverse tests can produce a programme-level total
- Sizes the arm without saying what the number will decide