How do you decide whether to keep a permanent 10% global holdout excluded from every experiment?
answer
- start from the decision it informs
- everyone else now shares 90% of traffic
- old code paths can never be deleted
- membership needs its own independent partition
- an owner, a size, an expiry
basics
~10 sWeigh one clean never-treated reference population against three costs: every experiment now draws from 90% of traffic, every shipped feature must keep its old code path alive, and held-out users get a stale product.
solid answer
~50 sStart from the decision it is supposed to inform. A permanent global holdout - a fixed slice of users excluded from every experiment - buys you a population that received none of what you shipped, so you can check whether the sum of individually measured wins matches the top line. That is worth real money at large scale and worth nothing if no one will act on it. Against that, price the costs honestly: every experiment now samples from the remaining 90%, so each one is slightly less precise; every launch must respect the holdout, which means old code paths cannot be deleted and drift becomes permanent maintenance; and the held-out users are receiving a deliberately worse product. Then set the governance: assign membership through its own partition with a dedicated salt so it is independent of any experiment's split, give it an owner, and give it an expiry date.
go deeper
Know what the term means: a slice of users deliberately excluded from every experiment, kept on the earlier experience so it can serve as a reference group.
Explain the mechanical cost - every experiment now samples from the remaining traffic and is slightly less precise - and that holdout membership must be assigned independently of any experiment's split.
Be ready to describe operating one: how membership is assigned and audited, how launches are prevented from leaking into it, and how you would notice that it has silently stopped being held out.
Own the call. State which decision the holdout will inform, price the traffic and engineering tax against that decision, set its size, owner and expiry, and be willing to refuse it at traffic levels or product velocities that cannot carry it.
## What a global holdout is A **global holdout** is a slice of the population - often described as 5% or 10% - that is deliberately excluded from every experiment and every launch. Its members keep experiencing the product as it was at the moment the holdout was created. It is different from a control arm inside one experiment: a control arm is specific to a single test and dissolves when that test ends, while a global holdout persists across all of them. ## What it buys **A reference for cumulative impact.** Individual experiments each measure one change against a contemporaneous control. Summing those measured wins over a year and expecting the total to appear in the top-line metric almost never works, for many legitimate reasons. A never-treated population gives a direct comparison: everything shipped, versus nothing shipped, over the same period. **A check on metric drift.** Metric definitions, logging and populations all change over time. A holdout that is measured with the same pipeline as everyone else gives a within-period comparison that is insulated from year-over-year definition changes. **A place to look when the numbers disagree.** When the product organisation believes it shipped a strong year and the business metrics disagree, a holdout converts an argument into a measurement. ## What it costs **Statistical cost.** Every experiment now randomises over 90% of traffic instead of 100%. That is a modest but permanent tax on precision for every test the company runs, and it compounds with any other population you carve out. Note what this does *not* do: it does not change the false-positive rate of a correctly analysed test - it shrinks samples, nothing more. **Engineering cost, which is usually the binding one.** Serving held-out users the previous experience means the previous code path has to keep working. Multiply that by every feature shipped in the holdout's lifetime and you have a second, unmaintained version of the product that no one tests and everyone must avoid breaking. Dead-code removal becomes impossible. This cost grows without bound the longer the holdout lives, which is the strongest argument for a fixed expiry. **Business and user cost.** Held-out users are customers. If the year's changes were improvements, that group is receiving a deliberately worse product, with the revenue and satisfaction consequences that implies. In some domains this crosses into a fairness or contractual question rather than a purely analytical one. **Decaying comparability.** As the product diverges, the holdout stops being a control and becomes a different product. The measured difference reflects the entire divergence, including infrastructure and content changes nobody intended to test. The longer it runs, the less the comparison isolates anything actionable. ## Making it correct if you do it **Assign membership independently.** Holdout membership should be its own partition, hashed with its own dedicated salt, so that being in the holdout is uncorrelated with where anyone would have landed in any experiment. If holdout membership were derived from an experiment's split, the excluded group would be systematically related to that experiment's arms. **Enforce it at the platform, not by convention.** Every treatment application must check holdout membership before doing anything. Leaving that to each team's discretion means the holdout quietly stops being held out, and the failure is invisible until someone audits it. Audit it: measure how many held-out users were actually exposed to something and treat a nonzero number as an incident. **Give it an owner, a size and an expiry.** Size it to what your traffic supports and what precision the comparison needs. Name the person who will read it and the decision it will feed. Set the date on which it is dissolved or deliberately renewed, so it does not become permanent by inertia. ## When the answer is simply no At low traffic the holdout is self-defeating: the 10% you set aside is exactly the sample your experiments need, and the holdout comparison will itself be too imprecise to answer anything. The same is true when the product surface is changing so fast that the held-out experience cannot be kept working, or when nobody has committed to acting on the result. In those cases spend the traffic on experiments and get your assurance from disciplined metric definitions and repeated A/A validation of the platform instead. ## What a strong answer sounds like It refuses to answer in the abstract. It asks what decision the holdout informs, prices the traffic and the engineering tax against that decision, names the governance that keeps the holdout honest, and is willing to say no at traffic levels or product velocities that cannot support it. Candidates who describe the holdout as free because it is only 10% have missed the code-path cost entirely.
- How should holdout membership interact with the bucketing scheme?It should be its own partition, hashed with a dedicated salt, so membership is independent of every experiment's split - otherwise the excluded group is systematically related to some experiment's arms. Experiments then bucket only the remaining population, and every treatment application checks holdout membership first. That check belongs in the platform, not in each team's code.
- What breaks first when a global holdout is left in place for years?The engineering cost. Every shipped feature has to keep a working previous path for held-out users, so old code can never be deleted and the two versions drift until the holdout is effectively an unmaintained product. Comparability erodes alongside it, because the measured difference reflects the whole divergence rather than anything anyone can act on.
- Would you run a global holdout at low traffic?No. The 10% you set aside is exactly the sample every experiment needs, and the holdout comparison would itself be too imprecise to answer anything. Spend the traffic on experiments and get your assurance from disciplined metric definitions and repeated A/A validation of the platform instead.
- How do you know a global holdout is still actually held out?Audit it rather than trust it. Measure how many held-out users were exposed to any treated code path over the period, and treat a nonzero count as an incident with a root cause. Silent leakage is the normal failure mode, because it is invisible in every dashboard until someone checks membership against exposure logs.
saying these in an interview costs you the question
- Calls a global holdout free because it is only 10%
- Forgets every experiment now draws from 90% of traffic
- Ignores the cost of keeping old code paths alive
- Assumes the holdout stays comparable indefinitely
- Leaves holdout enforcement to each team's discretion
- Sets one up with no owner, no expiry and no decision attached