When should adaptive allocation be your platform's default instead of fixed-horizon A/B tests?
answer
- segment the decisions, not the platform
- who consumes the effect size
- reward maturity versus decision horizon
- propensity logging, floor, holdout as entry price
- an effect-size library you stop accumulating
basics
~20 sDefault to adaptive allocation only where decisions repeat automatically, options are short-lived and the reward matures fast, and the org needs a winner rather than a measured effect. Keep fixed-horizon tests wherever a defensible lift drives the decision.
solid answer
~40 sTreat it as a segmentation of the decision space, not a platform-wide preference. Adaptive allocation should be the default on surfaces where the choice recurs automatically, arms are numerous and short-lived, the reward matures far faster than the decision horizon, and nobody consumes an effect size — ranking creative, content or ordering decisions. Fixed-horizon tests stay the default for one-shot permanent changes, guardrail monitoring, anything with a slow-maturing outcome, and anything whose magnitude feeds a business case or a model. The organisational cost is what people underestimate: adaptive allocation needs logged assignment probabilities, an exploration floor, a uniform holdout, and a rule about which numbers may be quoted as evidence. And a company that runs everything adaptively stops accumulating measured effect sizes, which is what makes future power calculations and roadmap forecasts possible.
go deeper
Be ready to name one decision type for each design: repeated short-lived content choices suit adaptive allocation, a one-off permanent change needs a fixed-horizon test.
Explain the conditions that must all hold for adaptive allocation — recurring automated decisions, fast-maturing reward, many arms, no consumer for the effect size, no interference between arms.
Demonstrate the operational entry price: probability logging, an exploration floor, a uniform holdout, a stale-feed fallback, and a clear line between serving decisions and quotable evidence.
Own the strategic tradeoff — the effect-size library the organisation stops accumulating, who is allowed to cite which numbers, and how you audit whether the default is set correctly per surface.
## Frame the decision as two different products A fixed-horizon test and an adaptive allocator are not two speeds of the same thing. They produce different outputs. A test produces an **estimate with uncertainty** — a number you can put in a business case, feed to a forecast, or defend in a review. An allocator produces a **serving policy** — traffic keeps flowing to whatever currently looks best, and there is no number to defend. Deciding the default means deciding which output each part of the business needs, surface by surface. ## Where adaptive allocation should be the default Every one of these should hold, not just one: - **The decision recurs automatically.** New candidates arrive continuously and no human signs off on each choice, so there is no consumer for a confidence interval. - **The reward matures much faster than the decision horizon.** Seconds-to-minutes feedback against hours-to-days of value. If the outcome takes as long to mature as the decision takes to matter, adaptivity is steering on censored data. - **There are many arms.** Powering a comparison for each candidate is wasteful when you only care which one to serve. - **Nobody consumes the magnitude.** If no downstream model, forecast or business case reads the effect size, you are not losing anything by not measuring it. - **Arms do not interfere.** Serving more of one arm must not change another's performance through shared inventory, marketplace effects or user-level spillover. Content ranking, creative selection, ordering of modules on a surface, and similar high-frequency choices usually satisfy all five. ## Where fixed-horizon must stay the default - **One-shot permanent changes.** A pricing change, a checkout redesign, a policy change: decided once, shipped forever, and someone must be able to say how much it moved the metric and with what uncertainty. - **Slow-maturing outcomes.** Retention, subscription, repeat purchase, anything measured over weeks. - **Guardrails.** You are looking for a *regression* you must detect, not a reward you are chasing; the design must be powered to catch it. - **Evidence for others.** Finance, legal, leadership, or a downstream model that consumes the effect size. - **Interference-prone settings.** Marketplaces and social graphs, where adaptive traffic shifts change the environment itself. ## The organisational costs people underestimate **Inference infrastructure is not optional.** To get anything trustworthy out of adaptively allocated traffic, you need per-unit assignment probabilities logged, an exploration floor keeping those probabilities away from zero, and ideally a permanent uniform-random holdout. Retrofitting any of this after a run is impossible. **Evidentiary discipline.** The single most common organisational failure is an adaptive run's logged rates appearing in a business review as measured lift. Those numbers are biased — the exploited arm flatters itself, the starved arms understate themselves. The platform needs an explicit rule about which outputs may be cited as evidence and which are only serving decisions, ideally enforced by the reporting layer rather than by etiquette. **Loss of the effect-size library.** This is the strategic cost and it is invisible for a year. Fixed-horizon tests accumulate measured magnitudes: how much a typical change on this surface moves this metric, what the metric's variance is, what minimum detectable effect a given traffic level supports. That library is what makes future power calculations, roadmap forecasts and prioritisation possible. An organisation that runs everything adaptively accumulates wins with no measured sizes and eventually cannot plan. **Skill and debugging load.** Adaptive systems fail quietly. Diagnosing lock-in, censored rewards, or drift requires people who understand the allocator; a fixed test that goes wrong is usually obvious. ## The hybrid default worth arguing for A good platform position is: fixed-horizon remains the default for anything a person decides once; adaptive allocation is opt-in per surface, gated on four requirements — a reward that matures inside the update cadence, logged probabilities with an exploration floor, a uniform holdout carved out for measurement, and a named owner for the reward metric. Additionally, cap how much allocation may move per update and fall back to an even split when data is sparse or the reward feed is broken; the worst adaptive failures come from an allocator confidently acting on a stale or empty signal. ## Answering the "tests take too long" argument Teams frequently ask for bandits because experiments feel slow. That is a category error worth naming directly: test duration is set by the effect size you want to detect, the metric's variance and the traffic available — not by the allocation rule. Adaptive allocation reduces the *cost of running* by serving fewer users the losing arm; it does not reduce the data needed to be confident. If speed is the real problem, the levers are variance reduction, a better-chosen primary metric, a larger minimum detectable effect, or more traffic. Handing the team an allocator instead gives them a faster decision with no measurement, which is only an improvement if they never needed the measurement. ## What to measure about the programme itself Track the share of decisions running under each design, the realised regret saved on adaptive surfaces, the coverage and width of the holdout intervals, and how often an adaptive run's winner survives a confirmatory fixed-horizon rerun. If winners frequently fail confirmation, the allocator is being steered by something other than arm quality and the default should tighten.
- What guardrails would you require before enabling adaptive allocation broadly?Logged assignment probabilities with an exploration floor, a permanent uniform-random holdout for measurement, a cap on how much share may move per update, a rule that the reward matures within the update cadence, and automatic fallback to an even split when the reward feed is stale or data is sparse.
- How do you answer a team that wants bandits because 'tests take too long'?Duration is set by the effect size, the metric's variance and available traffic — not by the allocation rule. Adaptive allocation lowers the cost of running, not the data needed for confidence. If speed is the real problem, the levers are variance reduction, a larger minimum detectable effect, or more traffic.
- What does an organisation lose if every decision runs adaptively?The library of measured effect sizes that future power calculations, forecasts and prioritisation depend on, plus comparable readouts across quarters and the ability to audit why something shipped. You accumulate wins with no magnitudes, and within a year planning conversations have nothing quantitative to stand on.
- How would you know the default is set wrong on a surface?Re-run a sample of adaptive winners as small confirmatory fixed-horizon tests. If winners frequently fail to confirm, the allocator is being steered by something other than arm quality — censored rewards, interference or drift — and that surface should move back to fixed allocation until the cause is fixed.
saying these in an interview costs you the question
- Adopts adaptive allocation because it 'needs no statistics'
- Promises bandits shorten every decision
- Quotes adaptive-run rates as measured lift in reviews
- Ships adaptive allocation with no holdout or probability logging
- Applies one default across every surface and decision type