skip to content

Sensitivity & Adaptive Designs

Getting more signal out of the same traffic: covariate adjustment that strips variance, tests that stay valid at every look, and allocation that shifts toward the winner as it learns.

on this pageshow

explore

questions

15

Would you use a bandit allocator or a fixed-horizon A/B test for twelve news headlines whose value expires within hours?

level: middleimportance: must knowfreq 66%

answer

  1. earn during the run or measure after
  2. how many arms, how fast the feedback
  3. value expires before a powered sample lands
  4. you need a winner, not a number
  5. keep a small uniform holdout for inference

basics

~20 s

Use adaptive allocation. With twelve options, clicks arriving in seconds and value gone in hours, the cost of serving losers dominates and no precise per-headline lift is needed. A fixed-horizon test would still be running when the headlines are worthless.

solid answer

~50 s

Adaptive allocation, because the decision fits every condition it is good for: you need a winner served now rather than a measured effect, there are twelve arms, feedback (clicks) arrives within seconds, and the whole decision repeats tomorrow with new headlines. A fixed-horizon test would need a powered sample per arm plus multiplicity control across eleven comparisons, and would report after the news cycle has passed — the answer arrives correct and useless. What you give up is inference: the click rates logged by the allocator are not unbiased estimates of each headline's quality, so do not publish them as measured lift. If you want anything defensible out of the run, hold back a small slice of traffic for uniform random assignment, or log each impression's assignment probability so the data can be reweighted later.

go deeper

for a junior

Be ready to pick one design and give a concrete reason: many short-lived options with instant feedback favour shifting traffic adaptively rather than waiting for a fixed sample.

for a middle

Explain both mechanics — the per-arm sample size and multiplicity cost of twelve arms in a fixed-horizon test, against what adaptive allocation gives up in estimate quality.

for a senior

Name the conditions that flip the choice, such as lagged rewards, one-shot permanent decisions or interference, and design the uniform holdout and exploration floor you would keep.

for a principal

Frame it as an organisational objective choice — earnings during the run versus a defensible effect size — and decide who may cite adaptive-run output as evidence.

## The question behind the question "Bandit or fixed-horizon test?" is really "what am I buying with this traffic — earnings during the run, or a measured effect afterwards?" Adaptive allocation buys earnings: it shifts traffic toward whichever arm currently looks best, so fewer users see a loser. A fixed-horizon test buys measurement: a pre-committed sample size, an allocation that does not depend on outcomes, and therefore an unbiased difference estimate with a valid confidence interval. You cannot maximise both with the same traffic. ## Score the decision on five axes **1. Objective.** Do you need to *serve* the best option, or to *know how much better* it is? Twelve headlines need serving. A pricing change that will ship permanently needs knowing. **2. Number of arms.** A fixed-horizon test with twelve arms means eleven comparisons against a control, or sixty-six pairwise ones. Controlling the family-wise error rate makes each comparison harder, and the required total sample grows roughly with the number of arms if you want each one powered. Adaptive allocation scales far more gracefully: it spends most impressions on the two or three plausible winners and only enough on the rest to rule them out. **3. Feedback latency versus decision horizon.** Adaptive allocation only works if the reward lands fast enough to steer the next allocation update. Headline clicks arrive within seconds against a horizon of hours — a ratio of thousands. If the reward were a subscription that lands next week, the allocator would be steering on censored data. **4. Repeatability.** Headline selection recurs every day, automatically, with new candidates. There is no human decision meeting to feed, so there is no consumer for a confidence interval. One-shot decisions that a person must sign off on are the opposite case. **5. Precision requirement.** Nobody needs to know that headline 7 beat headline 3 by 0.42 points. If someone does need that number — for a model, a forecast, or a business case — the fixed-horizon test is the only design that gives it cleanly. On all five, the headline problem points the same way. ## What the fixed-horizon test would actually cost here Suppose each headline's click rate is around 4% and you would care about a relative difference of 10%. Powering even a single such comparison takes on the order of tens of thousands of impressions per arm; with twelve arms and multiplicity control, the total climbs further. Meanwhile the headline is stale in hours. The test does not fail statistically — it fails *temporally*. Its answer would be perfectly valid and completely worthless, and every impression spent on the eleven losers during that window is pure loss. ## What adaptive allocation gives up The crucial cost is not statistical sophistication — it is that **allocation depends on observed outcomes**, which breaks ordinary inference. Arms that were starved carry few observations, and those observations are the ones that caused the starving. Their sample means understate their true rates. The arm that was exploited was picked for looking best, so its logged rate flatters it. Running a two-sample test on the raw log is invalid, and the numbers should never be reported as measured lift. There are practical repairs, and a mature setup applies them by default: - **An exploration floor.** Guarantee every arm some minimum share of traffic. This bounds how badly any arm can be starved and keeps assignment probabilities away from zero. - **A uniform holdout.** Route a small fixed fraction of traffic by pure random assignment. Analysed alone, that slice is an ordinary randomised experiment with unbiased estimates and valid intervals — a clean, if less precise, measurement running inside the adaptive system. - **Logged assignment probabilities.** If you record the probability with which each impression was assigned its arm, the log can be reweighted afterwards to recover unbiased estimates of each arm's mean. ## When the answer flips Switch to a fixed-horizon test when: the reward metric matures slower than your update cadence; the decision is one-shot and permanent; you need an effect size defensible to finance, legal or a model that consumes it; the arms interfere with each other so serving more of one changes another's performance; or the metric is a guardrail whose regression you must detect rather than a reward you are chasing. ## How to answer this in an interview Commit to a choice, name the two or three properties of the problem that drive it, and then volunteer the cost of your choice and the mitigation. Candidates who say only "bandits are more efficient" are giving a slogan; candidates who say "adaptive allocation, because feedback is seconds and value is hours and I need a winner rather than a number — and I will keep a 5% uniform holdout so I can still learn something honest" are giving an engineering answer.

  • What would change your answer to a fixed-horizon test here?
    A reward that matures slowly, so the allocator would steer on censored data; a one-shot permanent decision that someone must sign off on; a need for a defensible per-headline effect size; or interference between arms. Any of those makes the measurement worth more than the earnings saved during the run.
  • How do you still learn something trustworthy from an adaptive run?
    Two options that compose. Reserve a small fixed share of traffic for uniform random assignment and analyse only that slice as an ordinary experiment. Or log each impression's assignment probability, keep those probabilities bounded away from zero with an exploration floor, and reweight the log afterwards to recover unbiased per-arm means.
  • Does having twelve arms by itself favour adaptive allocation?
    Largely yes. A fixed-horizon design must power eleven comparisons and control the family-wise error rate, which inflates the total sample needed. Adaptive allocation concentrates impressions on the few plausible winners and only spends enough elsewhere to rule arms out, so cost grows with the number of *contenders*, not the number of candidates.

saying these in an interview costs you the question

  • Runs a twelve-arm fixed test with no multiplicity control
  • Picks bandits because they are 'always faster'
  • Reports bandit-logged click rates as measured lift
  • Ignores whether feedback arrives before the decision is needed
  • Assumes adaptive allocation removes the need for a reward metric

context

open as a page

What is CUPED, and how does a pre-experiment covariate reduce the variance of an A/B test?

level: middleimportance: must knowfreq 62%

basics

~20 s

CUPED replaces outcome Y with Y - theta*(X - mean X), where X is a pre-experiment covariate. With theta = cov(Y, X) / var(X), variance falls to (1 - rho^2) of the original, rho being the correlation.

open as a page

In a group-sequential A/B test, what does an alpha-spending function do?

level: middleimportance: must knowfreq 55%

basics

~20 s

An alpha-spending function states in advance how much of the total 0.05 false-positive budget each interim analysis may consume, as a function of how much of the planned data has arrived. The pieces sum to 0.05.

open as a page

Why is the sample mean of an arm abandoned early by a bandit allocator biased downward?

level: seniorimportance: must knowfreq 50%

basics

~20 s

Because allocation depended on outcomes: an arm gets starved precisely after a run of unlucky results, so the few observations it holds are the unlucky ones. Its sample mean understates its true mean, and ordinary confidence intervals under-cover.

open as a page

What does cumulative regret measure in a multi-armed bandit experiment?

level: juniorimportance: should knowfreq 42%

basics

~10 s

Cumulative regret is the total reward given up by not serving the best arm every round: summed over rounds, the best arm's true mean reward minus the mean reward of the arm actually served.

open as a page

Why does a post-exposure covariate bias a CUPED-adjusted treatment effect estimate?

level: middleimportance: should knowfreq 45%

basics

~20 s

CUPED subtracts theta times the arms' difference in the covariate, which averages to zero only for a covariate fixed before exposure. A covariate the treatment can move carries real effect, so subtracting it deletes part of that effect.

open as a page

How do O'Brien-Fleming and Pocock group-sequential boundaries differ?

level: middleimportance: should knowfreq 45%

basics

~20 s

O'Brien-Fleming boundaries are extremely stringent at early looks and relax to nearly the fixed-sample cutoff at the end. Pocock uses one constant, looser cutoff at every look, so it stops earlier but gives up power at the final analysis.

open as a page

Can CUPED reduce variance in a test that runs only on new users with no pre-period history?

level: seniorimportance: should knowfreq 38%

basics

~20 s

Essentially no. With no pre-period behaviour every user gets the same covariate value, so its variance is zero and theta is undefined. Only attributes known at assignment, such as acquisition channel or device, remain, and they buy little.

open as a page

What guarantee does an always-valid confidence sequence give that a fixed-horizon confidence interval does not?

level: seniorimportance: should knowfreq 38%

basics

~20 s

A confidence sequence covers the true parameter at every sample size simultaneously with probability at least 1 minus alpha, so it can be read or stopped at any moment. A fixed-horizon interval guarantees coverage only at one pre-specified sample size.

open as a page

When should adaptive allocation be your platform's default instead of fixed-horizon A/B tests?

level: principalimportance: should knowfreq 34%

basics

~20 s

Default to adaptive allocation only where decisions repeat automatically, options are short-lived and the reward matures fast, and the org needs a winner rather than a measured effect. Keep fixed-horizon tests wherever a defensible lift drives the decision.

open as a page

Would you make always-valid sequential testing the default on an experimentation platform?

level: principalimportance: should knowfreq 32%

basics

~20 s

Usually yes on a self-serve platform, because people read results whenever they like and sequential methods stay honest under that behaviour. The tradeoff is efficiency: the same conclusion needs more data than a fixed-horizon design.

open as a page

Why does a bandit allocator updating hourly lock onto a stale winner when conversions land three days later?

level: seniorimportance: nice to knowfreq 30%

basics

~20 s

Because exposures count immediately while conversions arrive days later, any recently promoted arm looks artificially bad at update time. The early leader, whose conversions have matured, keeps winning the comparison and reinforces its own head start.

open as a page

When would you use pre-experiment stratification on platform or country instead of a CUPED covariate?

level: seniorimportance: nice to knowfreq 27%

basics

~20 s

Stratify when no continuous pre-period metric exists and the metric differs sharply between a few large discrete groups. Its ceiling is the share of variance sitting between strata, which is usually well below what a user's own pre-period value explains.

open as a page

What does a futility boundary add to a group-sequential experiment, and what does it cost?

level: seniorimportance: nice to knowfreq 26%

basics

~20 s

A futility boundary stops an experiment early when the accumulated data make crossing the efficacy boundary implausible, freeing traffic and time. The cost is power: some experiments that would eventually have reached significance are killed before they get there.

open as a page

As experimentation lead, would you make CUPED the default analysis for every test on the platform?

level: principalimportance: nice to knowfreq 22%

basics

~20 s

Default it on only where the platform computes a locked, pre-registered covariate per metric and A/A validation confirms nominal error rates. Blanket promises fail because the gain is metric-dependent, and analyst-chosen covariates turn a variance tool into a p-hacking surface.

open as a page