skip to content

Why is the sample mean of an arm abandoned early by a bandit allocator biased downward?

level: seniorimportance: must knowfreq 50%

answer

  1. sample size was not fixed in advance
  2. who gets starved, and why they got starved
  3. conditioning on early bad luck
  4. winner's curse runs the other way
  5. log the assignment probability, keep a floor

basics

~20 s

Because allocation depended on outcomes: an arm gets starved precisely after a run of unlucky results, so the few observations it holds are the unlucky ones. Its sample mean understates its true mean, and ordinary confidence intervals under-cover.

solid answer

~50 s

In an adaptive run the number of observations an arm receives is itself random and correlated with those observations. An arm that ends with only 200 records got starved *because* its early draws came out low, so conditioning on "few observations" is conditioning on early bad luck, and the sample mean sits below the true mean. The mirror image applies to the exploited arm: it was selected for being the maximum of noisy estimates, so its logged rate is biased upward — the winner's curse. Because sample sizes and membership depend on outcomes, an ordinary two-sample test on the full log is invalid: both the point estimate and the standard error are wrong, and interval coverage falls below nominal. Recover valid numbers by analysing a uniform-random holdout, by reweighting with logged assignment probabilities kept above a floor, or by rerunning the finalists as a fixed-horizon confirmatory test.

go deeper

for a junior

Be ready to say that the allocator decided how much data each arm received based on that arm's own early results, so the leftover data is not a fair sample of that arm.

for a middle

Explain the mechanism precisely: per-arm sample size is random and correlated with the outcomes, so conditioning on a starved arm means conditioning on unlucky draws.

for a senior

Show how you recover usable numbers — a uniform holdout, logged probabilities with an exploration floor and reweighting, or a confirmatory fixed-horizon rerun of the finalists — and say what the intervals mean afterwards.

for a principal

Set the policy: which numbers may be quoted to the business, and the requirement that propensity logging and an exploration floor exist before any adaptive allocation ships.

## Where the bias comes from In a fixed-horizon experiment, assignment is decided by a rule that ignores outcomes — a coin flip, a hash, a pre-registered split. That independence is what makes the difference in sample means unbiased and makes the usual standard errors correct. An adaptive allocator breaks the independence on purpose. The probability that the next user sees arm `a` is a function of the rewards already observed for arm `a`. That single change has two consequences: 1. **The per-arm sample size is random**, not fixed in advance. 2. **The sample size is correlated with the observed rewards** — high early rewards buy more traffic, low early rewards buy less. When a quantity you condition on is correlated with the values you are averaging, the average stops being unbiased. ## The mechanism, concretely Take an arm whose true conversion rate is 4%. Run many parallel worlds. In the worlds where its first few hundred users happened to convert above 4%, the allocator keeps feeding it, and it accumulates a large sample whose mean converges toward 4%. In the worlds where its first few hundred users happened to convert below 4%, the allocator cuts it off, and it is frozen with a small sample carrying that unlucky mean. Now select the worlds where the arm ended with only 200 observations: you have selected the unlucky worlds. Averaged over them, the sample mean lands **below** 4%. Nothing here is a small-sample artefact of the kind that vanishes with a bias correction factor. It is *selection on the outcome*, and it does not go away by collecting more total traffic — collecting more traffic only feeds the arms that already look good. ## The other direction: the winner's curse The same logic runs upward for the arm the allocator settled on. That arm was chosen for having the highest observed mean among several noisy estimates. Taking the maximum of noisy estimates is a positively biased operation: whichever arm is genuinely best is *also* more likely to have been flattered by noise on the day it was selected. So the exploited arm's logged rate overstates its true rate, typically by more when the arms are close together and the per-arm samples are small. Put the two together and the picture is systematic: an adaptive log exaggerates the spread between the winner and the losers. That is why replaying an adaptive run's numbers as "headline A beat headline B by 18%" is not conservative — it is inflated in exactly the direction people want to believe. ## Why the usual test is invalid too The standard two-sample comparison assumes independent, identically distributed observations with sample sizes fixed independently of the data. Adaptive allocation violates that on both counts. The consequences are: - The **point estimate** is biased, as described above. - The **standard error** is computed as if `n` were a constant, when it is a random variable entangled with the data, so it understates the true sampling variability. - **Coverage falls below nominal**: a nominal 95% interval built this way contains the true value less than 95% of the time, so the reported uncertainty is optimistic in addition to the estimate being wrong. ## How to get valid numbers anyway **Keep a uniform-random holdout.** Route a fixed slice of traffic — say 5 or 10 percent — by pure random assignment for the whole run, independent of the allocator. Analysed on its own, that slice is an ordinary randomised experiment: unbiased differences, honest intervals. You pay a fixed, predictable amount of regret and get less precision than a full fixed-horizon test would give, but the inference is clean and needs no clever estimator. **Log the assignment probability and reweight.** If for every unit you record the probability `p` with which the allocator assigned it to the arm it got, you can weight each observation by `1/p` to build an unbiased estimator of an arm's mean under uniform assignment. This is only usable if you *also* impose an **exploration floor**, keeping every arm's probability bounded away from zero — otherwise a probability near zero produces an enormous weight and an estimator with unusable variance. The floor and the logging must be designed in before the run; they cannot be reconstructed afterwards. **Rerun the finalists.** When a decision genuinely needs a measured effect, take the two or three arms the adaptive run surfaced and run them as a small fixed-horizon confirmatory test. The adaptive phase then serves as screening, and the confirmatory phase supplies the number you publish. This also directly corrects the winner's curse, because the arm's effect is re-estimated on data that did not select it. ## What to say to stakeholders Be explicit about which numbers are quotable. "Which arm we serve" comes from the adaptive run; "how much better it is" comes from the holdout or the confirmatory test. If neither exists, the honest answer is that the run picked a winner and did not measure it — and the fix is a process change before the next run, not a cleverer analysis of this one.

  • Which direction is the winning arm's logged estimate biased, and why?
    Upward. The arm was selected for having the highest observed mean among noisy estimates, and taking a maximum over noise is positively biased — the winner's curse. Re-estimate its effect on data that did not select it, either a uniform holdout or a confirmatory fixed-horizon run, before quoting any magnitude.
  • How does keeping a uniform-random holdout restore valid inference?
    Within the holdout, assignment is independent of outcomes, so per-arm means are unbiased and ordinary confidence intervals have their stated coverage. You trade a fixed slice of regret and some precision for an analysis that needs no special estimator and is easy to explain to reviewers.
  • Why does simply collecting more total traffic not fix the bias?
    Extra traffic flows to the arms that already look good, because that is what the allocator does. The starved arm stays starved, so its sample stays small and selected on bad luck. Bias from outcome-dependent assignment is not a small-sample artefact that averages away.
  • What must be in place before the run for reweighting to be usable?
    Two things: per-unit logging of the assignment probability actually used, and an exploration floor keeping every arm's probability bounded away from zero. Without the floor, tiny probabilities produce huge inverse weights and an estimator too noisy to use; without the logging, the weights cannot be reconstructed at all.

saying these in an interview costs you the question

  • Says collecting more data will fix the bias
  • Runs a two-sample t-test on the full adaptive log
  • Calls the starved arm's mean unbiased, merely noisy
  • Treats the exploited winner's logged rate as trustworthy
  • Plans to reweight without having logged assignment probabilities

context