skip to content

If you stop an A/B test at the first significant look, what happens to the measured effect size?

level: seniorimportance: nice to knowfreq 31%

answer

  1. the decision is wrong less often than the number
  2. crossing early demands a big absolute effect
  3. you selected the flattering moment
  4. worst when true effect is small
  5. treat the readout as an upper bound

basics

~20 s

It is biased upward. Stopping when the statistic crosses the threshold selects the moments when noise happened to help, so the effect you report is systematically larger than the truth — and the earlier you stop, the worse the exaggeration.

solid answer

~50 s

Optional stopping does not only break the error rate; it corrupts the estimate. Crossing a significance boundary requires the observed difference to exceed roughly a fixed number of standard errors, and early in a test the standard error is large, so only an unusually flattering sample path crosses. Stopping there means you have conditioned on a favourable draw, and the estimate at the stopping time is biased away from zero. The bias is worst when the true effect is small relative to the noise and when you stop earliest — exactly the tests that most tempt an early call. Practically, this is why a portfolio of early-stopped wins under-delivers: the shipped effects were real less often than claimed and smaller than reported even when real. If a test was stopped early, treat its point estimate as an upper bound and re-measure before anyone forecasts revenue from it.

go deeper

for a junior

Know that a result captured the moment it first looked significant tends to overstate the effect, and that a later, properly run measurement usually comes back smaller.

for a middle

Explain the mechanism: the crossing bar is a large absolute effect early on, so only favourable sample paths cross, and stopping there conditions the estimate on good luck.

for a senior

Show you can act on it — treat the readout as an upper bound, re-measure before it enters a forecast, and resist explaining a shrinking rerun as effect decay before checking how the first number was obtained.

for a principal

Own the portfolio view: claimed lifts summed across a year versus actual top-line movement, and the organisational cost of planning against systematically inflated estimates.

## Two separate damages Most discussions of optional stopping stop at the false-positive rate — the chance of declaring a winner when the arms are identical. That is real, but there is a second, quieter damage that survives even when the effect **is** real: the number you report is too big. This matters commercially. A false positive gets caught eventually when the metric does not move. An inflated true positive never gets caught — the feature did help, just by a third of what the readout promised, and the forecast built on it silently misses. ## The selection mechanism Declaring significance means the observed difference cleared roughly a fixed number of standard errors: `observed difference > c * SE`. The standard error shrinks as the sample grows, so early in a test that bar is a **large absolute effect**. Late in the test, the same bar is a small one. Now consider what it means to stop at the first crossing. You are not sampling the estimate at a moment chosen independently of the data — you are sampling it at the moment it was most extreme, and specifically at a moment when it exceeded a high bar. Conditioning on "the estimate was big enough to cross" selects sample paths where the noise pushed in your favour. The average of those selected estimates sits above the true effect. This is a selection effect, not a computational error: the arithmetic on the day was correct, and the number is still too large. The severity scales in an intuitive way: - **Earlier stops are worse.** The bar in absolute terms is higher, so more of the crossing has to be supplied by luck. - **Small true effects are worse.** If the truth is a 0.2% lift and the early bar is a 2% lift, essentially all of the observed 2% is noise. - **Low-powered tests are worse.** When the design had little chance of detecting the true effect at all, the crossings that do occur are dominated by favourable noise. In the extreme — a genuinely tiny effect detected by a badly under-powered early look — the selected estimate can even carry the **wrong sign**, because the only paths that cross are extreme excursions and some of those come from the opposite direction of the truth. ## Why teams misread the follow-up When a stopped-early win is re-measured later and comes back smaller, the usual explanation offered is that the effect wore off, or that the launch was diluted, or that the second test was underpowered. Sometimes true. But the mechanical prediction of optional stopping is exactly this: **the first, selected estimate was inflated, and the second, unselected one is closer to the truth.** A candidate who can say "before I reach for behavioural explanations, I would ask whether the first estimate was taken at a stopping time chosen by the data" is demonstrating the judgment being probed here. The sharpest illustration is a stream where the estimate is watched over time: a test that would have been declared a solid win on day 3 and, from the very same accumulating data, read as a loss on day 10. Nothing about the world changed between those readings. What changed is that the day-3 reading was taken at a moment selected for being impressive, and the day-10 one was not. ## Practical consequences **Treat an early-stopped estimate as an upper bound.** If a decision has already been made on it, that is water under the bridge, but do not let the magnitude enter a forecast or a roadmap projection unqualified. **Re-measure before you plan around it.** A follow-up test at a fixed horizon, or a holdback measured after launch, gives an estimate that was not selected for being large. Expect it to come in lower, and say so in advance so the shrinkage does not read as a failure of the follow-up. **Audit the portfolio, not the test.** The strongest evidence that a team has an optional-stopping problem is aggregate: sum the claimed lifts from a year of shipped wins and compare against the actual movement in the top-line metric. A large gap that cannot be explained by interaction or seasonality is the signature of systematically inflated estimates. **Do not try to fix it by shrinking arbitrarily.** Halving every early-stopped estimate because it feels inflated is not a correction; the bias depends on how early you stopped, the true effect and the noise, none of which you know. The honest options are to re-measure, or to have designed for early stopping in the first place so that the estimate has a defined correction available. ## What an interviewer is listening for They want you to separate the two harms — calibration of the decision and bias of the estimate — and to notice that the second one is the expensive one at a company that ships a lot of experiments. A candidate who only recites "peeking inflates false positives" has half the picture; the other half is why the wins that *were* real still did not add up.

  • A test stopped early showed +6%, and a clean rerun shows +1.5%. How do you explain the gap to the business?
    The first number was measured at a moment chosen because it looked good, so it was selected for being flattering; the rerun measured at a fixed endpoint and is the better estimate. I would present the +1.5% as the planning number, note that the feature is still a genuine win, and be explicit that the gap is a known property of early stopping rather than a contradiction between the two tests.
  • Can you simply apply a fixed discount to every early-stopped estimate?
    No. The bias depends on how early you stopped, the true effect size and the noise level, and you do not know the first two. A flat discount is guesswork dressed as a correction and can overshoot on large real effects while undershooting on tiny ones. The reliable options are re-measuring at a fixed horizon or having chosen a design with a defined adjustment before launch.
  • How would you detect that a whole team has this problem rather than auditing one test?
    Compare the sum of claimed lifts from a period of shipped wins against the actual movement in the top-line metric over the same period. A persistent gap that seasonality and interaction effects cannot explain points at systematically inflated estimates. Pair it with a simple operational check: what fraction of tests were stopped before their planned end date.

It is like recording your fastest lap out of forty and calling it your typical pace. The stopwatch was honest; the lap was chosen for being the one where everything went right.

saying these in an interview costs you the question

  • Thinks stopping early only affects significance, not the estimate
  • Assumes the reported lift is unbiased because the arithmetic was right
  • Explains a shrinking rerun purely as the effect fading
  • Applies a flat discount and calls it a bias correction
  • Puts an early-stopped point estimate straight into a revenue forecast

context