skip to content

Why does stopping an A/B test as soon as its p-value drops below 0.05 inflate Type I error?

level: middleimportance: must knowfreq 74%

answer

  1. the 5% belongs to the procedure
  2. one test, one pre-set sample size
  3. each look is another chance to cross
  4. you are sampling a maximum, not a value
  5. looks share data, so not 1 - 0.95^k

basics

~20 s

A fixed-horizon p-value is calibrated for one analysis at one pre-set sample size. Each extra look gives noise another chance to cross the threshold, so stopping at the first p < 0.05 rejects far more than 5% of null tests.

solid answer

~50 s

The 5% error rate belongs to the whole decision procedure, not to any single look. A standard two-sample test assumes you analyse once, at a sample size fixed before launch. If instead you evaluate repeatedly and stop the moment the threshold is crossed, you have built a multiple-testing procedure in time: the cumulative test statistic wanders as data accrues, and you are sampling its maximum rather than its value at a fixed endpoint. The looks are correlated because they share data, so the inflation is not `1 - 0.95^k`, but it is large: five equally spaced looks at a nominal 0.05 give roughly a 14% false-positive rate, and looking daily across a two-week test pushes it past 20%. The fix is to fix the horizon before launch and honour it, or to use a design intended for continuous monitoring — not to peek and hope.

code

python · 16 lines
python
import random, math, statistics

def peeked_test(looks=5, batch=200):
    a, b = [], []
    for _ in range(looks):                     # A and B drawn from the SAME distribution
        a += [random.gauss(0, 1) for _ in range(batch)]
        b += [random.gauss(0, 1) for _ in range(batch)]
        se = math.sqrt((statistics.variance(a) + statistics.variance(b)) / len(a))
        if abs(statistics.mean(a) - statistics.mean(b)) / se > 1.96:
            return True                        # saw p < 0.05 at this look, called a winner
    return False

trials = 4000
peeking = sum(peeked_test() for _ in range(trials)) / trials
one_look = sum(peeked_test(looks=1, batch=1000) for _ in range(trials)) / trials
print(round(peeking, 3), round(one_look, 3))   # ~0.14 vs ~0.05

go deeper

for a junior

Be ready to say plainly that the 5% error rate assumes a single analysis at a sample size chosen before launch, and that checking repeatedly and stopping at the first green result breaks that assumption.

for a middle

Explain the mechanism, not just the rule: the cumulative statistic wanders, repeated looks sample its maximum, and the looks are correlated because they share data. Carry one concrete number, such as roughly 14% for five looks.

for a senior

An interviewer expects you to name the variants that break calibration the same way — extending a test that has not turned yet, informal dashboard-watching, quoting the flattering mid-flight reading — and to say how you would stop them happening on your team.

for a principal

Own the framing that the error rate is a property of the organisation's decision procedure, not of the analysis code. Argue for where pre-commitment lives, what it costs in speed, and when a monitoring-capable design is worth buying instead.

## What the 5% actually promises A significance level is a property of a **procedure**, not of a number on a screen. When you say "I will run a two-sample test at alpha = 0.05", the promise is: *if the two arms are truly identical, this procedure declares a winner at most 5% of the time.* That promise is derived under a specific protocol — collect a sample size decided in advance, compute the statistic once, compare it to the critical value once. That protocol is the **fixed-horizon assumption**. Optional stopping breaks the protocol. The data are the same; the procedure is not. ## The mechanism: sampling a maximum, not a value As users accumulate, the test statistic is recomputed on a growing sample. Its path over time is not a smooth march toward the truth — it is a noisy trajectory. Under a true null, the difference in means hovers around zero, but it hovers with a spread, and the critical boundary is a fixed number of standard errors away. A single fixed-horizon test asks: *is the trajectory outside the boundary at the one moment I chose?* The answer is yes 5% of the time. A peeking procedure asks: *does the trajectory ever go outside the boundary at any of the moments I looked?* That is a question about the **maximum** of the trajectory across many time points, and the maximum of a wandering series exceeds a threshold far more often than its value at one arbitrary point does. Every additional look is another chance to catch the series on an excursion. ## How much inflation The looks are strongly correlated — the sample at look 3 contains all the data from looks 1 and 2 — so they are nowhere near independent tests. A candidate who computes `1 - 0.95^5 = 22.6%` has the right instinct and the wrong model; correlation makes the true number smaller than that. The standard repeated-significance figures for equally spaced looks at a nominal two-sided 0.05 are roughly: | Looks | Overall false-positive rate | |---|---| | 1 | 5% | | 2 | 8% | | 3 | 11% | | 5 | 14% | | 10 | 19% | | 20 | 25% | So a team checking a dashboard every morning of a two-week test and shipping at the first green cell is running at roughly a one-in-five false-positive rate, not one in twenty. And in the limit — monitor continuously, with no cap on sample size — a null test is *guaranteed* to cross the boundary eventually. "Just run it a bit longer until it turns significant" is not a way to find the truth; it is a way to find noise with certainty. ## What is and is not the sin Looking is not the problem. Data does not know it is being observed. The problem is that the **decision rule depends on what you saw**. Formally, the sampling distribution of your test statistic at the moment you stop is not the fixed-horizon distribution, because the stopping time is itself a function of the data. This matters practically, because it identifies the real-world variants that also break calibration: - **Stopping early on a win.** The classic case. - **Extending a test that has not "turned" yet.** Same defect, opposite direction — you are letting the result choose the sample size. - **Informal peeking.** Watching a dashboard daily and "deciding when it looks stable" is a stopping rule, just an undocumented one. - **A result that flips.** The same data stream can look like a clear win on day 3 and a clear loss on day 10; both readings come from a trajectory that had not settled, and whichever one you happened to act on is the one you would have defended. By contrast, a team that renders the numbers daily but genuinely does not act until the pre-registered end date has *not* inflated anything. Their stopping time is fixed in advance, which is the only property the calibration needs. The practical difficulty is that humans rarely watch a number for two weeks and remain uninfluenced by it — which is why the durable fix is procedural, not a promise. ## What to do instead Decide the horizon before launch and write it into the test plan alongside the metric and the decision rule, then honour it when the result is inconvenient. If you genuinely need to act early — a high-traffic test where a week of a bad variant is expensive — that is a design requirement to satisfy before launch with a method built for continuous monitoring, not something you can retrofit by peeking at a fixed-horizon p-value and hoping. One clarification worth carrying into the interview: none of this makes the *final*, planned-horizon analysis invalid. If you looked, resisted, and analysed at the pre-set sample size, that analysis is exactly the 5% test you designed. The damage comes only from letting the looks change what you do.

  • If a team looks at the dashboard daily but only decides on the pre-registered end date, is their Type I error inflated?
    No. Calibration depends on the stopping time being fixed in advance, not on whether anyone looked. Observing data does not change its distribution. The catch is behavioural: people who watch a number for two weeks rarely stay uninfluenced by it, and "we would have stopped if it had looked bad enough" is a stopping rule. That is why the durable fix is hiding early significance rather than trusting restraint.
  • A test crossed p < 0.05 on day 4 and is still significant at the planned day-14 analysis. Is the final result valid?
    Yes, provided the mid-flight crossing changed nothing — you did not stop, extend, or alter the metric because of it. The day-14 analysis at the pre-set sample size is exactly the 5% test you designed. What is not valid is quoting the day-4 result, or having decided on day 4 to keep running only because the number was pleasing.
  • Why isn't the inflation from five looks equal to 1 - 0.95^5, about 23%?
    That formula assumes five independent tests. Sequential looks are heavily correlated because each one contains all the earlier data — the statistic at look 5 is mostly determined by the sample that produced look 4. The correlation drags the true rate down to roughly 14%. The instinct that repeated tests compound is right; treating them as independent overstates it.
  • What happens to the false-positive rate if you monitor continuously with no cap on sample size?
    It goes to 1. Under a true null the cumulative statistic keeps wandering, and with unbounded time it is certain to cross any fixed boundary eventually. "Run it until it turns significant" therefore always succeeds, which is precisely why it proves nothing. Any legitimate continuous-monitoring scheme has to tighten the boundary over time rather than hold it fixed.

A fixed-horizon test is one throw of a dart at a small target. Peeking is throwing every morning and stopping when you finally hit — the target never got bigger, but hitting it stopped meaning anything.

saying these in an interview costs you the question

  • Says more data can only improve accuracy, so peeking is harmless
  • Treats the p-value as a quantity that converges as the sample grows
  • Assumes the looks are independent, so five looks means 23%
  • Thinks stopping early only costs precision, not correctness
  • Says a p-value is the probability that the null hypothesis is true
  • Claims extending a flat test until it turns significant is just patience

context