skip to content

If you run 20 independent hypothesis tests at alpha = 0.05, how likely is at least one false positive?

level: juniorimportance: must knowfreq 76%

answer

  1. alpha is a per-test promise
  2. work with the complement
  3. 0.95 multiplied by itself 20 times
  4. roughly two chances in three

basics

~20 s

About 64 percent. Each test independently avoids a false positive with probability 0.95, so all 20 stay clean with probability 0.95^20 = 0.36, leaving roughly a 64 percent chance that at least one comes back significant by luck.

solid answer

~50 s

Alpha is a promise about one test, not about a batch. If all 20 nulls are true and the tests are independent, the chance that a single test does not produce a false positive is `0.95`, so the chance that none of the 20 does is `0.95^20 = 0.358`. The chance of at least one false positive is therefore `1 - 0.358 = 0.64`, about 64 percent. That quantity — the probability of at least one type I error anywhere in the set — is the family-wise error rate, and it grows toward 1 as you add tests. This is why a significant result plucked from twenty colours of jelly bean means little: the headline reports the one test that crossed the line and hides the nineteen that did not. The fix is to control error over the whole family rather than per test.

go deeper

for a junior

Recall that the 5 percent error rate is a promise about one test, not about a batch of them, and be able to say out loud why twenty tests are far riskier than one.

for a middle

Be ready to do the arithmetic on the spot: one minus 0.95 to the twentieth power, about 64 percent, and to name the assumption of independence that the product step requires.

for a senior

Expect a follow-up about where this bites in your own work - a forty-metric dashboard, a feature screen - and be ready to say which family you would define and what correction you would apply.

for a principal

Own the policy: decide what error rate the organisation is buying across a whole analysis programme, and weigh the cost of one false claim against the effects a strict correction will hide.

## The quantity being asked for A hypothesis test is set up so that, when the null hypothesis is true, the probability of wrongly rejecting it is at most alpha — conventionally 0.05. That is a per-test guarantee. It says nothing about what happens when you run the same procedure many times over. Suppose all 20 nulls are true (nothing real is going on anywhere) and the 20 tests are statistically independent. For a single test: - P(false positive) = 0.05 - P(no false positive) = 0.95 Because the tests are independent, the probability that *none* of the 20 produces a false positive is the product: ``` P(no false positive anywhere) = 0.95^20 = 0.3585 ``` So the probability of at least one false positive is the complement: ``` P(at least one) = 1 - 0.95^20 = 0.6415 ``` Roughly 64 percent. Running twenty null tests at the 5 percent level, you are more likely than not to end up with something that looks significant. ## Family-wise error rate The number just computed has a name: the **family-wise error rate (FWER)**, defined as the probability of making one or more type I errors across a specified family of tests. Contrast it with the per-comparison error rate, which is just alpha for each test taken alone. The whole multiple-comparisons problem is the gap between those two numbers: alpha controls the second, and the reader of your report almost always cares about the first. The general formula, for m independent tests all under true nulls, is: ``` FWER = 1 - (1 - alpha)^m ``` A few values at alpha = 0.05: m = 5 gives 23 percent, m = 10 gives 40 percent, m = 20 gives 64 percent, m = 100 gives 99.4 percent. The growth is fast at first and then saturates. By a hundred tests, at least one spurious hit is essentially guaranteed. ## Why the naive additive answer is wrong A common wrong answer is 20 x 0.05 = 1.00, i.e. certainty. Adding the alphas is the union bound, and the union bound is an upper bound, not an equality — it double-counts the scenarios where two or more tests fire at once. It happens to be a *useful* bound (the Bonferroni correction is built on exactly this inequality), but as an estimate of the actual FWER it overshoots, and it becomes absurd past m = 20 where it exceeds 1. The other common wrong answer is 5 percent — the belief that alpha somehow covers the whole analysis. It does not. Alpha is attached to one decision rule applied once. ## What independence buys, and what happens without it The 0.95^20 step needs independence. Real batches of tests are usually correlated: twenty metrics on the same users, twenty features that move together, twenty endpoints measured on the same patients. Positive correlation makes the tests fire together, which *reduces* the FWER relative to the independent case — in the extreme where all twenty tests are perfect duplicates of each other, the FWER is just 5 percent, because there is really only one test. So 64 percent is the independent-case figure, and correlated families sit somewhere between alpha and that number. What never happens is the FWER staying at 5 percent for genuinely distinct tests. Also note the conditioning: the formula assumes all 20 nulls are true. If only 5 of the 20 nulls are true and the other 15 effects are real, then only those 5 can produce a type I error, and the family-wise error rate is `1 - 0.95^5 = 0.23`. False positives can only come from true nulls. ## Why this matters in practice The scenario is not artificial. A dashboard with forty metrics, a screen of a thousand candidate features, a questionnaire with thirty items, a panel of gene-expression measurements — each is a family of tests being read as if each significant result stood alone. The canonical cartoon version is the lab that tests twenty colours of jelly bean against acne, finds that green beans are significant at p < 0.05, and puts *green jelly beans linked to acne* on the front page. Nineteen null results never appear. The response is to control error at the family level rather than the test level: shrink the per-test threshold so the family-wise rate lands back at 5 percent, or switch to a different, more permissive promise about the *proportion* of your discoveries that are false. Which of those you choose depends on how many tests you are running and how expensive a single false claim is. In an interview, the expected answer is the arithmetic (`1 - 0.95^20`, about 64 percent), the name of the quantity (family-wise error rate), and the recognition that alpha is a per-test guarantee that says nothing about a batch.

  • What changes if only 5 of the 20 null hypotheses are actually true?
    Only true nulls can generate type I errors, so the family-wise error rate is computed over those 5 tests: 1 - 0.95^5 = 0.23, about 23 percent. The 15 tests with real effects can produce type II errors instead, but never false positives. This is why the family-wise error rate is usually quoted for the worst case where every null is true.
  • How does correlation between the 20 tests change the 64 percent figure?
    Positive correlation pushes it down. Correlated tests tend to fire together, so the false positives pile onto the same scenarios rather than spreading out. In the extreme of twenty identical tests the family-wise rate is just 5 percent. So 64 percent is the independent-case value and correlated families fall between alpha and that. Corrections built on the union bound stay valid either way, just more conservative.
  • Why is adding the alphas, 20 x 0.05 = 1.00, the wrong calculation?
    That is the union bound, which is an upper bound rather than the exact probability. It double-counts scenarios where two or more tests are significant at once, so it overshoots, and past twenty tests it exceeds 1 and stops being meaningful. The exact independent-case answer uses the complement: 1 - 0.95^20 = 0.64.

A 5 percent chance of rain on any given day sounds safe; over twenty days, rain somewhere in the stretch is more likely than not.

saying these in an interview costs you the question

  • Says the overall error rate is still 5 percent
  • Adds alphas: 20 times 0.05 equals certainty
  • Treats each significant result as independently trustworthy
  • Thinks correction only matters for correlated tests
  • Forgets false positives arise only from true nulls

context