skip to content

Why is the Bonferroni correction called conservative when testing 15 endpoints at once?

level: middleimportance: must knowfreq 68%

answer

  1. divide the budget, not the evidence
  2. Boole's inequality, an upper bound
  3. no dependence assumption needed
  4. loosest exactly when endpoints correlate
  5. type I risk traded for type II

basics

~20 s

Bonferroni tests each of the 15 endpoints at 0.05/15 = 0.0033 instead of 0.05. That keeps the family-wise error rate under 5 percent whatever the dependence between endpoints, but the tiny per-test threshold makes real effects far harder to detect.

solid answer

~50 s

Bonferroni divides the target family-wise error rate by the number of tests: with 15 endpoints and a 5 percent target, each endpoint is judged against `0.05 / 15 = 0.0033`. The justification is the union bound — the probability of at least one false positive is at most the sum of the individual probabilities — so the guarantee holds under *any* dependence structure among the endpoints, which is why it is so widely used and so easy to defend. The cost is power. A threshold of 0.0033 needs a much larger effect or a much larger sample to clear, so genuine effects on secondary endpoints routinely fail to reach significance. It is doubly conservative when the endpoints are correlated, because the union bound assumes the worst case where false positives never overlap. Holm-Bonferroni gives the same guarantee with strictly more power, and controlling false-discovery rate instead gives much more.

go deeper

for a junior

Recall the rule itself: divide your significance level by the number of tests, so 0.05 across 15 endpoints becomes a threshold of about 0.0033 for each one.

for a middle

Be ready to justify it with the union bound and to explain that the guarantee needs no assumption about dependence, plus what the shrunken threshold does to your ability to detect real effects.

for a senior

Show you weigh both error types: describe a study where Bonferroni left every endpoint underpowered, and what you changed - primary endpoint designation, a step-down procedure, or false-discovery control.

for a principal

Own the framing that over-correction is not free caution but a conversion of one error into another, and be able to argue which error your organisation should be buying down for a given decision.

## The procedure Bonferroni is the simplest multiplicity correction there is. Fix the family-wise error rate you want — call it alpha, usually 0.05 — count the tests in the family, call it m, and reject any hypothesis whose p-value satisfies: ``` p_i <= alpha / m ``` With a clinical trial carrying 15 endpoints and alpha = 0.05, that threshold is `0.05 / 15 = 0.00333`. Equivalently you can leave the threshold at 0.05 and multiply each p-value by 15 (capping at 1); the two framings are identical decisions. ## Why it works, for any dependence The proof is one line of Boole's inequality, also called the union bound. Let A_i be the event that test i falsely rejects a true null. Then: ``` P(A_1 or A_2 or ... or A_m) <= P(A_1) + P(A_2) + ... + P(A_m) ``` If each true null is tested at level alpha/m, each term on the right is at most alpha/m, and there are at most m of them, so the whole sum is at most alpha. The family-wise error rate is therefore at most alpha. The important property is what the argument does *not* require: nothing at all about how the tests relate to one another. No independence, no particular correlation structure, no distributional assumption beyond each individual test being valid. That robustness is why regulators and reviewers accept Bonferroni without argument, and why it is a reasonable default when you cannot characterise the dependence among your tests. ## Where the conservatism comes from Two separate sources. **The union bound is loose.** Equality in Boole's inequality holds only when the error events are mutually exclusive — when no two tests can ever falsely fire at the same time. Real families do not behave that way. For independent tests the true family-wise rate at threshold alpha/m is `1 - (1 - alpha/m)^m`, which is slightly *below* alpha; for positively correlated tests it can be far below, because false positives clump onto the same draws instead of spreading out. Fifteen endpoints measured on the same patients are typically strongly correlated, so a procedure calibrated for the worst case is spending error budget it does not need. **The worst case assumes every null is true.** The bound counts all m tests. If several of the 15 endpoints carry real effects, only the true nulls can generate a type I error, so the actual family-wise rate is bounded by a smaller multiple of alpha/m than the procedure assumed. ## The power cost, concretely Power is the probability of detecting an effect that is really there, and it depends on the threshold you must clear. Moving from 0.05 to 0.0033 pushes the critical value further into the tail: for a two-sided z-test the critical value moves from about 1.96 to about 2.94 standard errors. An effect that would have been detected 80 percent of the time at 0.05 is detected far less often at 0.0033, and recovering the lost power requires a substantially larger sample. In a 15-endpoint trial this is exactly why the protocol usually names one or two *primary* endpoints and treats the rest as exploratory: spreading a 5 percent budget across fifteen equally would leave every endpoint underpowered. The trade is stark and worth stating plainly: Bonferroni buys a strong, assumption-free guarantee against any false claim, and pays for it in missed real effects. Over-correcting is not free caution — it converts type I error risk into type II error risk, and a study that detects nothing is also a failed study. ## What to use instead, and when - **Holm-Bonferroni** delivers exactly the same family-wise guarantee under exactly the same (absent) assumptions, but with a step-down sequence of thresholds that is uniformly at least as powerful. There is essentially no reason to prefer plain Bonferroni over Holm on power grounds; plain Bonferroni survives because it is a one-line calculation and is trivially explained. - **Sidak's correction** uses `1 - (1 - alpha)^(1/m)` instead of alpha/m, which is very slightly larger and therefore very slightly more powerful, but requires independence. The gain over Bonferroni is negligible — for m = 15 the thresholds are 0.00341 and 0.00333 — so it is rarely worth the extra assumption. - **False-discovery-rate control** changes the promise entirely, from *no false claims at all* to *a bounded proportion of false claims among the ones you make*. That is the right move when m runs into the thousands and Bonferroni's threshold becomes unreachable. ## The interview answer Say the rule (alpha/m, so 0.0033 across 15 endpoints), say the justification (union bound, valid under any dependence), and say the price (a far stricter threshold, badly reduced power, worse still when endpoints are correlated because the bound assumes non-overlapping errors). Naming Holm as the strictly better version of the same guarantee is what separates a solid answer from a textbook recital.

  • When is the Bonferroni bound tight rather than conservative?
    Only when the false-positive events are mutually exclusive, meaning no two tests can ever falsely fire on the same data. That essentially never happens with real tests, and it is furthest from true when the tests are strongly positively correlated, because then errors clump together and the true family-wise rate falls far below the target.
  • How does Sidak's correction compare with Bonferroni?
    Sidak uses a per-test level of 1 - (1 - alpha)^(1/m), which is derived from the exact independent-case family-wise formula rather than the union bound. It is marginally larger, so marginally more powerful, but valid only under independence. At m = 15 and alpha = 0.05 it gives 0.00341 versus Bonferroni's 0.00333 - not worth the extra assumption.
  • A trial has 15 endpoints but only one that matters. Should you still divide by 15?
    No. Name that endpoint primary in the analysis plan and test it at the full 0.05, reporting the other 14 as exploratory and uncorrected but explicitly labelled as such. Correction is owed on the family over which you claim an error guarantee. Dividing by 15 when you only intend to make one confirmatory claim throws away power for nothing.

saying these in an interview costs you the question

  • Claims Bonferroni requires independent tests
  • Says correction is always the safe choice
  • Multiplies alpha by m instead of dividing
  • Ignores the power loss entirely
  • Prefers plain Bonferroni over Holm without reason
  • Counts only the significant tests in m

context