skip to content

How does the Holm-Bonferroni step-down procedure improve on plain Bonferroni?

level: middleimportance: nice to knowfreq 31%

answer

  1. sort ascending, walk downward
  2. denominator shrinks as you go
  3. first failure ends the walk
  4. same guarantee, never fewer rejections
  5. first threshold equals alpha over m

basics

~20 s

Holm sorts the p-values smallest first and compares the i-th to alpha/(m - i + 1), stopping at the first failure. Thresholds loosen as you go, so Holm rejects at least as much as Bonferroni under the same guarantee.

solid answer

~50 s

Plain Bonferroni compares every p-value to the same fixed `alpha/m`. Holm sorts the p-values ascending and walks down the list, comparing the smallest to `alpha/m`, the second to `alpha/(m-1)`, the i-th to `alpha/(m - i + 1)`, and stopping at the first p-value that fails its threshold; everything from there on is retained. Take alpha = 0.05 and four p-values 0.010, 0.013, 0.030, 0.400. Bonferroni's fixed cut is 0.0125, so only 0.010 is rejected. Holm compares 0.010 to 0.0125 (reject), 0.013 to 0.05/3 = 0.0167 (reject), then 0.030 to 0.025 (fail, stop) — two rejections instead of one. The logic is that once a hypothesis is rejected there are fewer remaining hypotheses to protect, so the surviving budget can be spread over fewer tests. Holm controls family-wise error under any dependence, exactly like Bonferroni, and is uniformly at least as powerful — it never rejects less.

go deeper

for a junior

Recall that Holm is a sorted, sequential version of Bonferroni that ends up rejecting at least as many hypotheses, and that it is generally the better default of the two.

for a middle

Be ready to run the thresholds by hand on a short sorted list, name the divisor m minus i plus one, and state that the walk halts permanently at the first failure.

for a senior

Show you know what the procedure does and does not buy: identical family-wise protection, a free power gain, but no rescue at all once the number of tests reaches the thousands.

for a principal

Be ready to explain why an assumption-free step-down is easy to defend to a reviewer or regulator, and when you would instead argue for changing the error criterion entirely.

## The procedure, step by step Holm-Bonferroni is a **step-down** procedure: it starts at the most significant result and works toward the least, and it stops the first time a test fails. 1. Sort the m p-values ascending: `p(1) <= p(2) <= ... <= p(m)`. 2. Compare `p(1)` to `alpha / m`. If it fails, reject nothing and stop. 3. Otherwise reject that hypothesis and compare `p(2)` to `alpha / (m - 1)`. 4. In general compare `p(i)` to `alpha / (m - i + 1)`. 5. At the first index i where `p(i) > alpha / (m - i + 1)`, stop: retain that hypothesis and every one after it. The first threshold is identical to Bonferroni's. Every subsequent threshold is strictly larger. Nothing after the stopping point can be rejected even if its p-value happens to be small — the stopping rule is what preserves the guarantee, and skipping past a failure to grab a later small p-value breaks it. ## A worked example Four tests, alpha = 0.05, p-values 0.010, 0.013, 0.030, 0.400. - Bonferroni: fixed threshold `0.05/4 = 0.0125`. Rejections: 0.010 only. One discovery. - Holm: - `p(1) = 0.010` vs `0.05/4 = 0.0125` -> reject. - `p(2) = 0.013` vs `0.05/3 = 0.0167` -> reject. - `p(3) = 0.030` vs `0.05/2 = 0.025` -> fails, stop. - `p(4) = 0.400` retained automatically. Two discoveries. Same data, same error guarantee, one extra finding. That is the entire pitch. ## Why it still controls family-wise error The intuition: at the point where you are evaluating the i-th smallest p-value, you have already rejected i-1 hypotheses. In the worst case for the error rate — the case that matters, where the remaining hypotheses include all the true nulls you might falsely reject — there are at most `m - i + 1` of them left. Bonferroni's argument applied to that shrinking set justifies a threshold of `alpha / (m - i + 1)` rather than `alpha / m`. A careful version of that argument shows the family-wise error rate stays at or below alpha for **any** dependence structure among the tests, exactly the same assumption-free footing as plain Bonferroni. Holm's procedure is a closed testing argument at heart, and it controls error in the strong sense: the bound holds regardless of how many nulls are actually true. ## Uniformly more powerful Because Holm's first threshold equals Bonferroni's and every later one is larger, the set Holm rejects always contains the set Bonferroni rejects. It never rejects fewer hypotheses, and it often rejects more. Domination is uniform — no data set exists where plain Bonferroni wins. That means the choice between them is not a trade-off at all; plain Bonferroni survives mostly because it is a single division that a reader can verify in their head, and because with one test-of-interest the two coincide. The size of the gain depends on the shape of the p-value list. If only one p-value is small and the rest are large, both procedures reject the same single hypothesis and Holm adds nothing. If several p-values cluster just above `alpha/m`, Holm can pick up several extra rejections. ## Where it sits relative to other corrections Holm makes the same promise as Bonferroni: with 95 percent probability, **no** false rejection anywhere in the family. That is a strict promise, and its power still degrades badly as m grows — with thousands of tests, even Holm's loosest threshold near the top of the list is `alpha/m` for the very smallest p-value, so almost nothing clears the bar. When m is large the answer is usually to change the promise rather than to squeeze the procedure: control the expected *proportion* of false discoveries instead of the probability of any. There is also a step-**up** relative, Hochberg's procedure, which walks the sorted list from the largest p-value downward using the same thresholds and rejects everything from the first success onward. It is more powerful than Holm but needs a positive-dependence condition, whereas Holm needs nothing. ## The interview answer State the sorted thresholds `alpha/(m - i + 1)`, state the stopping rule (halt at the first failure, retain everything after), state that the guarantee is identical to Bonferroni's and equally assumption-free, and state that Holm uniformly dominates on power. A worked three- or four-p-value example, computed out loud, is what makes the answer land.

  • What happens if a p-value after the stopping point is smaller than its own threshold?
    You still cannot reject it. The stopping rule is part of the procedure, not a shortcut: once a hypothesis fails its threshold, it and everything after it are retained. Skipping past the failure to collect a later rejection invalidates the family-wise guarantee, because the argument for the loosened thresholds depends on every earlier hypothesis having been rejected.
  • Is there ever a case where plain Bonferroni rejects more than Holm?
    No. Holm's first threshold is exactly alpha/m, the same as Bonferroni's, and every later threshold is strictly larger, so Holm's rejection set always contains Bonferroni's. The domination is uniform across every possible data set, which is why Holm is preferred whenever anyone bothers to implement the sort.
  • Does Holm rescue power when m is in the thousands?
    Not meaningfully. The smallest p-value still faces alpha/m, so with thousands of tests the bar at the top of the list is brutal and only a handful of extremely small p-values can start the chain. When m is that large the fix is to change what you control - bound the proportion of false discoveries rather than the probability of any - not to switch step-down procedures.

Clearing a queue with a fixed error budget: every hypothesis you rule out leaves fewer left to protect, so the budget per remaining candidate goes up.

saying these in an interview costs you the question

  • Uses one fixed threshold for the sorted list
  • Continues past the first failing p-value
  • Sorts p-values descending for a step-down
  • Claims Holm needs independent tests
  • Thinks Holm controls false-discovery rate

context