skip to content

How do you decide which set of tests counts as one family needing multiplicity correction?

level: principalimportance: should knowfreq 38%

answer

  1. no formula supplies m
  2. the scope of the claim sets the scope
  3. primary, secondary, exploratory tiers
  4. too broad fails as badly as too narrow
  5. fix it before results are known

basics

~20 s

Define the family as the set of tests over which you claim an error guarantee, fix it in the analysis plan before seeing results, and keep it as small as the claim allows. No formula supplies it.

solid answer

~50 s

There is no statistical rule that hands you the family — the correction is only as meaningful as the boundary you draw, and you draw it. The workable principle is that the family is the set of tests over which a reader will interpret your claim, so the guarantee has to span exactly that set. In practice that means naming a small number of confirmatory tests in advance, testing those under a family-wise or false-discovery criterion, and labelling everything else exploratory and uncorrected but explicitly flagged as such. A 15-endpoint trial with one designated primary endpoint tests that endpoint at the full 0.05; it does not divide by 15 for a claim it never intended to make. The failure modes are symmetric: too narrow and the correction is theatre, too broad and you report nothing. Whichever you pick, write down the family size, criterion and level, and let the result be judged against it.

go deeper

for a junior

Recall that the number of tests you correct for is something a human chooses, not something the maths supplies, and that it has to be decided before the results are in.

for a middle

Be ready to explain how m enters every procedure you know, and why a correction computed over a family picked after seeing the results provides no guarantee at all.

for a senior

Show the working structure: a small named primary family carrying the claims, an ordered secondary set, and everything else reported and labelled exploratory rather than silently dropped or falsely confirmed.

for a principal

Own it as policy - which criterion applies to which class of decision, how many primary claims an experiment may declare, and how the family, level and procedure get recorded so any reader can audit the call.

## Why the question has no formula answer Every multiplicity procedure takes m as an input. Bonferroni divides by m, Holm's denominators count down from m, the false-discovery staircase is `(i/m)q`. None of them tells you what m is. That choice is entirely yours, it changes every conclusion, and it is not checkable from the data. This is the part of multiplicity that is a leadership problem rather than a computation. The orienting principle: **the family is the set of tests over which you are asserting an error guarantee**. If you are going to say *this result is significant* and expect a reader to act on it, the guarantee must cover every test that could have produced that sentence. If a reader would be equally happy with any of fifteen endpoints coming back positive, then all fifteen are in the family, because any of them could have generated the headline. ## Both failure modes are real **Too narrow.** The family is drawn around the tests that happened to look good, or redrawn after the fact so the correction is survivable. Correcting three tests when twenty were run gives a guarantee that means nothing — the arithmetic is right and the claim is still false. The tell is that m was determined after the results were known. **Too broad.** The family is every test anyone ran on the dataset, or every test in the quarter. Now m is in the hundreds, per-test thresholds are microscopic, and nothing is ever significant. This is often presented as rigour, but a procedure that can never detect anything is not conservative — it has simply converted all of its type I error risk into type II error risk and stopped being informative. Deciding to over-correct is a decision to miss real effects, and it should be defended as such rather than assumed to be free. ## The structure that actually works The convention that survives contact with real studies is a **hierarchy of claims**, fixed before analysis: - **Primary.** One, occasionally two, tests that the study exists to answer. These carry the confirmatory claim and get the full error budget — a single primary endpoint is tested at 0.05, not at 0.05 divided by the number of things you also measured. - **Secondary, confirmatory.** A named, ordered set that may also support claims. These are corrected as a family, or tested in a fixed sequence where each is only examined if the previous one succeeded, which spends no extra alpha. - **Exploratory.** Everything else. Reported without correction *and* explicitly labelled exploratory, generating hypotheses rather than conclusions. The label is the honest alternative to both silence and false confirmation. This works because it makes the family small where the stakes are high and admits that the rest is not a claim at all. It also removes the incentive to game m: the boundary was set when nobody knew which way the results would fall. ## Questions worth asking when drawing the boundary - **What single sentence will appear in the summary?** Everything that could have produced that sentence belongs in the family. - **Who bears the cost of a false positive?** A false claim that triggers a costly programme, a regulatory filing, or a public statement argues for family-wise control on a small primary family. A false lead on a shortlist that goes to a cheap validation step argues for a false-discovery criterion on a large family. - **What is the cost of a miss?** A screening study that returns nothing has failed. A confirmatory study that returns a false positive has failed worse. The asymmetry decides the criterion. - **Are these tests answering one question or many?** Fifteen ways of measuring the same underlying effect is closer to one claim than fifteen. Fifteen unrelated endpoints is genuinely fifteen. - **Can I write m down before looking?** If not, the correction will not mean what it says. ## Making the decision auditable Whatever you choose, the analysis plan should carry four things: the family membership, its size m, the criterion (family-wise or false-discovery), and the level. Publish those alongside the results. A reader can then disagree with your boundary and re-derive their own conclusion, which is the whole point — the choice is a judgment, and judgments are defensible when they are visible and made in advance. An adjusted p-value with no stated family is uninterpretable, because the same number means different things depending on what m was. At organisational scale, the useful move is to make this a standing convention rather than a per-analyst decision: a house rule about how many primary endpoints an experiment may declare, which criterion applies to screening work versus decision work, and a requirement that the plan is written before data is examined. That converts a repeated argument into policy, and it is what a lead is expected to own.

  • An analyst corrects only the tests that came back significant. What is wrong with that?
    The correction is computed from a family selected using the results, so m is chosen after the fact and the guarantee is void. The family has to include every test that could have produced a claim, including all the ones that failed. Correcting only the winners always makes the correction survivable and always makes it meaningless.
  • Fifteen endpoints all measure the same underlying construct. Is that one family of fifteen tests?
    Statistically they are fifteen tests and any of them could produce the headline, so the family is fifteen. The better design fix is upstream: combine them into a single composite or index and test that once, or name one as primary. Heavy correlation also means a union-bound correction will be badly conservative, which is another argument for consolidating rather than dividing.
  • How would you set a house policy for multiplicity across many teams?
    Fix conventions rather than adjudicating case by case: a cap on declared primary endpoints, family-wise control for decision-grade claims and false-discovery control for screening work, a written plan with family size and level before data is examined, and a requirement that exploratory results are labelled as such wherever they appear. Then review the policy against outcomes rather than per-analysis.
  • When is deciding not to correct at all defensible?
    When there is a single pre-named test carrying the claim, or when the results are being reported explicitly as exploratory and no confirmatory claim is attached to any of them. Both are defensible because the guarantee being offered matches what is stated. What is not defensible is running many tests, correcting nothing, and presenting the survivors as confirmed.

saying these in an interview costs you the question

  • Believes a formula determines the family size
  • Chooses the family after seeing results
  • Corrects only the tests that reached significance
  • Treats maximal correction as automatically rigorous
  • Never states family size or criterion in the writeup
  • Divides alpha across endpoints nobody will claim

context