Why run one ANOVA across five groups instead of all 10 pairwise t-tests?
answer
- each test spends alpha again
- five groups make ten pairs
- chance of at least one false win
- one minus 0.95 to the tenth
- pooled error estimate is a bonus
basics
~10 sEach pairwise test spends its own 5% false-positive budget, so 10 of them carry roughly a 40% chance of at least one spurious significant result. A single ANOVA asks one question at one alpha.
solid answer
~50 sFive groups give `5 * 4 / 2 = 10` pairs. If each pairwise test runs at alpha = 0.05 and you treat them as independent, the chance of no false positive is `0.95^10 = 0.60`, so the chance of **at least one** spurious win is about 40% — eight times the 5% you thought you were spending. In reality the tests share data and are correlated, so the true rate is somewhat below 40%, but it is still far above 5%. A one-way ANOVA replaces the pile with a single omnibus test held at alpha. It is also more powerful per comparison, because it estimates the error variance by pooling all groups, giving N-k error degrees of freedom instead of the handful available to any one pair. If you still need pairwise verdicts afterwards, you run a post-hoc procedure that holds error across the whole family of comparisons.
go deeper
Recall that running many tests raises the chance of a false positive, and that five groups mean ten pairs. Being able to say the risk is far above 5% is enough at this level.
Do the arithmetic on the spot: 1 minus 0.95 to the tenth is about 0.40. Then explain the second benefit — ANOVA pools the error variance across all groups instead of two.
Show the two-stage discipline in practice: omnibus gate, then a family-controlled post-hoc stage. Flag the correlation caveat honestly without letting it soften the conclusion.
Own the design call: whether the analysis should be an omnibus sweep at all, or a short list of pre-registered contrasts against a control. Deciding the comparison set before data collection is the real safeguard.
## The arithmetic of error inflation Alpha is a per-test promise: run one test when the null is true, and you have a 5% chance of calling it significant anyway. That promise says nothing about a *collection* of tests. With k groups the number of distinct pairs is `k * (k-1) / 2`. For k = 5 that is 10 pairs. Suppose every null is true — all five groups are identical — and you test each pair at alpha = 0.05. Treating the tests as independent: ``` P(no false positive) = 0.95^10 = 0.5987 P(at least one false positive) = 1 - 0.5987 = 0.4013 ``` About **40%**. You believed you were running at 5% risk and you were actually closer to a coin flip. And the growth is fast: three groups give 3 pairs and about 14%; six groups give 15 pairs and about 54%. ## The independence caveat, and why it does not rescue you The 10 pairwise comparisons among five groups are **not** independent — each group appears in four of them, so the tests share data and their results are positively correlated. The exact probability of at least one false positive is therefore somewhat lower than 0.40. Mentioning this is a good sign in an interview, as long as you land the conclusion correctly: the correlation dampens the inflation, it does not remove it. The realised error rate is still multiples of the nominal 5%. ## What ANOVA does instead A one-way ANOVA asks a single question — "are all five means equal?" — and produces a single p-value evaluated at a single alpha. Whatever risk you set is the risk you take. That is the headline reason, but there is a second, less-quoted one: **A pooled error estimate.** Every pairwise test estimates the noise from only the two groups involved. ANOVA estimates it once from all five, so the denominator of the statistic rests on `N-k` degrees of freedom rather than the much smaller count a single pair provides. A better-estimated denominator means a tighter reference distribution and more power to detect a real effect — assuming the equal-variance assumption is reasonable, which is precisely what licenses the pooling. ## "Haven't you just moved the problem?" A sharp interviewer will push: if the ANOVA is significant, you still want to know which groups differ, and now you are back to pairwise comparisons. True — and that is why the pairwise stage after a significant omnibus test uses a **post-hoc procedure** whose critical values are built for the whole family of comparisons at once, rather than a stack of ordinary two-sample tests. Tukey's HSD is the standard choice for "all pairs" after a one-way ANOVA. The two-stage structure is deliberate: one gate, then a controlled set of follow-ups. ## When pairwise comparisons are actually the right call The omnibus test is not always the question you care about. If the study was designed around a small set of **pre-planned comparisons** — each of three treatments against a single control, three comparisons rather than six — then testing those directly, with an appropriate adjustment for the three, is more powerful and more honest than sweeping across all pairs. Plan the comparisons before seeing the data; choosing them afterwards, based on which sample means look furthest apart, silently reintroduces exactly the selection effect the whole discipline exists to prevent. ## The sentence to have ready "Ten pairwise tests at 5% each carry about a 40% chance of at least one false positive, so the error rate I quote stops being the error rate I have. One ANOVA answers the omnibus question at the alpha I chose, pools a better variance estimate across all groups, and if it fires I follow it with a post-hoc procedure that keeps the pairwise comparisons controlled as a family."
- Is the roughly 40% figure exact for 10 pairwise comparisons among five groups?No. `1 - 0.95^10 = 0.40` assumes the 10 tests are independent, but each group appears in four of them, so the comparisons share data and are positively correlated. The true probability of at least one false positive is somewhat lower than 40%. It is still many times the nominal 5%, so the conclusion is unchanged.
- Besides error control, what makes the ANOVA F test more powerful than an isolated pairwise test?It pools the error variance across every group instead of estimating it from just the two being compared. That gives the denominator N-k degrees of freedom, so the reference distribution has lighter tails and a given mean difference clears the bar more easily. The pooling is only justified when the groups have comparable variance.
- How many pairwise comparisons are there among k groups, and how fast does the problem grow?`k * (k-1) / 2`, which grows quadratically. Three groups give 3 comparisons, five give 10, six give 15, and ten groups give 45. At 45 independent tests with alpha = 0.05, the chance of at least one false positive is above 90% — essentially guaranteed.
saying these in an interview costs you the question
- Says 10 tests at 5% each still give 5% overall risk
- Adds the alphas to get exactly 50%
- Thinks a significant ANOVA removes the need for post-hoc control
- Picks the pairs to test after seeing which means look furthest apart
- Believes the correlation between pairwise tests eliminates the inflation