Your one-way ANOVA across four groups is significant — how do you find which groups differ?
answer
- omnibus says something, not what
- all pairs judged together
- reuses the pooled within-group mean square
- studentized range critical value
- report simultaneous intervals, not stars
basics
~20 sA significant omnibus F only says at least one mean differs. Follow it with a post-hoc procedure such as Tukey's HSD, which compares all pairs at once while holding the error rate across the whole family of comparisons.
solid answer
~50 sThe omnibus F is a gate, not an answer, so the next step is a post-hoc pairwise procedure — for all-pairs comparisons after a one-way ANOVA, Tukey's Honestly Significant Difference is the standard choice. It compares every pair of group means using the pooled within-group mean square from the ANOVA, but takes its critical value from the **studentized range** distribution rather than a t distribution. That distribution describes the largest gap you would expect among k sample means drawn from identical populations, so the threshold automatically accounts for the fact that you are inspecting all pairs. For a balanced design the critical difference is `q * sqrt(MS_within / n)`; with unequal group sizes the Tukey-Kramer form uses `q * sqrt((MS_within / 2) * (1/n_i + 1/n_j))`. I would report simultaneous confidence intervals for the pairwise differences rather than bare p-values, since they show direction and magnitude, not just a verdict.
go deeper
Know that a significant ANOVA is not the end of the analysis and that a named post-hoc procedure comes next. Being able to say 'Tukey's HSD for all pairwise comparisons' already puts you ahead.
Explain the mechanics: the pooled within-group mean square supplies the noise estimate, and the studentized range distribution supplies a critical value sized for looking at every pair at once.
Demonstrate the operating judgment — reporting simultaneous intervals rather than stars, handling a significant omnibus with no significant pair, and refusing to fish for pairs after a non-significant F.
Own the choice of comparison family before the study runs. All-pairs comparison spends error budget on contrasts nobody will act on; a pre-registered set of treatment-versus-control contrasts usually answers the business question with more power.
## Why a post-hoc stage exists at all A one-way ANOVA across four fertilizers returns F(3, 36) = 6.2, p = 0.002. That means "the fertilizer factor moves yield." It does not mean fertilizer C beat fertilizer A. Two temptations must be resisted: reading off the largest sample-mean gap and calling it the answer, and running ordinary two-sample tests on all six pairs. The first ignores that the largest of several noisy gaps is systematically larger than any single gap you would have picked in advance. The second reintroduces the compounding false-positive problem the omnibus test was there to avoid. ## What Tukey's HSD does Tukey's HSD examines **all pairwise differences** among k group means and holds the probability of *any* false pairwise claim at your chosen alpha across the entire set. Its two ingredients are: 1. **The pooled error estimate.** It reuses `MS_within` from the ANOVA table, the within-group mean square with `N-k` degrees of freedom, rather than estimating noise separately for each pair. Every comparison therefore rests on the full sample's information about residual variability. 2. **The studentized range critical value.** Instead of a t critical value, it uses `q(alpha, k, N-k)` from the studentized range distribution — the distribution of `(largest sample mean - smallest sample mean) / standard error` when all k populations are identical. That is exactly the right reference for "how big can the biggest gap get by luck alone when I look at all of them," which is why the procedure needs no separate adjustment layered on top. ## The critical difference For a **balanced** design with n observations per group: ``` HSD = q(alpha, k, N-k) * sqrt(MS_within / n) ``` Any pair of group means further apart than HSD is declared different. Equivalently, each pairwise difference gets a simultaneous confidence interval `(mean_i - mean_j) +/- HSD`, and a pair is significant when its interval excludes zero. For **unequal** group sizes the Tukey-Kramer modification replaces the common standard error with a pair-specific one: ``` critical difference for pair (i, j) = q(alpha, k, N-k) * sqrt((MS_within / 2) * (1/n_i + 1/n_j)) ``` This reduces to the balanced formula when `n_i = n_j = n`. With unequal sizes the procedure is slightly conservative — the realised error rate sits at or below alpha rather than exactly at it. ## Reporting it well Prefer the **simultaneous intervals** over a grid of stars. "Fertilizer C exceeds A by 0.8 t/ha, 95% simultaneous interval 0.2 to 1.4" tells a decision-maker the direction, the size and the uncertainty in one line; "p = 0.011" tells them only that something crossed a threshold. Because the intervals are simultaneous, the stated coverage applies to the whole set of statements at once, which is the property you want when the reader will scan every row and act on whichever looks best. ## The awkward cases **Significant omnibus, no significant pair.** This is entirely possible and worth being ready for. The F test aggregates evidence across the whole pattern of means and can detect a diffuse effect — several groups slightly high, several slightly low — that no single pairwise contrast is powerful enough to resolve once the comparison is adjusted for all pairs. It is not a computational error, and it is not something to fix by dropping the adjustment. Report it honestly: the factor matters, but the data cannot pin the difference to a specific pair. **Non-significant omnibus.** The conventional discipline is to stop. Fishing for a significant pair after an omnibus test failed is exactly the selection behaviour the two-stage structure prevents. **You did not want all pairs.** If the real question is "does each treatment beat the control?", all-pairs comparison is wasteful — it spends its error budget on treatment-versus-treatment comparisons nobody asked about, costing power on the ones that matter. Procedures aimed specifically at many-against-one comparisons exist, and a small set of **pre-planned** contrasts fixed before the data arrived is generally the strongest design. Deciding what to compare afterwards, based on the sample means, is the failure mode. ## The sentence to have ready "The F test is a gate; it says the factor matters without naming a group. I follow it with Tukey's HSD, which compares every pair using the pooled within-group mean square and a studentized-range critical value so the error rate holds across the whole family, and I report simultaneous intervals so the size of each difference is visible — not just which cells got a star."
- Why not just run ordinary two-sample t-tests on each pair after a significant ANOVA?Because those critical values are built for one comparison, not for the whole set, so the chance of at least one false pairwise claim compounds. Tukey's HSD draws its critical value from the studentized range distribution, which already describes how far apart the extreme means among k groups drift by chance, so the family-wise error rate stays at alpha.
- What changes in Tukey's procedure when the groups have unequal sizes?You use the Tukey-Kramer form, which replaces the common standard error with a pair-specific one: `q * sqrt((MS_within / 2) * (1/n_i + 1/n_j))`. It reduces to the balanced formula when the two sizes match, and it is mildly conservative under imbalance, meaning the realised error rate sits at or below the nominal alpha.
- Can the omnibus F be significant while no Tukey pairwise comparison is?Yes, and it is not an error. The F test pools evidence across the full pattern of means and can pick up a diffuse effect that no single adjusted pairwise contrast has the power to resolve. The correct response is to report it as it stands — the factor matters, but the data cannot attribute it to one specific pair — not to drop the adjustment until something turns significant.
saying these in an interview costs you the question
- Declares the largest raw mean gap the winner without any test
- Runs unadjusted pairwise tests after a significant ANOVA
- Assumes a significant F guarantees a significant pair
- Uses each pair's own variance instead of the pooled within-group mean square
- Chooses which pairs to test after inspecting the sample means
- Continues to post-hoc comparisons after a non-significant omnibus test