skip to content

What does a significant Kruskal-Wallis test tell you about three or more groups?

level: middleimportance: should knowfreq 38%

answer

  1. one pooled ranking across every group
  2. the rank analogue of a one-way ANOVA
  3. how far the mean ranks scatter
  4. chi-square with k-1 degrees of freedom
  5. omnibus only, pairwise comes after

basics

~20 s

Only that the groups are not all alike — at least one tends to produce larger values than another. It is an omnibus test on pooled ranks with k-1 degrees of freedom and never says which pair differs.

solid answer

~40 s

Kruskal-Wallis is the rank analogue of a one-way analysis of variance: pool all N observations from the k groups, rank them together, and measure how far each group's mean rank sits from the overall mean rank `(N+1)/2`. The statistic `H` is that weighted spread, referred to a chi-square distribution with `k-1` degrees of freedom for large samples, with a correction when ties are heavy. A significant H is an omnibus verdict — comparing three teaching methods, it says the three rank distributions are not interchangeable, not which method wins. You follow it with pairwise rank comparisons under a multiplicity correction, such as Dunn's test, to locate the difference. As with any rank comparison, calling the difference a shift in medians needs the groups to have similar shapes.

go deeper

for a junior

Recognise it as the rank-based way to compare three or more independent groups when the data are ordinal or skewed, and know that a significant result means the groups are not all the same.

for a middle

Explain the machinery: one pooled ranking, mean ranks compared with the grand mean rank (N+1)/2, chi-square with k-1 degrees of freedom, and a tie correction. State that it reduces to the two-group rank comparison when k = 2.

for a senior

Show discipline about what comes next: a planned post-hoc procedure with multiplicity control, an effect size, group medians for direction, and a check that the shapes justify any location claim you make.

for a principal

Own the design decision — whether an omnibus screen is the right first step at all, or whether the pre-specified comparisons that actually drive the decision should be tested directly with the error budget allocated up front.

## The setup You have k independent groups — say three teaching methods, each tried on a different set of students — and an outcome that is ordinal or badly skewed. You want to know whether the groups differ at all before you look at any particular pair. ## The statistic Pool all N observations across the k groups, rank them from 1 to N with midranks for ties, and let `Rbar_i` be the mean rank of group i, with n_i observations in it. The overall mean rank is `(N + 1)/2`. Then ``` H = [12 / (N(N + 1))] * sum over i of n_i * (Rbar_i - (N + 1)/2)^2 ``` which is algebraically the same as `H = [12/(N(N+1))] * sum(R_i^2 / n_i) - 3(N + 1)`, using rank sums `R_i`. Read the first form: it is a weighted measure of how far the group mean ranks scatter around the grand mean rank. All groups behaving alike puts the mean ranks near `(N+1)/2` and H near zero; one group monopolising the high ranks pushes H up. When ties are present, H is divided by `1 - (sum of (t^3 - t)) / (N^3 - N)`, where t runs over the sizes of the tied groups. This inflates H, compensating for the reduced variability ties introduce. ## The reference distribution For reasonably sized samples H is compared with a chi-square distribution on `k - 1` degrees of freedom — one fewer than the number of groups, exactly as in the numerator degrees of freedom of a one-way analysis of variance. The approximation degrades with very small groups, where exact or permutation-based null distributions are preferable. A useful sanity anchor: with k = 2 the Kruskal-Wallis test reduces to the two-sided Mann-Whitney U test, and the two give the same p-value. ## What significance does and does not say A significant H is an **omnibus** result. It says: the k samples do not all come from the same distribution; at least one group tends to produce larger values than at least one other. It does not say which pair, and it does not by itself rank the groups. To locate the difference you run pairwise rank comparisons afterwards, and because the number of pairs grows as k(k-1)/2, you must control the error rate across them — Dunn's test, which compares mean ranks using the pooled ranking with an adjustment for multiplicity, is the standard rank-based follow-up. Reporting the omnibus p-value plus one eye-catching pair with no correction is a classic weak answer. The interpretive caveat carried over from the two-group case applies here too. The null is that the k distributions are identical. Reading a rejection as "the medians differ" requires that the groups share a shape and spread; if one group is much more dispersed than the others, H can be significant even with equal medians, and the honest statement is that the distributions differ. ## Assumptions 1. Observations are independent within and across groups. Students taught in the same classroom are not independent, and clustered data needs a method that models the clustering rather than a plain rank test. 2. The outcome is at least ordinal. 3. Similar shapes across groups, only if you intend a location interpretation. Normality of the outcome is not required, and neither is equal variance for the general distributional null — but equal spread does matter once you want to talk about location. ## Effect size and reporting Report H, the degrees of freedom `k - 1`, the group sizes and mean ranks, and the p-value. `epsilon-squared`, computed as `H/((N^2 - 1)/(N + 1))`, gives the proportion of rank variability attributable to group membership and runs from 0 to 1; eta-squared style variants exist as well. Then report the post-hoc comparisons with the correction you used, and the group medians so a reader can see the direction. ## A common trap Candidates sometimes say the test "compares the medians of the groups". Kruskal-Wallis never computes a median. It compares mean *ranks*. The distinction matters both for the equal-shape caveat and for interpreting mean ranks — a mean rank is a position within the pooled sample and has no meaning outside the specific data set that produced it.

  • What do you do after a significant Kruskal-Wallis result to find which groups differ?
    Run pairwise rank comparisons with a correction for the number of pairs — Dunn's test is the standard rank-based follow-up, comparing mean ranks from the pooled ranking with an adjusted significance threshold. Report every comparison you ran, not just the one that survived, and include group medians so readers can see direction.
  • What does Kruskal-Wallis reduce to when there are only two groups?
    The two-sided Mann-Whitney U test. With k = 2 the statistic H is the square of the standardised rank-sum statistic and is referred to a chi-square distribution on one degree of freedom, giving the same p-value. It is a good consistency check that you have understood both tests.
  • Can Kruskal-Wallis be significant when all three group medians are equal?
    Yes. The null is that the distributions are identical, so a large difference in spread or shape can drive a rejection even with matching medians. That is why a location claim requires the groups to have similar shapes, and why you should report the distributions and not just a p-value.

saying these in an interview costs you the question

  • Says a significant result identifies which group is best
  • Claims the test compares group medians unconditionally
  • Runs uncorrected pairwise comparisons after the omnibus test
  • Uses k degrees of freedom instead of k-1
  • Applies it to repeated measures on the same subjects

context