skip to content

How is the F statistic in a one-way ANOVA built, and what are its two degrees of freedom?

level: middleimportance: must knowfreq 68%

answer

  1. signal over noise
  2. two mean squares, one ratio
  3. numerator counts groups
  4. denominator is observations minus groups
  5. the two add up to N-1

basics

~20 s

F is the between-group mean square divided by the within-group mean square. With k groups and N observations the numerator has k-1 degrees of freedom and the denominator N-k. F sits near 1 when all group means are equal.

solid answer

~40 s

One-way ANOVA splits the total sum of squares into a between-group part and a within-group part: `SS_total = SS_between + SS_within`. Each becomes a mean square by dividing by its degrees of freedom — `MS_between = SS_between / (k-1)` and `MS_within = SS_within / (N-k)`, where k is the number of groups and N the total sample size — and the statistic is `F = MS_between / MS_within`. Both mean squares are estimates of the same error variance when the null of equal means holds, so F lands near 1; a real group effect inflates only the numerator, pushing F up. The reference distribution is F with (k-1, N-k) degrees of freedom, and the test uses only the upper tail. Note the degrees of freedom add up: (k-1) + (N-k) = N-1, the total.

go deeper

for a junior

Recall the shape of the statistic: between-group variability on top, within-group variability underneath, and a value near 1 when nothing is going on. Knowing k-1 and N-k by name is enough at this level.

for a middle

Derive it out loud: the sum-of-squares identity, dividing each by its degrees of freedom, and why both mean squares estimate the same error variance under the null. Expect to be asked for the df of a concrete design.

for a senior

Reason about what moves F in a real study — effect size, replication per group, measurement noise, and how many empty groups you dilute the numerator with — and use that to argue about power before data is collected.

for a principal

Own the design tradeoff: more levels of the factor versus more replication per level under a fixed budget. Extra levels spend numerator degrees of freedom and dilute the omnibus test, which is a strategic call, not an analysis detail.

## Partitioning the variability Start with N observations spread over k groups — say 40 plots of land, 10 under each of four fertilizers. Compute the **grand mean** over all 40 yields and each **group mean** over its own 10. Every observation's distance from the grand mean can be split in two: how far its group mean sits from the grand mean, plus how far the observation sits from its own group mean. Squaring and summing over all observations, the cross terms vanish and you get the exact identity: ``` SS_total = SS_between + SS_within ``` with - `SS_between = sum over groups of n_i * (group_mean_i - grand_mean)^2` — the **explained** part, larger when the group means are pulled apart. Note the weight `n_i`: a big group whose mean is off-centre contributes more. - `SS_within = sum over all observations of (value - its group mean)^2` — the **unexplained** or error part, the scatter that survives after each group is centred on its own mean. ## From sums of squares to mean squares Sums of squares are not comparable on their own — `SS_within` is built from far more terms than `SS_between`, so it is bigger almost by construction. Dividing each by its degrees of freedom turns them into variances on a common footing: - `MS_between = SS_between / (k-1)`. The numerator degrees of freedom are **k-1** because the k group means are free to vary around the grand mean, but once you know the grand mean and k-1 of them the last is determined. - `MS_within = SS_within / (N-k)`. The denominator degrees of freedom are **N-k** because one mean is estimated inside each of the k groups, spending one degree of freedom apiece out of N observations. The degrees of freedom partition just as the sums of squares do: `(k-1) + (N-k) = N-1`, which is the total degrees of freedom around the grand mean. In the four-fertilizer example with 40 plots, that is 3 and 36, summing to 39. ## The statistic and why it centres on 1 ``` F = MS_between / MS_within ``` `MS_within` is always an unbiased estimate of the common error variance sigma-squared, whether or not the means differ — centring each group on its own mean removes any group effect. `MS_between` estimates sigma-squared **only when the null of equal means is true**; otherwise its expectation is sigma-squared plus a positive term driven by how far apart the true means are. So under the null the ratio of two estimates of the same quantity hovers around 1, and under the alternative the numerator inflates while the denominator does not, driving F upward. (Strictly, the mean of an F distribution is `df2 / (df2 - 2)`, which is a little above 1 and approaches 1 as the error degrees of freedom grow. "About 1" is the right interview-grade statement.) ## Reading it as signal over noise F = 1.1 means the group means are no further apart than the within-group scatter would produce by chance. F = 9 with df (3, 36) means the between-group spread is nine times the noise level — far more separation than chance explains. Whether a given F is large enough is decided by the F(k-1, N-k) distribution, which is right-skewed and defined only on non-negative values. Only the **upper tail** is used: a large F is evidence against equal means, while a tiny F says the means are unusually close, which is not evidence for the alternative. ## What moves F - **Bigger true differences between the means** raise `SS_between` and so raise F. - **More observations per group at fixed means** also raise F, because `SS_between` carries the `n_i` weight while `MS_within` keeps estimating the same sigma-squared. This is just power: the same effect becomes detectable with more data. - **Noisier measurements** raise `MS_within` and shrink F. Reducing measurement noise or blocking out a nuisance source of variability is often a cheaper route to a detectable result than collecting more units. - **More groups with nothing going on** raise k-1 without raising `SS_between` proportionally, diluting F. ## Common slips worth avoiding Using `N-1` as the denominator degrees of freedom is the most frequent error — that is the *total*, not the error. Treating F > 1 as automatically significant is another: with df (3, 36) the 5% critical value sits well above 2.8, and small samples make the F distribution's right tail long. Finally, keep sums of squares and mean squares distinct; comparing raw `SS_between` to `SS_within` without dividing by degrees of freedom compares apples to oranges.

  • You hold the group means fixed and double the number of observations per group — what happens to F?
    F roughly doubles. `SS_between` weights each squared mean deviation by the group size, so it scales with n, while `MS_within` keeps estimating the same error variance regardless of sample size. The ratio therefore grows and the p-value falls. That is exactly why the same real effect can be non-significant in a small study and clearly significant in a larger one.
  • Why does the within-group mean square estimate the error variance whether or not the null is true?
    Because each observation is compared to its own group mean, not the grand mean. Centring inside the group removes any group effect entirely, so what remains is pure residual scatter. The between-group mean square has no such protection: it inherits a term proportional to the true spread of the means, which is what makes the ratio sensitive to the alternative.
  • How do the degrees of freedom decompose in a one-way ANOVA?
    Exactly as the sums of squares do: total N-1 splits into k-1 for between-groups and N-k for within-groups. Checking that (k-1) + (N-k) = N-1 is a quick sanity check on any ANOVA table you are handed, and a mismatch usually means the group count or the row count has been misread.

Several people speaking in a noisy room: the numerator is how far apart their voices are, the denominator is the background hiss. F is the ratio, and near 1 means you cannot tell them apart from the noise.

saying these in an interview costs you the question

  • Uses N-1 as the denominator degrees of freedom
  • Declares any F above 1 significant
  • Thinks larger within-group variance increases F
  • Compares sums of squares without dividing by degrees of freedom
  • Treats the F test as two-tailed

context