Why is the sum of squared deviations divided by n a biased estimator of the population variance?
answer
- the centre was fitted to this same data
- which value minimises the squared deviations
- compare deviations around xbar and around mu
- expected sum of squares is (n-1) sigma squared
- the deviations sum to zero, one constraint
basics
~20 sDeviations are measured from the sample mean, which is pulled toward the data and makes those deviations as small as they can possibly be. Dividing their squares by n therefore underestimates the population variance by the factor (n-1)/n.
solid answer
~50 sThe squared deviations are taken around the sample mean `xbar`, not the unknown population mean `mu`. Among all possible centres, `xbar` is precisely the value that minimises the sum of squared deviations, so the sum computed around it is systematically smaller than the sum around `mu` would have been. The exact algebra gives `sum (x_i - xbar)^2 = sum (x_i - mu)^2 - n*(xbar - mu)^2`, and taking expectations yields `E[sum (x_i - xbar)^2] = (n-1)*sigma^2`. Dividing by n therefore produces an estimator whose expected value is `(n-1)/n * sigma^2` — too small on average, badly so for small n. Dividing by n-1 instead cancels the factor exactly and gives an unbiased estimator. The same n-1 is the degrees of freedom: the n deviations from the sample mean must sum to zero, so only n-1 of them are free.
go deeper
Be ready to say that deviations are measured from the sample mean rather than the unknown population mean, that this makes them too small, and that n-1 corrects the shortfall exactly.
Expect to derive it: add and subtract mu inside the square, kill the cross term, and show the expected sum of squares is (n-1) times sigma squared. Then connect n-1 to the single constraint on the deviations.
Demonstrate you know when the correction stops mattering and when it bites — small-sample variance estimates feeding intervals and test statistics — and that unbiasedness of the variance does not carry over to the standard deviation.
Own the general principle: one degree of freedom is spent per parameter fitted from the same data before residual spread is measured. Be able to say where that reasoning reappears across the modelling stack.
## The claim Take a random sample `x_1, ..., x_n` from a population with mean `mu` and variance `sigma^2`. Form the sum of squared deviations around the sample mean: ``` SS = sum (x_i - xbar)^2 ``` The result to know is ``` E[SS] = (n - 1) * sigma^2 ``` So `SS / n` has expected value `(n-1)/n * sigma^2`, which is smaller than `sigma^2` — a downward bias — while `SS / (n-1)` has expected value exactly `sigma^2`. ## Why the bias points downward The intuitive argument is a minimisation argument, and it is the one to say out loud first. Consider the function `g(a) = sum (x_i - a)^2`. Differentiating and setting to zero gives `a = xbar`: the sample mean is the unique value that makes the sum of squared deviations as small as possible. The population mean `mu` is some other number (it coincides with `xbar` only by fluke), so ``` sum (x_i - xbar)^2 <= sum (x_i - mu)^2 ``` always, for every sample. The quantity we can compute is guaranteed to be no larger than the quantity we wish we could compute. Averaging over samples, it is strictly smaller. The sample mean has been fitted to this very data, and the data therefore look less spread out around it than they really are around the truth. ## The algebra Add and subtract `mu` inside the square: ``` sum (x_i - mu)^2 = sum ((x_i - xbar) + (xbar - mu))^2 = sum (x_i - xbar)^2 + 2*(xbar - mu)*sum (x_i - xbar) + n*(xbar - mu)^2 ``` The cross term vanishes because `sum (x_i - xbar) = 0` by construction. Rearranged: ``` sum (x_i - xbar)^2 = sum (x_i - mu)^2 - n*(xbar - mu)^2 ``` Now take expectations term by term. Each `E[(x_i - mu)^2] = sigma^2`, giving `n * sigma^2` for the first piece. The second piece is `n` times the variance of the sample mean, which is `sigma^2 / n`, giving `sigma^2`. Hence ``` E[SS] = n*sigma^2 - sigma^2 = (n - 1)*sigma^2 ``` One unit of `sigma^2` was consumed by the act of estimating the centre from the same data. ## Degrees of freedom, said properly The n deviations `x_i - xbar` are not n free numbers: they satisfy one linear constraint, `sum (x_i - xbar) = 0`. Tell me any n-1 of them and the last is determined. So they carry n-1 independent pieces of information about spread, and dividing the sum of their squares by n-1 is dividing by the amount of information actually present. This is the general rule — one degree of freedom is spent per parameter estimated from the data before measuring residual spread — and it is why the correction generalises far beyond this one formula. ## How large is the effect The bias factor `(n-1)/n` is 0.5 at n = 2, 0.9 at n = 10, 0.99 at n = 100 and 0.999 at n = 1000. So the correction is decisive for tiny samples and cosmetic for large ones. Note also that the bias shrinks to zero as n grows: the n-divisor version is biased at every finite sample size yet still converges to `sigma^2` — biased but consistent. ## Three traps **"n-1 makes it bigger, and bigger is safer."** Wrong reasoning even though the arithmetic direction is right. The correction is not a safety margin; it is the exact factor that makes the expected value land on `sigma^2`. **"So the sample standard deviation is unbiased too."** It is not. The n-1 divisor makes the *variance* estimator unbiased; taking a square root is a non-linear operation, and the resulting standard deviation is biased slightly low. Unbiasedness does not pass through square roots. **"You always divide by n-1."** Only when the centre was estimated from the same data. If the population mean `mu` were genuinely known, `sum (x_i - mu)^2 / n` is already unbiased for `sigma^2` — nothing was fitted, so nothing was spent. That contrast is the cleanest way to show an interviewer you understand the mechanism rather than the ritual. ## Saying it in an interview The compact answer: the sample mean is fitted to the data and sits at the point that minimises squared deviations, so deviations around it understate the true spread; the expected sum of squares is `(n-1)*sigma^2`, and dividing by n-1 rather than n corrects that exactly. If pressed, run the add-and-subtract-mu algebra, then connect n-1 to the single constraint the deviations obey.
- If the population mean were genuinely known in advance, would you still divide by n-1?No. With `mu` known, `sum (x_i - mu)^2 / n` is already unbiased for `sigma^2`, because each term has expected value `sigma^2` and nothing was estimated from the sample. The n-1 correction exists solely to pay for fitting the centre from the same data. That contrast shows the divisor is a consequence of a mechanism, not a fixed ritual.
- How much does the n divisor actually matter in practice?The n-divisor estimate is too small by the factor (n-1)/n: 50% too small at n = 2, 10% at n = 10, 1% at n = 100. For large samples the choice is cosmetic and either divisor is defensible. It matters when n is small, and small-n variance estimates are exactly the ones that feed into intervals and test statistics, where an understated spread produces overconfident conclusions.
- Does the n-1 correction depend on the data being normally distributed?No. The derivation uses only linearity of expectation, the definition of variance and the fact that the sample mean has variance `sigma^2 / n`. It holds for any population with a finite variance, whatever its shape. Normality is needed for the *distribution* of the resulting statistic, not for its unbiasedness.
Measuring how far a class sits from the seat you chose after seeing where they sat makes them look tightly clustered. You picked the seat to make the distances small, so the number understates how spread out they really are.
saying these in an interview costs you the question
- Says n-1 is just a convention with no derivation
- Claims the correction only applies to normal populations
- Thinks n-1 exists to make estimates conservatively larger
- Asserts the sample standard deviation is therefore unbiased
- Uses n-1 even when the population mean is known
- Cannot say which direction the n divisor errs in