When computing a variance, when do you divide by n and when by n-1?
answer
- ask who the number is meant to describe
- whole group, or a draw from one
- the divisor follows that answer
- smaller divisor means a larger result
basics
~20 sDivide by N when your numbers are the whole population you describe. Divide by n-1 when they are a sample standing in for a larger population's spread. The n-1 version is always the larger of the two.
solid answer
~50 sThe divisor follows from what the rows represent, not from taste. If the data is the entire group you are describing — every employee, every transaction in the closed month — use `sigma^2 = (1/N) * sum (x_i - mu)^2`. If the rows are a subset standing in for something bigger, use `s^2 = (1/(n-1)) * sum (x_i - xbar)^2`. The deviations are measured from the sample's own mean, which is exactly the point that makes the sum of squares as small as possible for those numbers, so the n divisor runs systematically small as a description of the wider population; the smaller divisor lifts it. The two answers differ by a factor of `n/(n-1)`, which is 25% on the variance at n = 5 and under 4% at n = 30. Tools expose both under similar names, so always know which one yours returned.
go deeper
Memorise the pairing and be able to say it cleanly: whole population means divide by N, sample standing in for a population means divide by n-1. Practise computing both on five or six numbers by hand.
Explain the mechanism: deviations are taken from the sample's own mean, which minimises the sum of squares, so the n divisor understates the wider spread. Know the n/(n-1) ratio and that the corrected value is always larger.
Be able to reconcile two conflicting numbers on the same column and rank the divisor against filtering and missing-value differences as candidate causes. Know when the difference is too small to be worth a sentence.
Set the convention so results are comparable across teams and time: which form the standard reporting path uses, when sample size must be published alongside a spread, and how a small-n figure is caveated before an executive reads it.
## The two formulas and the question each answers Population form, for N values with population mean `mu`: `sigma^2 = (1/N) * sum (x_i - mu)^2` Sample form, for n values with sample mean `xbar`: `s^2 = (1/(n-1)) * sum (x_i - xbar)^2` Use the population form when the rows in front of you *are* the group you intend to describe and you are making no claim beyond them: every employee on the payroll, every order placed in a closed month, every pixel in one image. Use the sample form when the rows are a draw from something larger and the number is supposed to characterise that larger thing: 400 surveyed customers standing in for the customer base, an hour of traffic standing in for the service's behaviour. The honest way to decide is to finish the sentence "this variance describes the spread of ___". If the blank is the rows themselves, divide by N. If the blank is a wider population the rows were drawn from, divide by n-1. ## Same data, two answers Take the eight values 2, 4, 4, 4, 5, 5, 7, 9. Their mean is 5. Deviations: -3, -1, -1, -1, 0, 0, 2, 4. Squares: 9, 1, 1, 1, 0, 0, 4, 16. Sum of squared deviations: 32. - divide by n = 8: variance 4, standard deviation 2 - divide by n - 1 = 7: variance about 4.571, standard deviation about 2.138 Same eight numbers, two defensible answers, about 7% apart on the standard deviation. This is the scenario behind the interview question: someone reports 2, someone else reports 2.14, and the reconciliation is the divisor. ## Why the n-1 answer is always larger Algebraically, `s^2 = (n/(n-1)) * (sum of squared deviations / n)`, and `n/(n-1) > 1` for every n > 1, so the sample form never returns the smaller number. The intuition is that the deviations are measured from the sample's own mean, and the sample mean is precisely the value that minimises the sum of squared deviations for those particular numbers. Measured against the unknown population mean, the sum would come out larger. So dividing by n produces a figure that runs low as a description of the wider population, and shrinking the divisor lifts it back up. The formal treatment of why n-1 is exactly the right correction is a separate topic; what matters when computing a variance is knowing which situation calls for which formula and which direction the gap runs. ## When the difference matters The correction factor on the variance is `n/(n-1)`: - n = 5: 1.25 on the variance, so about 11.8% on the standard deviation - n = 10: about 1.11, roughly 5.4% on the standard deviation - n = 30: about 1.034, under 2% on the standard deviation - n = 1000: about 1.001, entirely negligible So this is a small-sample argument. At n = 8 in a lab notebook it changes the reported number visibly; at n = 50,000 in a warehouse table it is invisible. Two edge cases are worth naming: at n = 1 the sample variance is undefined, since the sum of squared deviations is zero and so is the divisor, while the population form would report zero spread — an honest reflection of the fact that a single observation says nothing about scatter. ## Know what your tooling returned Spreadsheets and statistical environments expose both forms, often under near-identical names, and their defaults are not consistent with each other. When two analysts report different standard deviations for the same column, the divisor is the first thing to check, ahead of filtering differences and missing-value handling. On small samples, state in the write-up which form you used; on large ones, say nothing, because the difference is below the precision anyone is reading. ## One caution about the standard deviation The n-1 correction is a correction on the variance scale. Taking the square root of the corrected variance does not carry it through cleanly, because the square root is a nonlinear step, so the resulting standard deviation still tends to run slightly low as a description of the population's spread. Everyone reports it anyway, and at any reasonable sample size the residual effect is negligible — but knowing that the fix lives on the variance and not on its root is the kind of detail that separates a memorised rule from an understood one. ## Common mistakes Saying "always n-1, that's what statistics does" is a memorised rule that collapses the moment the interviewer describes a full census. Claiming n-1 exists to compensate for outliers is wrong: it is about the gap between a sample and the population it represents, and it does nothing about extreme values. Saying the n-1 answer is smaller inverts the inequality.
- At what sample size does the choice of divisor stop mattering in practice?The correction factor on the variance is `n/(n-1)`. At n = 5 that is 1.25, a visible 12% change in the standard deviation. By n = 30 it is 1.034, under 2% on the standard deviation, and by n in the thousands it is well below the precision anyone reads off a dashboard. So it is a small-sample concern: worth stating explicitly in a lab-scale write-up, not worth a sentence on a warehouse-scale table.
- Two analysts report different standard deviations for the same column — how do you reconcile it?Check the divisor first, because it is the cheapest explanation and the tools disagree on defaults. If the ratio between the two numbers is close to `sqrt(n/(n-1))`, that is the whole story. If it is not, move on to the real suspects: different row filters, different handling of missing values, one side deduplicating and the other not, or the two working on different time windows.
- What is the sample variance of a single observation?It is undefined. The sum of squared deviations from the sample's own mean is zero, and the divisor n-1 is also zero, so the formula gives 0/0. That is the right behaviour: one observation carries no information about spread. The population form would return zero, which is a true statement about that one value but a misleading one if you meant it to describe anything wider.
saying these in an interview costs you the question
- Says always use n-1 without asking what the rows represent
- Claims n-1 corrects for outliers or skew
- Says the n-1 variance is smaller than the n version
- Cannot say which divisor their own tool used
- Applies n-1 to a full census of the group being described