skip to content

What is the difference between variance and standard deviation?

level: juniorimportance: must knowfreq 85%

answer

  1. start from the units you report in
  2. average of the squared deviations
  3. dollars-squared means nothing to a reader
  4. one is the square root of the other

basics

~20 s

Variance is the average squared deviation from the mean, so it is measured in squared units. Standard deviation is the square root of the variance, which puts the number back into the data's original units and makes it readable.

solid answer

~40 s

Both measure how far values sit from the mean, and they carry exactly the same information. Variance averages the squared deviations: `sigma^2 = (1/N) * sum (x_i - mu)^2` for a population. Standard deviation is its square root. Squaring is what makes the deviations stop cancelling out, but it also squares the units, so a salary column with variance 144,000,000 has a variance in dollars-squared — a quantity nobody can picture. Its square root, a standard deviation of 12,000 dollars, sits on the same scale as the salaries and can be compared directly with the mean. That is why variance lives inside the algebra and standard deviation goes into the report. A standard deviation is never negative, and it is zero only when every value equals the mean.

go deeper

for a junior

Be ready to define both in one breath and convert between them: square the standard deviation to get the variance, take the root to go back. Say the word 'units' out loud — that is the point interviewers listen for.

for a middle

Explain why deviations are squared at all, since raw deviations sum to zero, and why the squaring forces the units up. Be able to compute both by hand on three or four numbers without hesitating.

for a senior

Show judgment about which one belongs in front of a reader and what caveats travel with it. Know that the share of data inside one standard deviation depends on distribution shape and say so before quoting any percentage.

for a principal

Own the reporting convention. Decide once whether teams publish spread in original units alongside every headline mean, and make sure dashboards never surface a squared-unit quantity to a business audience.

## What both numbers measure Every measure of spread answers the same question: how far do individual values sit from the centre? Variance and standard deviation answer it the same way — by measuring distance from the arithmetic mean — and differ only in the scale the answer is reported on. Variance takes each observation, subtracts the mean, squares that difference, and averages the squares. Standard deviation is the square root of the variance. ## The formulas For a population of N values with mean `mu`: - variance: `sigma^2 = (1/N) * sum (x_i - mu)^2` - standard deviation: `sigma = sqrt(sigma^2)` For a sample of n values with sample mean `xbar`: - variance: `s^2 = (1/(n-1)) * sum (x_i - xbar)^2` - standard deviation: `s = sqrt(s^2)` Which divisor applies is its own decision, but the shape is the same either way: centre, square, average, and — for the standard deviation — take the root. ## Why square the deviations Raw deviations from the mean always sum to exactly zero, because the positives cancel the negatives, so their average is useless as a spread measure. Squaring removes the sign. It also weights far points more heavily than near ones: a value three units away contributes nine times as much as a value one unit away. That is a modelling choice rather than a law of nature — averaging absolute deviations is a perfectly legitimate alternative — but squares are smooth and differentiable, which is why they sit at the centre of least-squares fitting. ## Units are the practical difference Squaring the deviations squares the units too. A salary column measured in dollars has a variance measured in dollars squared. Nobody can picture 144,000,000 squared dollars. Take the square root and you get a standard deviation of 12,000 dollars: a number on the data's own scale, comparable against the mean, against a threshold, or against last quarter's figure. This is why reported spreads and error bars are almost always standard deviations, while variance stays in the derivation. ## Worked example Two classes sit the same exam. Class A scores 30, 50, 70. Class B scores 49, 50, 51. Both have a mean of exactly 50, so the mean alone declares the classes identical. Class A: deviations -20, 0, 20; squares 400, 0, 400; sum 800. Dividing by N = 3 gives a population variance of about 266.7 and a standard deviation of about 16.3. Class B: deviations -1, 0, 1; sum of squares 2; population variance about 0.67 and a standard deviation of about 0.82. (With the n-1 divisor the figures are 400 and 20 for class A, 1 and 1 for class B.) One class is wildly uneven, the other nearly uniform, and only the spread measure tells you which is which. ## Reading a standard deviation A standard deviation is a typical distance from the mean in the data's own units — more precisely the root-mean-square distance, which is always at least as large as the plain average absolute distance. How much of the data actually falls inside one standard deviation depends on the shape of the distribution. For roughly bell-shaped data, about 68% of values fall within one standard deviation of the mean and about 95% within two. For any distribution whatsoever, Chebyshev's inequality guarantees only that at least `1 - 1/k^2` of the values lie within k standard deviations for k > 1 — at least 75% within two, at least 88.9% within three. The assumption-free guarantee is far weaker, and that is exactly the point: quoting the 68% figure for a skewed column is an assumption, not a fact. ## Properties worth having ready - Both are non-negative. A standard deviation of zero means every value equals the mean; there is no other way to reach zero. - Standard deviation carries the data's units; variance carries their square. - Both are built around the mean, so both inherit the mean's sensitivity to extreme values. - Conversion is exact in both directions, so neither contains information the other lacks. ## Common mistakes The frequent interview error is quoting a variance as if it lived on the data's scale ("salaries vary by 144 million"), which is wrong by a squaring. The second is calling the standard deviation the average distance from the mean; it is the root-mean-square distance, and the two agree only in degenerate cases. The third is treating a large variance as evidence of dirty data — a genuinely wide distribution produces a large variance with no bad rows in it at all.

  • What does a standard deviation of exactly zero tell you about a column?
    That every value in the column is identical and equal to the mean. A standard deviation is the root of an average of squares, so it can only reach zero when every squared deviation is zero. In practice a zero standard deviation on a column that should vary usually means a constant default was written, a single value was broadcast across rows, or the column was filtered down to one distinct value.
  • Why square the deviations rather than averaging their absolute values?
    Averaging absolute deviations is a valid spread measure and is easier to explain, since it really is the mean distance from the centre. Squaring wins on mathematical convenience: it is smooth and differentiable everywhere, it makes least-squares fitting tractable, and it produces the quantity that the rest of statistical theory is written in. The cost is that squares weight far-away points heavily, so squared measures react strongly to extremes.
  • Can standard deviation exceed the mean of a column?
    Yes, easily. For a strictly positive, heavily right-skewed column such as request latency or purchase value, a long tail can push the standard deviation above the mean. It signals that spread dominates the level, so the mean is a weak summary of a typical value. It is not an error and not proof of bad data; a skewed distribution simply looks like that.

Variance is an area, standard deviation is the side of the square: both describe the same figure, but only one is a length you can hold a ruler against.

saying these in an interview costs you the question

  • Reports variance in the data's original units
  • Says standard deviation is exactly the average absolute distance from the mean
  • Claims a standard deviation can be negative
  • Treats a large variance as proof the data is dirty
  • Cannot convert between variance and standard deviation on demand

context