skip to content

Why does filling missing numeric values with the column mean understate the standard deviation?

level: middleimportance: must knowfreq 58%

answer

  1. count the rows, then count the deviations
  2. each filled point sits on the mean
  3. numerator unchanged, denominator grows
  4. n minus 1 replaces m minus 1

basics

~20 s

Every filled value sits exactly at the mean, so it adds nothing to the sum of squared deviations while still adding to the row count. The numerator is unchanged and the denominator grows, so variance and standard deviation shrink.

solid answer

~50 s

Sample variance is `s^2 = sum (x_i - xbar)^2 / (n - 1)`. If `k` of the `n` rows are filled with the mean of the `m` observed values, each filled row contributes a deviation of exactly zero, so the numerator stays at the observed sum of squares while the denominator rises from `m - 1` to `n - 1`. The filled variance is therefore the honest one multiplied by `(m - 1) / (n - 1)`. With 100 rows and 40 missing that factor is 59/99, about 0.60, so the standard deviation comes out roughly 23% too small. The damage compounds in the standard error `s / sqrt(n)`, which now divides a shrunken `s` by a larger `n` — confidence intervals get far too narrow. And the mean itself is untouched by mean filling, so it corrects no bias; it only manufactures confidence.

go deeper

for a junior

Be ready to say that every filled value equals the mean, so the spread shrinks and any interval built on it looks tighter than the data earns. The intuition plus one sentence of reasoning is enough at this level.

for a middle

You are expected to write the variance formula and point at exactly which part changes: the sum of squared deviations stays put while the divisor grows from the observed count to the full count. Then carry it through to the standard error.

for a senior

Demonstrate that you trace the consequence downstream — narrowed confidence intervals, deflated test statistics, attenuated correlations — and that you would publish the missingness rate alongside any statistic computed after filling.

for a principal

Own the standard for what leaves the team: whether filled values may enter a reported metric at all, and whether propagating uncertainty is mandatory rather than optional when a convenient single number is available.

## The arithmetic Let a column have `n` rows, of which `m` are observed and `k = n - m` are missing. Mean filling replaces every missing cell with `xbar_obs`, the mean of the `m` observed values. Start with the mean. The filled column's total is `m * xbar_obs + k * xbar_obs = n * xbar_obs`, so the filled mean is `xbar_obs` exactly. Mean filling does not move the centre at all — which is the first thing to say in an interview, because it makes clear that the technique buys no accuracy whatsoever. Now the spread. Sample variance is `s^2 = sum (x_i - xbar)^2 / (n - 1)`. Each filled value equals `xbar`, so `(x_i - xbar)^2 = 0` for all `k` of them. The numerator is therefore identical to the observed-only sum of squares, while the denominator has grown from `m - 1` to `n - 1`. So `s^2_filled = s^2_obs * (m - 1) / (n - 1)` With 100 rows and 40 missing, that is `59/99 ≈ 0.596`. The standard deviation, being the square root, comes out at about 77% of the honest value — roughly 23% too small. ## Why the interval damage is worse than the spread damage The standard error of the mean is `SE = s / sqrt(n)`. Mean filling attacks both terms in the wrong direction at once: `s` has been deflated, and `n` has been inflated from `m` to the full row count, as though the filled rows carried information. In the 100-row example the filled `SE` is around 60% of the standard error you would honestly report from the 60 observed values, so a 95% confidence interval built on it is about 40% too narrow, and any test statistic using it is correspondingly too large. The result is a number that looks more certain the more data you were missing, which is precisely backwards. ## What else it distorts - **Shape.** A histogram of the filled column shows a spike at the mean that exists nowhere in reality. Kurtosis rises, and any distributional check on the column is now testing an artefact. - **Associations.** Fill column A at its mean and the filled rows sit at one fixed A value regardless of their B value. Those rows contribute nothing to the covariance numerator while still counting toward the sample size, so the estimated correlation is attenuated toward zero. A real relationship can be flattened into apparent noise. - **Groups.** Filling with a single global mean when the column differs by segment drags every segment toward the overall centre, compressing exactly the between-group differences you were probably trying to measure. - **Quantiles.** Medians, percentiles and any interval-based summary shift toward the mean as the pile of filled values grows. ## What mean filling does and does not assume Because the filled mean equals the observed-case mean, mean filling inherits whatever bias complete-case analysis had. If the missingness is MCAR, that bias is zero and the only sin is the fake precision. If it is MAR or MNAR, the centre is already off and mean filling adds a false claim of certainty on top of an already wrong number. That is why it is often described as the worst of both worlds: it does not fix accuracy, and it destroys the honest measure of uncertainty that would have warned you. ## Better single fills, and why they are still not enough A point fill of any kind understates spread; the mean is merely the worst case, because it puts every filled value at the exact centre. Filling from a conditional model plus a random residual, or hot-deck sampling a similar row's actual value, preserves roughly the right variance and roughly the right associations. But any *single* completed dataset still presents guessed values as though they were measurements, so downstream standard errors remain too small. The only treatment that carries the uncertainty through is to build several completions with random draws and pool the results, so that the disagreement between completions enters the reported error. ## What to say in the room Write the variance formula, point at the numerator and the denominator, and say which one moved. Follow with the standard error, because that is where the practical damage lands. Then close with the honesty point: the mean is unchanged, so nothing was gained; the spread shrank, so something was lost. If you ever must report a statistic computed after filling, report the missingness rate beside it.

  • Does mean filling change the sample mean itself?
    No. Filling the gaps with the observed mean leaves the overall mean exactly equal to the observed-case mean, so mean filling inherits whatever bias complete-case analysis already had. It buys nothing on accuracy and costs you an honest measure of spread, which is the worse half of the trade.
  • What does mean filling do to the correlation between two columns?
    It attenuates the correlation toward zero. Filled rows sit at one column's mean regardless of what the other column says, so they add a flat band of points with no association. The covariance numerator gains nothing while the row count grows, and a real relationship can be flattened into apparent noise.
  • If you must put a single value in each gap, what is less damaging than the mean?
    A draw rather than a point. Filling from a conditional model with a random residual added, or copying an actual value from a similar row, keeps the spread and the associations roughly right. Any single completed dataset still understates uncertainty, though, since guessed values are presented as measurements.

Imagine telling forty people in a hundred-person lineup to stand at exactly the group's average height. The average height is unchanged, but the lineup now looks far more uniform than the crowd really is.

saying these in an interview costs you the question

  • Says mean filling is safe because the mean is unchanged
  • Reports the post-fill standard deviation as the real spread
  • Thinks imputation adds information rather than assumptions
  • Ignores that intervals built after filling are too narrow
  • Fills with a global mean across clearly distinct groups

context