Why can a biased estimator have lower mean squared error than an unbiased one?
answer
- unbiasedness zeroes only one of two terms
- total error, not just the centre
- squared bias plus variance
- small samples, pull toward the pool
- trade a little shift for much less wobble
basics
~20 sMean squared error is variance plus squared bias. Accepting a little bias can cut variance far more than the squared bias adds, so total error drops. Shrinking a noisy small-sample average toward a pooled average is the standard example.
solid answer
~50 sThe error criterion that matters is usually total squared error, and it splits cleanly: `MSE(theta_hat) = Var(theta_hat) + Bias(theta_hat)^2`. Unbiasedness zeroes only the second term, and there is no rule saying the estimator that does so has small variance. If a biased alternative buys a large variance reduction for a small squared bias, its MSE is lower and it wins on the criterion you actually care about. The workhorse case is shrinkage: a team's average handling time computed from six tickets is centred but wildly noisy, so you report `w * team_average + (1 - w) * grand_average` with `w` below one. That estimate is biased toward the grand average, yet for small team samples the variance saved dwarfs the squared bias introduced, and typical error falls. The judgement call is the weight and the direction of the pull — and whether a systematically shifted number is acceptable for the decision it feeds.
go deeper
Be ready to write mean squared error as variance plus squared bias and say why zeroing only the bias term does not minimise total error.
Explain the mechanics of the trade: a weight below one multiplies variance by its square while introducing a bias toward the pooled centre, and show why small samples make that trade favourable.
Demonstrate operating judgement — pick the weight from how noisy each unit is, and name the reports where shrinkage is wrong, such as outlier hunting or audited figures. Say how you would disclose the choice.
Own the policy: which published metrics may be stabilised, who signs off on the loss function being optimised, and how the organisation avoids two differently-computed versions of the same headline number.
## The decomposition For an estimator `theta_hat` of a parameter `theta`, mean squared error is the expected squared distance from the truth: ``` MSE(theta_hat) = E[(theta_hat - theta)^2] ``` Write `E[theta_hat] = m`. Then ``` E[(theta_hat - theta)^2] = E[((theta_hat - m) + (m - theta))^2] = E[(theta_hat - m)^2] + 2*(m - theta)*E[theta_hat - m] + (m - theta)^2 ``` The middle term vanishes because `E[theta_hat - m] = 0`, leaving ``` MSE = Var(theta_hat) + Bias(theta_hat)^2 ``` Two non-negative pieces. Unbiasedness sets the second to zero and leaves the first entirely unconstrained. That is the whole argument: minimising one component is not the same as minimising the sum. ## The trade, made concrete Suppose you report average handling time per support team. A team with six tickets last week has a sample mean that is unbiased for its true average but has a standard error roughly `sigma / sqrt(6)` — huge. Rank teams on that number and the top and bottom of the league table are mostly noise; the smallest teams appear at both extremes, which is the tell. A shrunken estimator pulls each team toward the grand average across all teams: ``` theta_hat_team = w * team_average + (1 - w) * grand_average ``` with `0 < w < 1`, and with `w` chosen close to 1 for teams with lots of data and close to 0 for teams with almost none. This estimator is **biased** for any team whose true average differs from the grand average: its expectation is `w * theta_team + (1 - w) * grand_average`, which is systematically pulled inward. Its variance, however, is `w^2` times the variance of the raw average — for `w = 0.5` that is a 75% cut. Whether the trade pays depends on how far the team truly sits from the pool. Nearby teams gain enormously; a genuine outlier team pays a real bias cost. Averaged across many teams, when most are close to the pool and each has little data, total squared error falls — often dramatically. ## Why this is not a trick It is worth being explicit that this is not sleight of hand. There are known results in which a biased estimator has strictly lower mean squared error than the natural unbiased one *for every value of the parameter*, not merely on average over some assumed spread of parameters — the James-Stein estimator of a multivariate normal mean, which dominates the sample mean when three or more means are estimated at once, is the celebrated example. The mere existence of such a result kills the reflex that unbiasedness is a requirement rather than a preference. ## When bias is the wrong purchase Shrinkage is not free, and a senior answer names the cases where you refuse it: - **You have plenty of data.** With large n the raw estimator is already tight; shrinking buys little variance and imports bias for nothing. - **The extremes are the point.** If the purpose of the report is to find genuinely unusual units — fraud detection, safety outliers — an estimator that systematically drags outliers toward the middle is actively harmful. - **The bias direction is contested.** Shrinking toward a pooled average is defensible when units are exchangeable. If they are not — different products, different customer bases — the pooled centre is not a sensible place to pull toward, and the bias is not "small" in any principled sense. - **The number is contractual or audited.** A figure that must be defended as the plain average of what happened cannot be quietly shrunk, whatever it does to MSE. - **Downstream consumers assume unbiasedness.** If the estimate is summed or differenced across many units, systematic pulls can accumulate in a direction nobody accounted for. ## The disclosure obligation When you do shrink, say so. Two numbers labelled "average handling time" that are computed differently — one shrunken, one raw — will be compared by someone eventually, and the discrepancy will be read as a bug. Publishing the estimator alongside the estimate, and keeping the raw average available, is what turns a defensible statistical choice into a defensible organisational one. ## Saying it in an interview Lead with the decomposition, in symbols: MSE is variance plus squared bias, and unbiasedness only zeroes the second term. Give the shrinkage example with a real, small n. Then show judgement by naming the conditions under which you would refuse the trade — abundant data, outlier-hunting reports, non-exchangeable units, audited figures. Interviewers are looking for the second half at least as much as the first.
- How would you choose the shrinkage weight in practice?Let it depend on how noisy each unit's own estimate is relative to how much the units genuinely differ. A unit with many observations keeps most of its own average; a unit with very few is pulled almost entirely to the pooled centre. The spread between units can be estimated from the data itself, and the weight can be validated by holding out later periods and comparing predicted against observed.
- When would you refuse to shrink and report the raw average instead?When the extremes are the point — safety outliers, fraud, incident hunting — because shrinkage systematically drags exactly those units toward the middle. Also when the units are not exchangeable, so the pooled centre is not a meaningful target; when data is abundant enough that variance is already small; and when the number is audited or contractual and must be defensible as the plain average of what happened.
- Does a lower-MSE biased estimator also give you valid uncertainty intervals?Not automatically. An interval built as the estimate plus or minus a multiple of its standard error is centred on a systematically shifted point, so its actual coverage of the true parameter can fall below the nominal level. If you shrink the point estimate you have to account for the bias when quantifying uncertainty, rather than reusing the interval machinery designed for the unbiased version.
A dart player whose throws land all over the board but average on the bullseye scores worse than one who lands tightly a centimetre off-centre. Total distance from the target is what counts, not whether the average lands on it.
saying these in an interview costs you the question
- Insists an unbiased estimator is always preferable
- States MSE as variance plus bias, not bias squared
- Cannot name a situation where shrinkage is harmful
- Treats shrinkage as adjusting the data rather than the estimator
- Shrinks outlier-detection metrics toward the pooled centre
- Reports a shrunken figure without disclosing the method