skip to content

How does consistency of an estimator differ from unbiasedness?

level: middleimportance: should knowfreq 56%

answer

  1. one is about now, one about the limit
  2. centre of the distribution versus its collapse
  3. an estimator that ignores extra data
  4. bias to zero plus variance to zero
  5. neither property implies the other

basics

~20 s

Unbiasedness is a finite-sample property: the estimator is centred on the parameter at every sample size. Consistency is an asymptotic one: the estimator converges to the parameter as the sample grows. Neither implies the other.

solid answer

~50 s

Unbiasedness says `E[theta_hat] = theta` holds for every n, however small. Consistency says `theta_hat` converges in probability to `theta` as n goes to infinity: for any tolerance you name, the chance of missing by more than that tolerance goes to zero. They are logically independent. Estimating the population mean by the first observation alone is unbiased at every n but never converges, since its spread is that of a single data point however much data you collect. Conversely, dividing the sum of squared deviations by n is biased low at every finite n, yet that bias and its variance both vanish as n grows, so it is consistent. A useful bridge: bias to zero plus variance to zero implies consistency. In practice consistency is the more basic requirement — an estimator that will not converge is not worth having.

go deeper

for a junior

Be ready to state that unbiasedness holds at any sample size while consistency is about what happens as the sample grows without bound, and that neither one implies the other.

for a middle

Expect to supply a concrete estimator on each side: one unbiased at every n that never converges, and one biased at every n that does. Then give the bias-to-zero plus variance-to-zero bridge.

for a senior

Show the applied ordering — you would refuse an estimator that will not converge, but knowingly accept small bias for a large variance reduction — and flag that asymptotic guarantees say nothing about the n you actually have.

for a principal

Own the standard for what the organisation is allowed to ship: which asymptotic arguments count as evidence, at what sample sizes they may be invoked, and how finite-sample behaviour gets checked before a metric goes live.

## Two different questions Both properties are about the sampling distribution of an estimator, but they ask different things. **Unbiasedness** asks: *where is this distribution centred, right now, at the sample size I have?* Formally `E[theta_hat] = theta` for every admissible value of the parameter. It is a statement about a fixed n, and it must hold at n = 5 as much as at n = 5000. **Consistency** asks: *what happens as I keep collecting data?* Formally `theta_hat_n` converges in probability to `theta`: for any tolerance `epsilon > 0`, `P(|theta_hat_n - theta| > epsilon)` tends to 0 as n grows. It is a statement about a whole sequence of estimators, one per sample size, and it says nothing about how the estimator behaves at any particular n. A compact way to hold them apart: unbiasedness is about **location** at fixed n; consistency is about **collapse** onto the target as n grows. ## Unbiased but not consistent Take the estimator that reports the first observation and discards the rest: `theta_hat = X_1`. Its expectation is `mu`, so it is unbiased at every sample size. But its variance is `sigma^2` no matter what n is — collecting a million observations does not narrow it by a hair. The probability of missing `mu` by more than any fixed tolerance stays constant, so it never converges. It is unbiased and useless. The lesson: unbiasedness constrains only the centre of the distribution. Nothing in the definition forces the distribution to tighten. ## Consistent but biased at every finite n Take the sum of squared deviations about the sample mean divided by n. Its expected value is `(n-1)/n * sigma^2`, so its bias is `-sigma^2 / n` — strictly negative for every finite n, never zero. Yet that bias shrinks toward zero as n grows, and so does its variance, so the estimator settles onto `sigma^2`. It is biased at every sample size you will ever actually have, and consistent. This pair of examples is the whole answer. Interviewers ask the question specifically to see whether you can produce one estimator of each kind. ## The bridge between them There is a sufficient condition worth knowing. Mean squared error decomposes as ``` MSE(theta_hat) = Var(theta_hat) + Bias(theta_hat)^2 ``` If both the bias and the variance tend to zero as n grows, the mean squared error tends to zero, and an estimator whose mean squared error vanishes converges in probability to the parameter — it is consistent. This is why asymptotically unbiased estimators with shrinking variance are consistent, and it explains both examples above: the first-observation estimator has zero bias but non-vanishing variance, so it fails; the n-divisor variance has non-zero bias that vanishes and vanishing variance, so it succeeds. Note the condition is sufficient, not necessary. An estimator can be consistent without having a finite mean at all, in which case its bias is not even defined. Consistency is the weaker, more robust requirement. ## Asymptotic unbiasedness — a third, distinct thing Candidates often conflate consistency with **asymptotic unbiasedness**, the statement that `E[theta_hat_n]` tends to `theta`. These are not the same. An estimator can be asymptotically unbiased while its variance refuses to shrink, in which case it is not consistent. Keep three labels apart: unbiased (centred now), asymptotically unbiased (centred eventually), consistent (collapses onto the target eventually). ## Which matters more In applied work, **consistency is close to non-negotiable and unbiasedness is negotiable**. An estimator that does not converge to the truth as data accumulates is broken in a way no sample size fixes. Small bias, by contrast, is routinely accepted — often deliberately purchased — in exchange for a large drop in variance, because what you actually care about is total error, not the location of a theoretical average. There is also a practical caveat: consistency is an asymptotic promise, and it says nothing about the sample size you actually have. "It is consistent" is not a defence of an estimate computed from twelve observations. When you invoke consistency, be ready for the interviewer's next question, which is how fast it converges and whether your n is anywhere near that regime. ## Saying it in an interview Define both crisply, insist they are logically independent, then hand over the two examples: the first-observation estimator (unbiased, not consistent) and the n-divisor sum of squares (biased at every n, consistent). Close with the bridge — bias to zero plus variance to zero implies consistency — and the applied verdict that you would not ship an inconsistent estimator but would happily ship a slightly biased one.

  • Give an estimator that is asymptotically unbiased but still not consistent.
    Average the first observation with a term that fades, for instance `X_1 + c/n`. Its expectation tends to `mu`, so it is asymptotically unbiased, yet its variance stays at `sigma^2` forever because it still rests on one data point. It never collapses onto the parameter. This shows asymptotic unbiasedness and consistency are separate claims: one is about the centre in the limit, the other about the whole distribution.
  • Why is consistency usually considered the more important of the two?
    An inconsistent estimator cannot be rescued by collecting more data, which is the one lever an analyst always has. Bias, by contrast, is often small, sometimes computable, and frequently worth accepting for a large variance reduction. Interviewers want to hear that you would reject an estimator that fails to converge outright, while treating unbiasedness as one desirable property to be traded against spread.
  • Is "it is consistent" a good defence of an estimate computed from twelve observations?
    No. Consistency is a promise about the limit and says nothing about any particular sample size. The honest questions at n = 12 are how large the finite-sample bias is, how wide the spread is, and whether the asymptotic behaviour has kicked in at all. Quoting an asymptotic property to justify a tiny-sample number is one of the more common ways candidates overstate what they know.

saying these in an interview costs you the question

  • Says consistent is just another word for unbiased
  • Claims every unbiased estimator is automatically consistent
  • Cannot produce an example of either property failing
  • Confuses asymptotic unbiasedness with consistency
  • Treats consistency as a guarantee at the sample size on hand
  • Thinks a biased estimator can never converge to the truth

context