What does the Cramer-Rao lower bound say about the variance of an unbiased estimator?
answer
- a floor on precision, not a guarantee
- reciprocal of the information in the sample
- only holds for unbiased estimators
- Cauchy-Schwarz between estimator and score
- support must not depend on the parameter
basics
~20 sIt sets a floor on precision: under regularity conditions any unbiased estimator of a parameter has variance at least the reciprocal of the Fisher information in the sample. No unbiased estimator can beat that floor, and some attain it exactly.
solid answer
~50 sFor `n` i.i.d. observations and an unbiased estimator `T` of `theta`, the Cramer-Rao inequality states `Var(T) >= 1 / (n * I_1(theta))`, where `I_1` is the Fisher information carried by one observation. The bound says information is a hard budget: you cannot squeeze more precision out of the data than the model's curvature allows, only waste it. An estimator that meets the bound is called efficient in the exact, finite-sample sense. The Bernoulli sample proportion attains it — its variance is `p(1-p)/n`, exactly the reciprocal of `n / (p(1-p))`. Many estimators do not; the usual unbiased variance estimator for a Normal has variance `2*sigma^4/(n-1)`, above the bound `2*sigma^4/n`. Two caveats matter: the bound applies only under regularity conditions, notably a support that does not depend on `theta`, and only to unbiased estimators — a biased estimator can have smaller mean squared error.
go deeper
Recall the shape of the claim: more Fisher information means a lower floor on the variance of an unbiased estimator. Getting the direction of the inequality right already puts you ahead.
Be ready to write the bound as one over n times the per-observation information, name the unbiasedness and regularity preconditions, and give one estimator that attains it.
Show you know when the bound is silent: biased or penalised estimators are outside its scope, and models whose support depends on the parameter break the derivation outright.
Frame it as a budget conversation. When the bound says the design cannot deliver the precision a decision needs, the answer is more or better-targeted data, not a cleverer estimator.
## The statement Let `T` be an estimator of a scalar parameter `theta` computed from `n` independent, identically distributed observations, and suppose `E[T] = theta` for all `theta` (unbiasedness). Under regularity conditions, the **Cramer-Rao lower bound** says `Var(T) >= 1 / I_n(theta) = 1 / (n * I_1(theta))` where `I_1(theta)` is the Fisher information in a single observation — the variance of the score, equivalently the expected negative curvature of the one-observation log-likelihood. The inequality generalises. If `T` is not unbiased for `theta` but estimates some smooth function, with `E[T] = psi(theta)`, the bound becomes `Var(T) >= psi'(theta)^2 / I_n(theta)`. And for a vector parameter, the covariance matrix of an unbiased estimator minus the inverse information matrix is positive semi-definite, which in particular bounds each variance on the diagonal from below. ## Where it comes from The proof is one application of the Cauchy-Schwarz inequality. The score `U(theta)` has mean zero, and differentiating the unbiasedness identity `E[T] = theta` under the integral sign yields `Cov(T, U) = 1`. Cauchy-Schwarz then gives `1 = Cov(T, U)^2 <= Var(T) * Var(U) = Var(T) * I_n(theta)`, and rearranging is the bound. That derivation also tells you exactly when equality holds: Cauchy-Schwarz is tight only when the two random variables are perfectly linearly related, so an estimator attains the bound precisely when `T - theta` is proportional to the score. This is the structural reason only a narrow family of models admits bound-attaining estimators. ## Attaining it, and not **Attained.** For Bernoulli data with success probability `p`, one observation carries information `I_1(p) = 1/(p(1-p))`, so the bound for any unbiased estimator of `p` is `p(1-p)/n`. The sample proportion has exactly that variance, so it is efficient in the strict finite-sample sense — no unbiased alternative can do better at any sample size. **Not attained.** For Normal data with unknown mean, the standard unbiased estimator of the variance `sigma^2` has variance `2*sigma^4/(n-1)`, while the Cramer-Rao bound for `sigma^2` is `2*sigma^4/n`. The estimator sits strictly above the floor. That is not a defect to fix — it reflects that no unbiased estimator here is a linear function of the score. **Bound does not apply.** The regularity conditions require the support of the distribution to be free of the parameter, so that you may differentiate under the integral sign. For observations uniform on `(0, theta)`, the support *is* the parameter, the derivation collapses, and estimators based on the sample maximum achieve precision that shrinks like `1/n` rather than the `1/sqrt(n)` the naive bound would suggest. Quoting a Cramer-Rao bound for such a model is a classic error. ## Two boundaries people trip over **Unbiasedness is a precondition, not a decoration.** The bound constrains variance only among unbiased estimators. A shrunken or penalised estimator accepts some bias and can have strictly smaller mean squared error than the Cramer-Rao floor. There is no contradiction: it is simply outside the class the theorem quantifies over. **The bound is about variance, not about achievability by any particular procedure.** It says what is impossible, not what a given method delivers. A high bound (little information) means no procedure will save you; a low bound does not promise your estimator reaches it. ## The asymptotic connection The bound is the exact-sample counterpart of the large-sample story: as `n` grows, the maximum-likelihood estimate is approximately Normal with variance `1/(n * I_1(theta))`, so it meets the bound in the limit even in models where no finite-sample estimator attains it. That is what people mean when they say maximum likelihood is asymptotically efficient. Notice the direction of the logic: the bound is a fact about the model and its information, and asymptotic likelihood theory says the maximum-likelihood estimate eventually saturates it. ## Interview framing A good answer states the inequality with the correct direction (variance is at least the reciprocal of information), names the two conditions (unbiased, regular), gives one example that attains it and one that does not, and closes with the asymptotic link. A weak answer flips the inequality, forgets unbiasedness, or asserts that every maximum-likelihood estimate attains the bound at finite `n`.
- Can any estimator have variance below the Cramer-Rao bound?Yes, if it is biased. The bound quantifies only over unbiased estimators, so a shrunken or penalised estimator can sit below it and can even have smaller mean squared error, paying for the variance reduction with bias. The bound is also inapplicable when regularity fails, such as when the support depends on the parameter.
- What has to be true for an estimator to attain the bound exactly?The proof runs through Cauchy-Schwarz, which is tight only under perfect linear dependence. So equality holds exactly when the centred estimator is proportional to the score function at every data point. That happens in a narrow family of models — the Bernoulli sample proportion is the standard example — which is why most unbiased estimators sit strictly above the floor.
- How does the bound change when the parameter is a vector?The covariance matrix of an unbiased estimator minus the inverse Fisher information matrix is positive semi-definite. Reading the diagonal gives a lower bound on each individual variance, but the matrix form says more: any linear combination of the parameters also has its variance bounded below by the corresponding quadratic form in the inverse information.
saying these in an interview costs you the question
- States the inequality with the direction reversed
- Drops the unbiasedness condition and claims a universal floor
- Says maximum likelihood always attains the bound at finite n
- Applies the bound to a uniform model whose support depends on the parameter
- Confuses the bound on variance with a bound on mean squared error