skip to content

Under what assumptions is OLS the best linear unbiased estimator?

level: middleimportance: should knowfreq 55%

answer

  1. four substantive conditions, one technical
  2. zero conditional mean is the load-bearing one
  3. same spread, no correlation between errors
  4. best only within a restricted class
  5. normality is conspicuously absent

basics

~20 s

The Gauss-Markov conditions: linearity in the coefficients, errors with zero mean given the predictors, constant variance, and no correlation between them, plus non-redundant predictors. Then OLS has the smallest variance among linear unbiased estimators. Normality is not required.

solid answer

~50 s

Gauss-Markov needs four substantive conditions plus one technical one. The model must be linear in the parameters; the errors must satisfy `E[e | X] = 0`, so no predictor carries information about the average disturbance; the errors must be homoscedastic, `Var(e_i | X) = sigma^2` for every observation; and they must be mutually uncorrelated, `Cov(e_i, e_j | X) = 0` for `i` not equal to `j`. Technically the predictor matrix must have full column rank so the estimator is defined. Under those conditions OLS is **BLUE**: among all estimators that are linear functions of `y` and unbiased for the coefficients, it has the smallest variance — for each coefficient and for every linear combination of them. Notice what is absent: no assumption that the errors are normal. Normality is what licenses the exact small-sample distributions of the usual test statistics; it is not needed for OLS to be BLUE.

go deeper

for a junior

Learn the acronym and what it expands to: best linear unbiased estimator. Be able to say that constant error variance and uncorrelated errors are required and that normality is not.

for a middle

Be ready to list all the conditions precisely, including zero conditional mean of the errors, and to explain what 'best' means — smallest sampling variance within the linear unbiased class.

for a senior

Show you can classify a violation you have actually met: does it cost efficiency only, or does it bias the estimates? Explain why extra data rescues one and not the other.

for a principal

Own the question of whether unbiasedness is even the right target. Argue when a biased estimator with lower mean squared error serves the decision better, and when the effort belongs in the study design rather than the estimator.

## The statement The Gauss-Markov theorem says: under a specific set of conditions on the linear model, the ordinary least squares estimator is the **B**est **L**inear **U**nbiased **E**stimator of the coefficients. Each word in that acronym is load-bearing and each one is also a restriction. ## The assumptions, one at a time **1. Linearity in the parameters.** The model is `y_i = b0 + b1*x_i1 + ... + bk*x_ik + e_i`. The predictors themselves may be transformed however you like — squares, logs, interactions — as long as the coefficients enter linearly. `y = b0 + b1*log(x)` qualifies; `y = b0*exp(b1*x)` does not. **2. Zero conditional mean of the errors: `E[e_i | X] = 0`.** Knowing the predictors tells you nothing about the average size of the disturbance. This is the assumption that carries the most content. It fails when a variable that affects `y` is omitted and is correlated with an included predictor, when `y` feeds back into `x`, or when a predictor is measured with error. Failure here makes OLS **biased**, and no amount of extra data fixes it. **3. Homoscedasticity: `Var(e_i | X) = sigma^2` for all `i`.** The spread of the disturbance is the same at every combination of predictor values. Failure is common in practice — spending, revenue and count-like outcomes often have variance that grows with their level. **4. No correlation between errors: `Cov(e_i, e_j | X) = 0` for `i` not equal to `j`.** Failure is the norm in time series and in clustered or repeated-measures data. **5. (Technical) The predictor matrix has full column rank.** Without this, the estimator is not uniquely defined at all, so there is nothing to be best. Assumptions 3 and 4 are often stated together as *spherical errors*. ## What each word of BLUE means **Linear** — the comparison class contains only estimators that can be written as a fixed weight matrix times `y`, with weights that do not depend on `y`. OLS is one of these: `bhat = (X'X)^-1 X'y`. Any estimator that uses `y` non-linearly is outside the theorem's scope entirely. **Unbiased** — the comparison class contains only estimators whose expectation equals the true coefficient vector. Biased estimators are excluded from the contest by the rules, not beaten in it. **Best** — smallest sampling variance. Formally, for any competing linear unbiased estimator, the difference between its covariance matrix and the OLS covariance matrix is positive semidefinite, which means OLS has no larger variance for every individual coefficient and for every linear combination of them. So the theorem is a statement about the *winner of a restricted race*, and the restrictions are as important as the result. ## What Gauss-Markov does not say - **It does not say normality is required.** No distributional shape is assumed anywhere in the conditions. Normality is a separate, additional assumption used to get exact small-sample sampling distributions for test statistics; drop it and OLS is still BLUE. - **It does not say no estimator can beat OLS.** A *non-linear* unbiased estimator falls outside the class and is not ruled out by the theorem. And a deliberately *biased* estimator can have a smaller mean squared error, because MSE is bias-squared plus variance and a small bias can buy a large variance reduction. - **It does not say the model is correct.** BLUE is a conditional promise: *if* the assumptions hold, *then* OLS is efficient in that class. If assumption 2 fails, the estimator is biased and its efficiency is beside the point. - **It does not say anything about how good the fit is.** Efficiency of the estimator and explanatory power of the model are unrelated questions. ## When the assumptions fail The failures split cleanly into two tiers, and knowing which tier a violation lands in is what interviews are probing. *Efficiency-only failures.* If the errors are heteroscedastic or correlated but `E[e | X] = 0` still holds, OLS remains **unbiased** — it just stops being minimum-variance among linear unbiased estimators. A weighted or generalised least squares estimator that accounts for the error structure takes the BLUE title instead. The estimates are still pointing at the right target; they are simply noisier than they need to be, and the textbook variance formula that assumes constant, uncorrelated errors no longer describes their true sampling variability. *Bias failures.* If `E[e | X]` is not zero — an omitted confounder correlated with a predictor, simultaneity, or measurement error in a predictor — OLS is biased and remains biased however large the sample grows. This is the failure that invalidates causal claims, and it cannot be repaired by adjusting how variability is computed. A candidate who can put a violation into the right tier is demonstrating the practical content of the theorem; one who lists normality among the assumptions is signalling that they memorised the wrong list.

  • Does Gauss-Markov require the errors to be normally distributed?
    No. The conditions are zero conditional mean, constant variance and no correlation between errors, plus linearity in the parameters — nothing about distributional shape. Normality is an extra assumption used to derive exact small-sample distributions for the usual test statistics. Without it, OLS is still the best linear unbiased estimator; you only lose the exact finite-sample inference.
  • If Gauss-Markov holds, does that mean no estimator can beat OLS?
    No. The theorem only rules out linear unbiased competitors. A non-linear unbiased estimator is outside the class the theorem covers, and a deliberately biased estimator can achieve a lower mean squared error by trading a little bias for a large variance reduction. BLUE is a statement about a restricted comparison, not a universal optimality claim.
  • Which Gauss-Markov assumption actually threatens your conclusions in applied work?
    Zero conditional mean of the errors. Violations of constant variance or of uncorrelated errors cost efficiency — the estimates stay unbiased, they are just noisier than necessary. A violation of zero conditional mean, from an omitted confounder correlated with a predictor, from simultaneity, or from measurement error in a predictor, makes the estimates biased no matter how much data you collect.

BLUE is winning a race with strict entry rules: only estimators that are linear in the data and unbiased are allowed to compete. OLS wins that race, which says nothing about competitors who were never allowed on the track.

saying these in an interview costs you the question

  • Lists normality among the Gauss-Markov assumptions
  • Says OLS has the lowest variance of every possible estimator
  • Treats BLUE as proof the model is correctly specified
  • Thinks heteroscedasticity makes the coefficient estimates biased
  • Cannot separate efficiency failures from bias failures

context