Does heteroscedasticity bias OLS coefficient estimates, or only their standard errors?
answer
- estimates versus the uncertainty around them
- unbiasedness rests on a different assumption
- the coefficient does not move at all
- t-statistics and intervals are the casualty
basics
~20 sHeteroscedasticity leaves OLS coefficient estimates unbiased and consistent; it breaks the conventional standard-error formula. Every t-statistic, p-value and confidence interval built from those standard errors is therefore untrustworthy, and OLS is no longer the minimum-variance linear estimator.
solid answer
~50 sOnly the standard errors. Provided the mean function is correctly specified, OLS coefficients stay unbiased and consistent under any pattern of error variance - unbiasedness follows from the errors having zero mean given the predictors, not from their variances being equal. What breaks is the conventional variance formula, which is derived assuming a single common error variance; under heteroscedasticity it is inconsistent, so every t-statistic, p-value and interval computed from it is wrong, and it is frequently too small rather than too large. You see this concretely when a slope's t-statistic drops from 2.6 to 1.7 after switching to HC1 sandwich standard errors while the coefficient itself does not move a single digit: the effect was never significant, the original standard error was simply understated. OLS also stops being BLUE - still unbiased, but no longer the minimum-variance linear unbiased estimator.
go deeper
Memorise the headline and be able to say it cleanly: estimates stay unbiased, standard errors do not. Know that t-statistics and confidence intervals are the things that become untrustworthy.
Explain why by naming the assumption behind each property - zero conditional error mean gives unbiasedness, equal variances give the simple variance formula - and describe what a sandwich estimator replaces.
Demonstrate the working protocol: refit with robust standard errors, compare, and judge whether any conclusion actually flips. Be able to argue that the direction of the standard-error distortion is not fixed in advance.
Own the reporting standard for the team - when robust inference is the default, what gets documented when robust and conventional results disagree, and how to keep the distinction between bias and precision from being blurred in write-ups.
## The claim to get right The single most common mistake on this topic is the sentence *heteroscedasticity biases the coefficients*. It does not. This is worth understanding at the level of where each property comes from, because then you never have to memorise it. ## Where unbiasedness comes from The OLS estimator can be written as the true parameter vector plus a term that is linear in the errors: `b_hat = b + (X'X)^-1 X'e`. Taking expectations conditional on the predictors, the second term has expectation zero as soon as `E(e | X) = 0`. That is the only condition unbiasedness needs. Nothing in that argument mentions the variance of `e`, so nothing about unequal variances can disturb it. Consistency follows similarly from the errors being uncorrelated with the predictors. Bias comes from a broken *mean* assumption - omitted variables correlated with a regressor, a mis-specified functional form, simultaneity, measurement error in a predictor. Heteroscedasticity is a *second-moment* problem and lives in a different part of the argument. ## Where the damage actually lands The conditional variance of the estimator is, in general, `Var(b_hat | X) = (X'X)^-1 * (X' Omega X) * (X'X)^-1`, where `Omega` is the covariance matrix of the errors. Under homoscedasticity `Omega = sigma^2 * I`, the middle collapses, and the whole thing simplifies to the familiar `sigma^2 * (X'X)^-1` - the formula behind the standard errors most software prints by default. When the variances differ, `Omega` is a diagonal matrix with unequal entries, the collapse is invalid, and the familiar formula estimates the wrong quantity. Critically it is not merely imprecise - it is **inconsistent**: more data does not make it converge to the truth. So: - standard errors are wrong, - t-statistics built from them are wrong, - p-values are wrong, - confidence intervals have the wrong coverage, - F-tests using the same variance estimate are wrong. ## Which direction is the error? There is no universal direction, and claiming there is one is a red flag. The conventional formula understates the true sampling variability when the large error variances sit at observations with unusual predictor values, which is the common case - variance rising with `x` puts the noisiest points at the ends of the predictor range, exactly where they have the most pull on the slope. That produces standard errors that are too small and false positives. But the reverse configuration - large variances concentrated near the centre of the predictor distribution - makes the conventional formula too *large* and costs you power. Test it rather than assume. ## The efficiency loss, separately The Gauss-Markov theorem says OLS is the Best Linear Unbiased Estimator - minimum variance among linear unbiased estimators - **under** homoscedastic, uncorrelated errors. Drop constant variance and the *unbiased* part survives while the *best* part does not. Intuitively, OLS gives every observation equal weight, but under unequal variances the low-noise observations carry more information about the line than the high-noise ones and deserve more weight. That is a real, if usually modest, loss: a correctly weighted estimator would have genuinely narrower sampling distribution. Note this is a different complaint from the broken standard errors, and it has a different fix. ## The fix, and what it does not do Heteroscedasticity-consistent (sandwich) standard errors estimate the middle of that variance expression directly from the squared residuals rather than assuming it away. The variants differ in their small-sample correction: HC0 is the raw form, HC1 applies a simple degrees-of-freedom scaling, and HC2 and HC3 additionally adjust each squared residual for how much the fit was pulled toward that observation, with HC3 the most conservative and the usual recommendation in modest samples. All of them are asymptotic arguments - in very small samples they can themselves be unreliable. What they emphatically do **not** do is change the point estimates. The coefficient vector is computed from the same normal equations either way; only the reported uncertainty moves. That is exactly why the diagnostic story is so clean: refit with robust standard errors, and if a coefficient's significance evaporates while its value is byte-for-byte identical, you have learned that your original confidence was an artefact of the variance formula, not evidence about the world. ## How to answer in the room Say the headline first - estimates unbiased, inference broken - then justify it by naming which assumption each property depends on. Add that the direction of the standard-error error is not fixed, and finish with the practical protocol: quantify it by comparing conventional against robust standard errors and see whether any conclusion actually changes.
- If the coefficients are unbiased anyway, why is it worth fixing?Because almost every decision made from a regression rests on inference, not on the raw coefficient: is this effect real, how large could it plausibly be, do we ship. Understated standard errors manufacture significance that will not replicate, and overstated ones hide effects you should have acted on. The estimate being fine is cold comfort if the error bar is fiction.
- Are conventional standard errors always too small under heteroscedasticity?No, and asserting it is a common overreach. The direction depends on how the error variance lines up with the predictor values: variance concentrated at extreme predictor values understates the true sampling variability, while variance concentrated in the middle can overstate it. Compare conventional against robust standard errors rather than guessing.
- What does the sandwich standard error estimate that the conventional one assumes away?The middle of the variance expression. Conventional standard errors assume the error covariance matrix is a single variance times the identity, which lets the expression collapse to a simple form. The sandwich estimator instead plugs in the observed squared residuals, so it stays consistent whatever pattern the variances follow, at the cost of being an asymptotic argument.
It is like a scale that weighs correctly on average but whose needle jitters more for heavy objects. Your recorded weights are still right; the tolerance printed on the label is the thing that lies.
saying these in an interview costs you the question
- Claims heteroscedasticity biases the coefficient estimates
- Says the conventional standard errors are always too small
- Believes robust standard errors change the point estimates
- Thinks a significant heteroscedasticity test invalidates the whole model