What does heteroscedasticity mean for the error terms in a linear regression?
answer
- concerns error spread, not error mean
- one shared variance across rows, or not
- spending spread widens as income rises
- the constant-variance Gauss-Markov condition fails
basics
~20 sHeteroscedasticity means the variance of the regression errors is not the same for every observation: it changes systematically, usually with a predictor. Homoscedasticity, the classical assumption, is the opposite - one common error variance for every row.
solid answer
~50 sHeteroscedasticity means the variance of the error term is not the same for every observation - it changes systematically, typically with a predictor or with the level of the fitted response. Regress household spending on income and you see it plainly: low-income households deviate from the fitted line by small absolute amounts, while high-income households deviate by large ones, so the error variance grows with income. The opposite condition, homoscedasticity, is one of the classical Gauss-Markov assumptions: a single common variance shared by every row. It is a statement about the unobserved population errors, which we judge indirectly from the residuals, either by eye or with a formal test such as Breusch-Pagan or White. It says nothing about the mean of the errors, which is still assumed zero, and nothing about whether errors are correlated with each other - that is a separate assumption.
go deeper
Be ready to state the definition in one sentence and give one concrete example where spread grows with a predictor. Know that homoscedasticity is the assumption and heteroscedasticity is its failure.
Explain that the assumption constrains the variance of the unobserved errors, not their mean or their mutual correlation, and separate it cleanly from non-normality and from dependence between observations.
Show that you know which processes generate it by construction - aggregated group means, counts, proportions, scale-driven outcomes - so you can anticipate it from the data design before fitting anything.
Own the framing that the assumption is about inference quality, not about model correctness, and be able to tell a team when unequal variance is worth engineering around and when it is a rounding error on the decision.
## The assumption being described A linear regression model writes each observation as `y_i = b0 + b1*x_i1 + ... + bk*x_ik + e_i`, where `e_i` is an unobserved error. The classical Gauss-Markov conditions place three separate demands on those errors: their mean is zero given the predictors, they are uncorrelated with one another, and they all share **one common variance**, `Var(e_i) = sigma^2` for every `i`. That third condition is *homoscedasticity* - literally, same scatter. **Heteroscedasticity** is its failure: `Var(e_i) = sigma_i^2`, a quantity that differs across observations. Usually the differences are systematic rather than random - the variance is some function of the predictors, of the fitted level, or of a grouping variable. ## What it is not Three confusions come up constantly, and interviewers probe all of them. - It is **not** about the *mean* of the errors. Non-zero conditional mean is a different (and far more serious) failure - that one really does bias the coefficients. Heteroscedasticity leaves the mean structure alone. - It is **not** about **normality**. Errors can be perfectly normal and wildly heteroscedastic, or non-normal with a rock-steady variance. The two assumptions are independent. - It is **not** about **correlation between errors**. Errors that are correlated across time or across members of a group violate the independence condition, which is a distinct problem with a distinct fix. ## A worked intuition Take household spending regressed on household income. At an income of 20,000 a year, spending is heavily constrained: nearly all of it goes on necessities, so actual spending sits within a narrow band of whatever the fitted line predicts. At an income of 500,000, spending is discretionary: one household saves aggressively, another buys a boat. The *average* relationship may still be perfectly linear - the fitted line can be exactly right at every income level - but the *spread* around it fans out as income rises. That is heteroscedasticity in its most common form: variance increasing with a predictor. ## Why it arises so often Some data-generating processes produce it by construction, not by accident: - **Aggregated rows.** If a row is an average over a group - a city mean computed from 400 households, another from 40,000 - then the variance of that average is inversely proportional to the group size. Unequal groups guarantee unequal variances. - **Bounded or count outcomes.** For a count with a Poisson-like generating process the variance equals the mean, so units with larger expected counts vary more. For a proportion estimated from `n` trials the variance is `p*(1-p)/n`, which depends on both `p` and `n`. - **Scale effects.** Anything measured in money, volume or size tends to vary proportionally rather than absolutely: a 10% deviation is a bigger absolute deviation for a bigger unit. - **Mixed populations.** If the sample pools two subgroups with genuinely different noise levels, the pooled errors are heteroscedastic even if each subgroup on its own is not. ## Population errors versus observed residuals The assumption is about `e_i`, which nobody ever observes. What we have are the fitted residuals `e_hat_i = y_i - y_hat_i`. These are not a clean stand-in: even when the true errors are perfectly homoscedastic, the fitted residuals have slightly unequal variances by construction, because fitting the line uses up information unevenly across the observations. That is exactly why formal tests standardise the residuals before judging them, and why a small visual wobble is not evidence of anything. ## Why anyone cares The short version, which the next question develops: heteroscedasticity does **not** bias the coefficient estimates. Ordinary least squares stays unbiased and consistent as long as the mean function is right. What it damages is the estimated *precision* of those coefficients - the standard errors, and therefore every t-statistic, p-value and confidence interval computed from them. It also costs efficiency: OLS is no longer the minimum-variance linear unbiased estimator once the variances differ. So the practical stakes are entirely about inference. If you only need a fitted line for prediction of the conditional mean, heteroscedasticity is close to harmless for the point predictions - though a single constant residual standard deviation will make prediction *intervals* too narrow in the noisy region and too wide in the quiet one. If you need to say whether a coefficient is real, it matters a great deal.
- Is heteroscedasticity a property of the observed residuals or of the unobserved errors?Of the unobserved population errors. Residuals are only estimates of them, and even under perfect homoscedasticity fitted residuals carry slightly unequal variances by construction, because the fit absorbs different amounts of information at different points in the predictor space. That is why formal tests standardise residuals before judging constant variance.
- Name a data-generating situation that produces heteroscedasticity by construction.Aggregated rows are the cleanest case: if each row is a group mean built from a different number of underlying units, the variance of that mean is inversely proportional to the group size, so unequal groups force unequal variances. Count and proportion outcomes do it too, since their variance is a function of their mean.
- Does heteroscedasticity affect prediction intervals as well as coefficient inference?Yes. A prediction interval built from a single pooled residual standard deviation applies the same width everywhere, so it is too narrow in the high-variance region and too wide in the low-variance one. Point predictions of the conditional mean stay fine; the stated uncertainty around them does not.
Think of a rifle whose grouping is tight at short range and scatters widely at long range. The aim is still true on average at every distance, but how much you can trust a single shot depends on where you are standing.
saying these in an interview costs you the question
- Says heteroscedasticity means the errors are not normally distributed
- Confuses widening spread with a curved mean relationship
- Thinks it means the errors are correlated with each other
- Claims it makes the fitted line systematically wrong