How does the Breusch-Pagan test check for non-constant error variance in a regression?
answer
- look at the squared residuals
- regress them on the predictors
- the statistic is n times R-squared
- chi-square, one df per auxiliary predictor
basics
~20 sBreusch-Pagan regresses the squared OLS residuals on the model's predictors. Its statistic, n times the auxiliary R-squared, follows a chi-square distribution with one degree of freedom per auxiliary predictor; a small p-value rejects the null of constant error variance.
solid answer
~50 sIt turns the question into a second regression. Fit the model by OLS, take the residuals, square them, and regress those squared residuals on the original predictors - the *auxiliary* regression. If the error variance really is constant, nothing in the predictors should explain the squared residuals, so the auxiliary R-squared should be near zero. The Lagrange-multiplier statistic is `n * R_aux^2`, compared against a chi-square distribution with degrees of freedom equal to the number of predictors in the auxiliary regression, excluding the intercept. On household spending regressed on income, that is a single auxiliary predictor and hence one degree of freedom; a statistic of 20 against a one-degree-of-freedom chi-square is decisively significant, so you reject constant variance. The null is homoscedasticity, so a small p-value is evidence *against* it. The studentised variant is preferred because the original form leans on normally distributed errors.
go deeper
Know that the test exists, that its null is constant error variance, and that a small p-value means you reject constant variance. Getting the direction of the null right is most of the credit at this level.
Walk through the auxiliary regression from memory: square the residuals, regress them on the predictors, form n times the auxiliary R-squared, compare to chi-square with one degree of freedom per auxiliary predictor.
Show the limits: the test only sees variance related to the variables you supply, it saturates in large samples, and a rejection should trigger a magnitude check against robust standard errors rather than a reflex correction.
Be ready to argue whether the team should test at all versus defaulting to robust inference everywhere, and to set the standard for what evidence justifies restructuring a model rather than adjusting its standard errors.
## The idea in one line If the error variance is constant, then nothing you have measured should predict how big an error is. Breusch-Pagan operationalises exactly that sentence. ## The procedure 1. Fit the model of interest by ordinary least squares and keep the residuals `e_i`. 2. Square them. A squared residual is a one-observation estimate of that observation's error variance - noisy, but unbiased up to a constant. 3. Run an **auxiliary regression** of `e_i^2` on a set of variables suspected of driving the variance. The default choice is the original predictors themselves, but any set of candidate drivers works, including the fitted values. 4. Read the auxiliary R-squared and form the statistic `LM = n * R_aux^2`. 5. Compare it against a chi-square distribution with `q` degrees of freedom, where `q` is the number of regressors in the auxiliary regression not counting the intercept. The hypotheses are: null - the error variance does not depend on those variables (homoscedasticity); alternative - it does. A **small p-value rejects constant variance**. Getting this direction backwards is a classic stumble under pressure. ## A worked example Regress household spending on income. There is one predictor, so the auxiliary regression of squared residuals on income has one slope and `q = 1`. Suppose `n = 500` and the auxiliary regression achieves `R_aux^2 = 0.04`. The statistic is `500 * 0.04 = 20`. The 5% critical value of a chi-square with one degree of freedom is about 3.84, so 20 is far into the tail and you reject: the spread of spending genuinely scales with income. Notice how weak an auxiliary R-squared of 0.04 looks and how emphatic the verdict is. That is not an error - squared residuals are extremely noisy, so a small systematic component is meaningful, and with 500 rows it is easy to detect. It also shows exactly why the test alone is not a decision rule. ## The two versions The original formulation builds the statistic as half the explained sum of squares from a regression of `e_i^2 / sigma_hat^2` on the candidate variables, and derives its chi-square distribution assuming the errors are **normally distributed**. If the errors are heavy-tailed, that version over-rejects - it fires on non-normality rather than on unequal variance. The **studentised** (Koenker) version uses the `n * R_aux^2` form given above and does not require normality. It is the version implemented as the default in most modern tooling and the one to name in an interview. If asked why two versions exist, the honest answer is that the original traded robustness for power under a normality assumption that data rarely honours. ## What the test cannot see Breusch-Pagan looks for variance that is a **linear function of the variables you gave it**. Three consequences follow: - If the variance is driven by a variable not in the auxiliary regression, the test is blind to it. Put the suspects in deliberately rather than accepting the default set without thought. - If the variance is U-shaped in a predictor - small at both extremes, large in the middle - a linear auxiliary term captures none of it and the test can pass while the assumption is badly violated. Adding a squared term to the auxiliary regression is the cheap fix, and taken to its logical end that is essentially White's test. - A non-significant result is **not** proof of homoscedasticity. It is failure to detect, which in a small sample is unremarkable. ## The large-sample trap Run this test on a few hundred thousand rows and it will almost always reject, because with enough data any microscopic dependence of variance on a predictor is statistically detectable. Significance answers *is there any* rather than *does it matter*. The decision-relevant follow-up is a magnitude check: compute robust standard errors alongside the conventional ones and see whether the inference you actually care about changes. If a t-statistic moves from 6.1 to 5.9, the rejection is real and irrelevant. This is why many practitioners skip the test in applied work and report heteroscedasticity-robust standard errors unconditionally. Robust standard errors cost very little when the errors happen to be homoscedastic - a modest efficiency loss in finite samples - and protect you when they are not. Testing then becomes something you do to *understand* the data (which variable drives the spread, and does that suggest a missing term or a transformation) rather than to decide whether to correct. ## Answering well Name the auxiliary regression, name the statistic and its degrees of freedom, state the null in the right direction, and finish with the judgement layer: the test tells you the variance is not constant, not whether your conclusions are at risk. That last sentence is what separates a mechanical answer from a practitioner's one.
- Why can the Breusch-Pagan test miss real heteroscedasticity?Because it only detects variance that is a linear function of the variables placed in the auxiliary regression. Variance driven by a variable you left out is invisible, and variance that is U-shaped in a predictor contributes almost nothing to a linear auxiliary slope, so the test can comfortably pass while the assumption is badly broken.
- A Breusch-Pagan test rejects at p below 0.001 on 400,000 rows. Must you act?Not automatically. At that sample size any trivial dependence of variance on a predictor is detectable, so significance answers whether there is any effect, not whether it matters. Compare conventional and robust standard errors on the coefficients you care about; if the conclusions are identical, the rejection is real and inconsequential.
- Is a non-significant Breusch-Pagan result evidence that the errors are homoscedastic?No - it is a failure to detect, which in a modest sample is entirely unremarkable. The test has limited power, especially when the variance depends on something you did not include. Treat it as absence of evidence rather than evidence of absence, and let the cost of being wrong decide whether you use robust standard errors anyway.
saying these in an interview costs you the question
- Regresses the raw residuals instead of the squared residuals
- States the null as heteroscedasticity and reads the p-value backwards
- Uses residual degrees of freedom instead of the auxiliary predictor count
- Treats a non-significant result as proof of constant variance
- Acts on a rejection in a huge sample without checking whether inference changes