skip to content

How does White's test for heteroscedasticity differ from the Breusch-Pagan test?

level: middleimportance: nice to knowfreq 30%

answer

  1. same recipe, richer right-hand side
  2. adds squares and cross-products
  3. no functional form assumed for the variance
  4. term count grows quadratically with predictors

basics

~20 s

White's test uses the same squared-residual auxiliary regression as Breusch-Pagan but adds the squares and all cross-products of the predictors, so it catches variance patterns of any shape. The price is many more degrees of freedom and lower power.

solid answer

~40 s

Same recipe, richer right-hand side. Both regress the squared OLS residuals on a set of variables and compare `n * R_aux^2` to a chi-square distribution. Breusch-Pagan uses the predictors themselves, so it detects variance that moves roughly linearly with them. White's test adds every predictor's square and every pairwise cross-product, which lets it pick up non-monotone patterns - variance that is large at both extremes of a predictor, say - and interaction-driven variance that Breusch-Pagan is blind to. Two costs follow. The auxiliary term count grows quadratically: with `k` predictors it is `k*(k+3)/2` terms, so 10 predictors means 65 degrees of freedom, which drains power in anything but a large sample. And because a curved mean relationship also leaves structure in the squared residuals, a rejection may indicate functional-form misspecification rather than unequal variance.

go deeper

for a junior

It is enough to know the two tests exist, share the squared-residual auxiliary regression idea, and both use constant variance as their null hypothesis.

for a middle

Explain the actual difference in the auxiliary regression - predictors alone versus predictors plus squares plus cross-products - and why the wider set can detect non-monotone and interaction-driven variance.

for a senior

Show that you weigh the degrees-of-freedom cost against the sample size, and raise unprompted that a rejection may be pointing at functional-form misspecification rather than unequal variance.

for a principal

Own the judgement that neither test changes a number in the model, and be able to argue for a default reporting policy rather than a per-model testing ritual that invites data-dredging.

## Both tests share a skeleton Square the OLS residuals, regress them on some set of variables, and ask whether that auxiliary regression explains anything. The statistic is `n` times the auxiliary R-squared, referred to a chi-square distribution whose degrees of freedom equal the number of auxiliary regressors excluding the intercept. The null in both cases is constant error variance, and a small p-value rejects it. The tests differ only in what goes on the right-hand side of the auxiliary regression - which is to say, in what shapes of heteroscedasticity they are capable of noticing. ## What each one looks for **Breusch-Pagan** puts the original predictors there. The implicit variance model is roughly linear: variance rising or falling steadily with a predictor. That covers the most common real pattern - spread growing with the scale of a driver - and it is cheap, needing only one degree of freedom per suspect. **White** puts the predictors, their squares, and all pairwise cross-products there. Because a smooth function of the predictors can be approximated by such a second-order expansion, the test does not commit to a functional form for the variance at all. It therefore catches: - **non-monotone variance** - large at both ends of a predictor's range and small in the middle, a pattern to which a linear auxiliary term is essentially blind because the positive and negative contributions cancel; - **interaction-driven variance** - noise that is high only for the combination of two conditions, invisible to either main effect alone. ## The degrees-of-freedom price With `k` original predictors, White's auxiliary regression has `k` linear terms, `k` squared terms and `k*(k-1)/2` cross-products, giving `k*(k+3)/2` regressors in total. For `k = 2` that is 5. For `k = 5` it is 20. For `k = 10` it is 65. Two problems arise as that number grows. First, spreading the test across 65 degrees of freedom dilutes it: if the real signal lives in one or two of those terms, the chi-square threshold for 65 degrees of freedom is far away and you fail to reject something that a targeted test would have caught easily. Second, you may not have the rows to fit the auxiliary regression stably at all - 65 terms on 300 observations is a fragile fit whose R-squared is inflated by chance. This makes the choice sample-dependent rather than a matter of one test being better. With few predictors and plenty of rows, White's generality is nearly free. With many predictors and a modest sample, Breusch-Pagan aimed at two or three genuine suspects is the more informative test. A middle path is to run Breusch-Pagan with hand-chosen terms - a specific predictor plus its square, for instance - which buys the shape flexibility exactly where you suspect it and spends only two degrees of freedom. ## Ambiguity of a rejection This is the subtlest point and the one worth raising unprompted. A squared residual can be inflated for two different reasons: the observation genuinely comes from a high-variance region, or the mean function is wrong there and the residual is absorbing systematic misfit. If the true relationship is curved but you fit a straight line, the residuals are systematically large at both ends and small in the middle - and that is a pattern in the *squares* that correlates with the squared predictor, which is precisely a term in White's auxiliary regression. So White's test is really a joint test against heteroscedasticity *and* certain kinds of specification error. A rejection tells you something is wrong with the second moments of your residuals; it does not tell you which of the two problems you have. The correct response is not to reach immediately for robust standard errors, because robust standard errors do nothing whatsoever about a mis-specified mean function - they would just give you correctly measured uncertainty around a wrong line. Check the functional form first, and only then treat the remaining unequal variance. Breusch-Pagan run on the linear predictors alone is somewhat less exposed to this, since it lacks the squared terms that curvature most readily loads on - though it is not immune either. ## When to use which - Concrete hypothesis about which variable drives the spread: **Breusch-Pagan** with that variable, possibly plus its square. - No hypothesis, few predictors, ample data: **White**, accepting the wide net. - Many predictors, modest data: neither test is very informative - reporting robust standard errors is the pragmatic route, since the correction costs little when the errors happen to be well behaved. In all three cases the test is diagnostic rather than decisive. Neither test changes a single number in the model, and neither answers the question you actually have, which is whether your conclusions survive a variance-agnostic standard error.

  • Why is a rejection from White's test ambiguous?
    Because a squared residual grows either from genuine high variance or from systematic misfit. A curved true relationship fitted with a straight line leaves large residuals at both ends, which loads directly on the squared predictor terms in White's auxiliary regression. So the test is joint against unequal variance and certain specification errors, and a rejection does not say which one you have.
  • With 10 predictors and 300 rows, would you run White's test?
    Probably not. Ten predictors generate 65 auxiliary terms, which on 300 rows is a fragile fit spread across far too many degrees of freedom to have useful power. Either run Breusch-Pagan aimed at two or three specific suspects, or skip testing and report heteroscedasticity-robust standard errors, which cost little when the errors turn out to be well behaved.
  • How would you catch U-shaped variance without paying White's full price?
    Run the Breusch-Pagan auxiliary regression with a hand-picked term set: the suspect predictor plus its square. That gives the curvature flexibility exactly where you suspect it for two degrees of freedom instead of the full quadratic expansion across every predictor, and keeps the test focused enough to retain power.

saying these in an interview costs you the question

  • Thinks White's test assumes a specific functional form for the variance
  • Ignores that a rejection may signal a mis-specified mean function
  • Runs the full quadratic expansion with many predictors on a small sample
  • Believes running the test corrects the standard errors

context