skip to content

With 5,000 rows, a regression's residuals are clearly non-normal in the Q-Q plot. Are the coefficient p-values still usable?

level: seniorimportance: should knowfreq 42%

answer

  1. which assumption, for which output
  2. normality is the least load-bearing one
  3. estimator's sampling distribution, not the errors
  4. estimation error shrinks, the noise does not
  5. prediction interval keeps the error's shape

basics

~20 s

Usually yes. With thousands of observations the sampling distribution of a least-squares coefficient is close to normal whatever the error distribution looks like, so its p-values and confidence intervals are approximately right. Prediction intervals for individual observations get no such protection.

solid answer

~50 s

Normality of the errors is the least load-bearing linear-model assumption for coefficient inference. Unbiasedness needs the mean function to be right, not normal errors, and the usual standard-error formula needs constant variance and independence, again not normality. Normality only buys *exact* t and F distributions in small samples; with 5,000 rows the estimator's sampling distribution is approximately normal anyway, provided the errors have finite variance and the other assumptions hold, so the p-values are fine to a good approximation. Two caveats. Heavy tails make least squares inefficient, with a few rows dominating the fit. And a prediction interval for one new observation inherits the error distribution's shape directly — its width is driven by the error standard deviation, which does not shrink with n — so at n = 15 with skewed errors a nominal 95 percent interval will not cover 95 percent.

go deeper

for a junior

Know that the normality assumption in linear regression is about the errors, checked on the residuals, and that it matters much less than the model having the right shape. Do not claim regression is invalid the moment residuals look non-normal.

for a middle

Explain what normality actually buys: exact small-sample t and F results, and correctly shaped prediction intervals. Be able to say that unbiasedness and the standard-error formula do not depend on it.

for a senior

Show you separate the outputs: coefficient inference at large n is approximately safe, prediction intervals are not, and heavy tails mean an inefficient fit carried by a few rows. Say what you would investigate rather than declaring the model fine.

for a principal

Decide how much a distributional departure should change what the organisation does. Weigh the cost of respecifying against the decisions at stake, and set the norm that diagnostics are judged by consequence rather than run as ritual.

## Rank the assumptions by what they damage A linear model leans on four things, and they are not equally important: 1. **The mean function is right.** If it is not, the coefficients are biased and predictions are systematically wrong by region. Nothing about sample size repairs this. This is the assumption to worry about first. 2. **The errors are independent.** Violating it usually leaves the estimate roughly centred but makes the reported standard errors badly wrong, often far too small, which is worse than it sounds because you get confident wrong answers. 3. **The error variance is constant.** Violating it leaves the estimate unbiased but makes the usual standard-error formula the wrong formula. 4. **The errors are normal.** This one buys exactness in small samples and correct prediction intervals; it is not needed for the estimate to be unbiased, and at large n it is not needed for approximately correct coefficient inference either. The classical optimality result for least squares — that among linear unbiased estimators it has the smallest variance — requires the first three and says nothing about normality. That is the fact to reach for when someone insists all regression inference needs normal errors. ## Why n = 5,000 protects the coefficients A least-squares coefficient is a weighted sum of the outcomes: a linear combination of many independent contributions. Large-sample theory says such a combination has an approximately normal sampling distribution as the sample grows, as long as the errors have finite variance, are independent, and no small handful of observations carries most of the weight in the design. So the object that has to be normal for a p-value to be right is not the error distribution but the **sampling distribution of the estimator**, and at n = 5,000 the second is close to normal even when the first is visibly not. The t-distribution used for the test converges to the normal at those degrees of freedom too, so the arithmetic barely changes. The p-value is approximate rather than exact, but the approximation error is typically negligible next to everything else that is uncertain about the model. The conditions matter, though. 'No small handful of observations carries most of the weight' is a real requirement: with a very lopsided predictor distribution, a few rows can dominate, and the effective sample size for that coefficient is much smaller than 5,000. Extremely heavy tails, where variance is barely finite, also slow the approximation down. ## Why prediction intervals are not protected There are two intervals people conflate. A **confidence interval for the mean response** at a given set of predictor values expresses uncertainty about where the regression line sits. Its width is driven by the uncertainty in the estimated coefficients, which shrinks roughly like `1 / sqrt(n)`. Large n makes it narrow and approximately correct. A **prediction interval for a single new observation** expresses uncertainty about where one future point will land. Its width has two parts: the same shrinking estimation uncertainty, plus the irreducible spread of a single error, of size about the error standard deviation `sigma`. That second part does not shrink at all. As n grows, the interval converges to roughly `yhat` plus or minus a multiple of `sigma` — and which multiple gives 95 percent coverage depends entirely on the **shape** of the error distribution. If the errors are right-skewed, the symmetric normal-based interval is wrong in a specific way: it is too wide on the low side and too short on the high side, so the actual coverage is not 95 percent and, worse, the misses are asymmetric. This is exactly the situation in which practitioners are surprised, because they know 'large n fixes normality' as a slogan and apply it to the wrong quantity. ## The small-sample case At n = 15 the picture reverses. The estimator's sampling distribution is not guaranteed to be close to normal, so the t-based p-value and confidence interval are only exact if the errors really are normal. With visible skew or heavy tails at that sample size, both the coefficient inference and the prediction interval are suspect. Sensible responses are to model the response on a scale where the errors behave, to use an interval procedure that does not assume a normal error shape, or to be honest that the model supports a direction rather than a calibrated number. ## Non-normality as a symptom Even when the p-values survive, a clearly non-normal residual picture at large n is information worth using. - **Strong right skew** often means the response is multiplicative or bounded below, and the model would be better on a different scale. - **Heavy tails** often mean a mixture of subpopulations treated as one, or a rare regime the model has no term for. Least squares squares the errors, so those rows quietly buy a lot of influence over the fit. - **Bimodality** in the residuals almost always means a missing categorical distinction. So the professional answer is not just 'it is fine at this sample size'. It is: the coefficient inference is approximately valid, the prediction intervals are not, and the picture is telling me something about the model I should investigate before I ship it.

  • Why does a large sample not rescue a prediction interval?
    Because it covers one future observation, not an average. Its width is the estimation uncertainty, which shrinks with n, plus the spread of a single error, which never shrinks. In the limit the interval is the fitted value plus or minus a multiple of the error standard deviation, and the right multiple depends on the error distribution's actual shape.
  • Which assumption would you worry about far more than normality, and why?
    Whether the mean function is right, because that biases the estimates themselves and no sample size fixes it. Independence comes next: correlated errors leave the estimate roughly centred but can make standard errors badly too small, which produces confident wrong conclusions rather than merely imprecise ones.
  • What might clearly non-normal residuals at n = 5,000 be telling you, even if the p-values are fine?
    Often that the response belongs on a different scale, that a subpopulation is missing from the model, or that the process has a heavy-tailed regime. Heavy tails also mean least squares is inefficient and a small number of rows carry the fit, so the estimate is noisier and more fragile than the standard error suggests.
  • Is there a case where large n makes the normality question harder rather than easier?
    Yes, in the sense that at large n any departure becomes clearly visible, so a plot that would have looked fine at n = 50 now shows an unmistakable bend. The right response is to judge the size of the departure and what it affects, rather than to treat visible non-normality as automatically disqualifying.

Averaging smooths out the shape of what you average, but predicting a single future draw means facing that shape head-on, however much history you have.

saying these in an interview costs you the question

  • Says all regression inference requires normally distributed errors
  • Checks normality of the response instead of the residuals
  • Assumes a large sample also fixes prediction intervals
  • Treats any visible non-normality at large n as fatal
  • Confuses a confidence interval for the mean with a prediction interval

context