What is the residual standard error in a fitted regression, and what are its degrees of freedom?
answer
- typical residual size, not a percentage
- you spent data fitting coefficients
- one lost dimension per estimated coefficient
- n minus k minus 1, intercept included
- estimated sigma means t, not the normal value
basics
~20 sThe residual standard error estimates typical error size around a fitted line: the root of the residual sum of squares over n minus the number of estimated coefficients. With 4 predictors plus an intercept, that divisor is n - 5.
solid answer
~50 sThe residual standard error is `s = sqrt( SSE / (n - k - 1) )`, where `SSE` is the sum of squared residuals, `n` the number of rows and `k` the number of predictors — so the divisor counts the intercept too. It estimates `sigma`, the error standard deviation, and is the scale factor in every regression interval: change `s` and every prediction interval scales with it. The degrees of freedom are `n - k - 1` because each estimated coefficient uses up one dimension in which the residuals could have varied; dividing by `n` instead would bias the estimate downward. Because `sigma` is estimated rather than known, the interval uses a `t` multiplier on those same degrees of freedom. With `n = 20` and 4 predictors, `df = 15` and the 95% multiplier is about `2.13` instead of `1.96` — roughly 9% wider than a naive normal-based interval.
go deeper
Recall that this quantity is the typical size of a prediction miss, expressed in the units of the response, and that it comes from the squared residuals rather than from a correlation.
Be ready to write s = sqrt(SSE / (n - k - 1)), count the intercept in the degrees of freedom, and explain why the divisor is not n. Connect it to the t multiplier used in the interval.
Show judgment about when the number lies: non-constant residual spread makes one global value misleading, and a small model gives both a noisier estimate and a fatter multiplier. Note that adding weak predictors can raise it.
Frame this as the quantity that bounds what the model can promise. Be ready to argue whether a reported error scale is small enough for the decision it feeds, and to push back on comparisons of this number across models with different responses.
## What the quantity is Fit a regression and you get a fitted value for every row. The residual for row `i` is `e_i = y_i - y_hat_i`, the vertical distance from the observed point to the fitted surface. The residual sum of squares is `SSE = sum of e_i^2`. The residual standard error is `s = sqrt( SSE / (n - k - 1) )` with `n` rows and `k` predictors. It lives in the units of the response: for a delivery-time model it is a number of minutes, and it answers "how far off is a typical prediction?" A residual standard error of 7.5 minutes means individual orders scatter about 7.5 minutes around the fitted line, roughly speaking. It is an *estimate* of a parameter, `sigma`, the standard deviation of the true error term. Keeping that distinction straight matters: `sigma` is unknown and fixed; `s` is computed from your sample and would come out differently on another sample. ## Why the divisor is not n Residuals are not the errors. The errors are the distances to the *true* line; the residuals are the distances to the *fitted* line, and the fitted line was chosen to make exactly those distances as small as possible. So residuals are systematically a little smaller than the errors they stand in for, and dividing `SSE` by `n` gives an estimate of `sigma^2` that is biased downward. The correction is to divide by the residual degrees of freedom, `n - k - 1`. The intuition: fitting an intercept plus `k` slopes imposes `k + 1` linear constraints on the residual vector, so it is free to vary in only `n - k - 1` independent directions. Dividing by that count makes `s^2` an unbiased estimator of `sigma^2`. Count the intercept. A model with 4 predictors on 20 rows has `df = 20 - 4 - 1 = 15`, not 16. Forgetting the intercept is the single most common slip on this question. ## Why the interval uses t rather than a normal multiplier If you somehow knew `sigma`, a 95% interval would use the normal multiplier `1.96`. You do not know it — you plugged in `s`, itself a noisy estimate. The `t` distribution on `n - k - 1` degrees of freedom is exactly the correction for that extra uncertainty: it has heavier tails, so its critical value is larger, and the interval is wider by the right amount. How much wider depends on the degrees of freedom: - `df = 15`: the 95% multiplier is about `2.13`, roughly 9% wider than `1.96`. - `df = 30`: about `2.04`. - `df = 200`: about `1.97`, essentially the normal value. So in a 20-row model the correction is real and worth making; in a 20,000-row model it is invisible. The practical rule is that `t` is never wrong — it converges to the normal multiplier as degrees of freedom grow — so there is no reason to reach for the normal value in a small model just to save arithmetic. ## How s propagates into prediction intervals Every regression interval has the shape `y_hat +/- t * s * (something)`. For a new case at `x0` in simple regression the something is `sqrt( 1 + 1/n + (x0 - xbar)^2 / Sxx )`. Two consequences: - **`s` sets the scale.** Halve the residual standard error and every interval halves. It is the single number that most controls how useful your predictions are. - **`s` has its own uncertainty.** In a small model, `s` is estimated from few residuals and can be well off. The `t` multiplier accounts for that, which is why small-sample intervals are wide in two ways at once: a larger multiplier and a noisier scale. ## The trap: more predictors does not always mean smaller s Adding a predictor to an OLS model can never increase `SSE` — the fit has more freedom, so the residual sum of squares stays the same or falls. But the divisor `n - k - 1` falls by one at the same time. If the new predictor explains little, the numerator barely moves while the denominator shrinks, and `s` goes **up**. This is precisely why `s` (and adjusted measures built from it) behaves more honestly than raw `R^2`, which cannot decrease when you add predictors. A candidate who says "more predictors always lowers the residual standard error" is confusing it with the unadjusted fit statistic. ## Reading it correctly A few honest caveats to have ready: - `s` describes typical scatter only if that scatter is roughly constant. If the residual spread grows with the fitted value, one global `s` overstates the noise at the low end and understates it at the high end, and every interval built from it inherits that error. - `s` is not a goodness-of-fit percentage. It is an absolute quantity in response units, so it is comparable across models of the same response but not across different responses. - `s` says nothing about whether the model form is right. A badly curved relationship fitted with a straight line can still yield a modest `s` while being systematically wrong in places.
- Why not simply divide the residual sum of squares by n?Because residuals are distances to the fitted line, and the fit was chosen to minimise exactly those distances, so they are systematically smaller than the true errors. Dividing by `n` therefore underestimates `sigma^2`. Dividing by `n - k - 1`, the number of directions in which the residuals are still free to vary, makes the estimate unbiased.
- How much does the 95% multiplier change between 15 and 200 degrees of freedom?It moves from about `2.13` down to about `1.97`, converging on the normal value `1.96`. So the correction is worth roughly 9% of interval width in a 20-row model and is negligible in a large one. Since `t` converges to the normal multiplier anyway, using `t` everywhere costs nothing.
- Does adding a predictor always reduce the residual standard error?No. The residual sum of squares cannot increase when a predictor is added, but the degrees of freedom drop by one. If the new predictor explains almost nothing, the divisor shrinks faster than the numerator and `s` rises. That behaviour is exactly what makes it more honest than unadjusted goodness of fit.
saying these in an interview costs you the question
- Divides the residual sum of squares by n rather than the degrees of freedom
- Confuses residual standard error with a goodness-of-fit percentage
- Forgets the intercept when counting estimated coefficients
- Uses a normal multiplier in a model fitted on 20 rows
- Claims extra predictors always lower the residual standard error