Why do OLS residuals sum to exactly zero when the model includes an intercept?
answer
- look at the first-order conditions
- the constant column of ones
- orthogonality, one condition per coefficient
- the intercept can always absorb an offset
- no constant term, no zero-sum guarantee
basics
~20 sBecause the intercept's normal equation is literally the condition that the residuals sum to zero. Fitting a constant term forces that, which also makes the fitted line pass through the point of sample means. Drop the intercept and the guarantee disappears.
solid answer
~50 sDifferentiating the squared-error sum with respect to the intercept gives `-2 * sum (y_i - b0 - b1*x_i) = 0`, that is `sum e_i = 0`. So the zero sum is an algebraic identity produced by including a constant column, not evidence about the data. The slope's normal equation gives `sum x_i*e_i = 0`, so the residual vector is also orthogonal to every predictor — and hence to the fitted values. Two consequences follow: the mean of the fitted values equals the mean of y, and the fitted line passes through `(xbar, ybar)`. In regression through the origin there is no constant column, only `sum x_i*e_i = 0` holds, and the residuals generally sum to something non-zero. Crucially, `sum e_i = 0` holds for *any* dataset with an intercept, so it says nothing about whether the model fits.
go deeper
Recall the fact and its cause in one sentence: the intercept's own first-order condition is sum of residuals = 0, so the fitted line passes through the point of sample means.
Be ready to derive it — show the partial derivative with respect to the intercept, and state the companion condition that residuals are orthogonal to each predictor.
Demonstrate the diagnostic judgment: because the zero sum is mechanical, use residual plots for shape rather than level, and explain what breaks when someone forces a fit through the origin.
Own the modelling call on whether an intercept belongs at all. Forcing a line through zero must be justified by domain theory, and you should be able to explain the cost — a biased slope and summary statistics that no longer compare across models.
## The two normal equations, read as orthogonality conditions For `y_i = b0 + b1*x_i + e_i`, setting the partial derivatives of the residual sum of squares to zero produces ``` sum e_i = 0 (from the intercept) sum x_i * e_i = 0 (from the slope) ``` where `e_i = y_i - b0hat - b1hat*x_i` is the residual. Both are statements that the residual vector is **orthogonal** — has zero inner product — with a column of the predictor matrix. The first column of that matrix is the column of ones that represents the intercept, and taking an inner product with a column of ones is exactly summing. That is the whole answer: residuals sum to zero because the model contains a constant term whose first-order condition says so. The deeper intuition: the intercept is a free vertical dial. If the residuals averaged to some non-zero amount `c`, you could shift the intercept by `c` and reduce the residual sum of squares — so at the minimum, the average residual must be zero. A quantity that could still be improved cannot be at the optimum. ## Worked check Use five months of advertising spend and sales in thousands of dollars: ``` x: 1 2 3 4 5 y: 3 8 7 10 12 ``` The OLS fit is `yhat = 2 + 2x`, so fitted values are 4, 6, 8, 10, 12 and residuals are -1, +2, -1, 0, 0. - `sum e_i = -1 + 2 - 1 + 0 + 0 = 0`, as the intercept condition requires. - `sum x_i*e_i = 1*(-1) + 2*(2) + 3*(-1) + 4*(0) + 5*(0) = -1 + 4 - 3 = 0`, as the slope condition requires. Both conditions hold exactly, to the last decimal, because they are what defined the estimates in the first place. ## The consequences worth naming **The fitted values have the same mean as the observations.** Since `y_i = yhat_i + e_i` and the residuals sum to zero, `sum y_i = sum yhat_i`, so `ybar = mean of yhat`. In the example, both are 8. **The fitted line passes through the point of sample means.** Dividing the intercept condition by `n` gives `ybar = b0hat + b1hat*xbar`, so `(xbar, ybar)` lies on the line. In the example `(3, 8)` satisfies `8 = 2 + 2*3`. **Residuals are uncorrelated with the fitted values.** Fitted values are a linear combination of the predictor columns, and the residuals are orthogonal to each of those columns, so `sum yhat_i * e_i = 0`. Combined with the zero mean of residuals, the sample correlation between residuals and fitted values is exactly zero. This is why a residual-versus-fitted plot has no *linear* trend by construction — anything you see in it is curvature, changing spread, or outliers, never slope. **In multiple regression, the same holds column by column.** Every estimated coefficient contributes one orthogonality condition, so the residuals are orthogonal to every predictor in the model at the same time. ## Regression through the origin Drop the intercept and fit `y_i = b1*x_i + e_i`. Now there is only one normal equation, `sum x_i*e_i = 0`, giving `b1hat = sum x_i*y_i / sum x_i^2`. There is no condition forcing the residuals to sum to zero, and in general they do not. On the same five points: `sum x*y = 140` and `sum x^2 = 55`, so `b1hat = 140/55 = 28/11`, about 2.545. The residuals then sum to `sum y - b1hat * sum x = 40 - (28/11)*15 = 40 - 420/11 = 20/11`, roughly 1.82 — clearly not zero. The consequences above all fail with it: the mean fitted value no longer matches ybar, and the line no longer passes through `(xbar, ybar)` (it passes through the origin instead, by construction). The usual decomposition of total variation around the mean also stops being valid, which is why summary statistics computed the ordinary way are not comparable across intercept and no-intercept fits. This is one of the strongest arguments for keeping an intercept unless theory genuinely forces the line through zero: without it, a systematic average offset in y has nowhere to go except into the slope, biasing it. ## The trap: what the zero sum does not mean Because `sum e_i = 0` holds for any data you feed in — random noise included — it carries no evidence about model quality. Three specific misreadings to avoid: 1. *It is not caused by the assumption that the errors have mean zero.* The **errors** are unobservable, and their mean-zero property is an assumption about the population. The **residuals** sum to zero as arithmetic. The two are separate statements that happen to rhyme. 2. *It does not imply the linear form is right.* A strongly curved relationship fitted with a straight line still has residuals summing to zero, with a clear U-shape when plotted against the predictor. 3. *It does not imply the residuals are independent, identically distributed, or well behaved.* In fact the zero-sum and orthogonality conditions mean the residuals satisfy `k+1` linear constraints, so they are mildly dependent on each other by construction even when the underlying errors are independent.
- Why does the OLS line always pass through the point of sample means?Divide the intercept's normal equation by n: it becomes `ybar = b0hat + b1hat*xbar`, which is exactly the statement that `(xbar, ybar)` lies on the fitted line. Equivalently, residuals summing to zero means the fitted values have the same mean as the observations. This is another reason it fails for a no-intercept fit, which is pinned to the origin instead.
- What does the second normal equation say about the residuals?It says `sum x_i*e_i = 0` — the residual vector is orthogonal to the predictor. In multiple regression this holds for every predictor column simultaneously, and since the fitted values are a linear combination of those columns, the residuals are orthogonal to the fitted values too. Any apparent linear trend in a residual-versus-fitted plot is therefore impossible by construction.
- If residuals always average zero, what can a residual plot still reveal?Everything except the average. Curvature against a predictor exposes a wrong functional form; a fanning or funnel shape exposes non-constant error variance; isolated far-off points expose outliers or high-influence observations; and structure against time or group exposes correlated errors. The zero mean is mechanical, so read the shape, never the level.
The intercept is a free vertical dial on the fitted line. If the residuals averaged above zero, turning the dial up would reduce the squared error, so at the best fit the dial has already used up any average offset.
saying these in an interview costs you the question
- Says residuals sum to zero because the errors have mean zero
- Treats a zero residual sum as evidence the linear model fits
- Assumes residuals still sum to zero without an intercept term
- Confuses residuals summing to zero with residuals being independent
- Claims it holds only when the data are approximately linear