How do you derive the OLS slope and intercept from the normal equations?
answer
- two unknowns, two derivative conditions
- one equation per coefficient, not per row
- cross-products over squared x-deviations
- intercept pinned by the sample means
- matrix form: X'X b equals X'y
basics
~20 sSet both partial derivatives of the squared-residual sum to zero to get the two normal equations. Solving them gives slope = sum of (x - xbar)(y - ybar) over sum of (x - xbar) squared, and intercept = ybar minus slope times xbar.
solid answer
~40 sWrite `SSR(b0, b1) = sum (y_i - b0 - b1*x_i)^2` and differentiate. The derivative with respect to `b0` gives `sum y_i = n*b0 + b1*sum x_i`; the derivative with respect to `b1` gives `sum x_i*y_i = b0*sum x_i + b1*sum x_i^2`. Those two linear equations in two unknowns are the normal equations. Solving them gives `b1hat = sum (x_i - xbar)(y_i - ybar) / sum (x_i - xbar)^2` and `b0hat = ybar - b1hat*xbar`. Because SSR is a convex quadratic — a bowl — the stationary point is the global minimum, not a maximum or a saddle. In matrix form the same derivation collapses to `X'X b = X'y`, one equation per estimated coefficient, with closed-form solution `bhat = (X'X)^-1 X'y`.
go deeper
Be ready to recall the two formulas and apply them to a small table of numbers: slope from cross-products over squared x-deviations, then intercept from the sample means.
You are expected to produce the derivation on a whiteboard — differentiate the squared-error sum, set both partials to zero, state the two normal equations, and solve. Know that there is one equation per coefficient.
Show you understand why the solution exists and is unique: the objective is convex, the stationary point is global, and the fit degenerates precisely when a predictor carries no variation.
Frame the closed form as a design choice with limits. Argue when an exact algebraic solution is the right tool and when the estimation problem should be posed differently — penalised, weighted, or iteratively solved at scale.
## Setting up the minimisation For simple regression with an intercept, the objective is a function of two unknowns: ``` SSR(b0, b1) = sum over i of (y_i - b0 - b1*x_i)^2 ``` This is a smooth quadratic surface in `(b0, b1)`. Minimising it means finding where both partial derivatives vanish. **Derivative with respect to the intercept.** Differentiating term by term and using the chain rule: ``` dSSR/db0 = sum over i of 2*(y_i - b0 - b1*x_i)*(-1) = -2 * sum (y_i - b0 - b1*x_i) ``` Setting this to zero and dividing by -2: ``` sum (y_i - b0 - b1*x_i) = 0 => sum y_i = n*b0 + b1*sum x_i (1) ``` **Derivative with respect to the slope.** The same procedure, but the inner derivative is `-x_i`: ``` dSSR/db1 = -2 * sum x_i*(y_i - b0 - b1*x_i) = 0 => sum x_i*y_i = b0*sum x_i + b1*sum x_i^2 (2) ``` Equations (1) and (2) are the **normal equations**. They are exact conditions, not approximations, and there is one of them per estimated coefficient — never one per observation. ## Solving the system Divide (1) by `n` to get `ybar = b0 + b1*xbar`, so ``` b0hat = ybar - b1hat*xbar ``` Substituting that into (2) and simplifying gives ``` b1hat = (n*sum x_i*y_i - (sum x_i)(sum y_i)) / (n*sum x_i^2 - (sum x_i)^2) ``` which, after centring, is the form most people memorise: ``` b1hat = Sxy / Sxx, where Sxy = sum (x_i - xbar)(y_i - ybar) Sxx = sum (x_i - xbar)^2 ``` The two expressions are algebraically identical; the centred one is easier to reason about, the raw-sums one is easier to evaluate from a table of totals. ## Worked example Five months of advertising spend and sales, both in thousands of dollars: ``` x (ad spend): 1 2 3 4 5 y (sales): 3 8 7 10 12 ``` Totals: `n = 5`, `sum x = 15`, `sum y = 40`, `sum x*y = 3 + 16 + 21 + 40 + 60 = 140`, `sum x^2 = 55`. Raw-sums form: ``` b1hat = (5*140 - 15*40) / (5*55 - 15^2) = (700 - 600) / (275 - 225) = 100 / 50 = 2 b0hat = (40 - 2*15) / 5 = 10 / 5 = 2 ``` Centred form, as a cross-check: `xbar = 3`, `ybar = 8`; x-deviations are -2, -1, 0, 1, 2 and y-deviations are -5, 0, -1, 2, 4, so `Sxy = 10 + 0 + 0 + 2 + 8 = 20` and `Sxx = 4 + 1 + 0 + 1 + 4 = 10`, giving `b1hat = 20/10 = 2` and `b0hat = 8 - 2*3 = 2`. The fitted line is `yhat = 2 + 2x`. Interpretation in the units of the problem: each extra thousand dollars of advertising is associated with two thousand dollars more sales, and the fitted line predicts two thousand dollars of sales at zero advertising. ## Why the stationary point is a minimum A vanishing gradient alone does not prove a minimum. The second derivatives are `d2SSR/db0^2 = 2n`, `d2SSR/db1^2 = 2*sum x_i^2`, and the cross term `2*sum x_i`; the resulting matrix is `2*X'X`, which is never negative in any direction, because for any vector `v` the quantity `v'X'Xv` is a sum of squares. So the surface is a bowl and the unique stationary point sits at its bottom. ## More than one predictor With `k` predictors and an intercept, stack the data into a matrix `X` whose first column is all ones. The same differentiation produces `k+1` equations, written compactly as ``` X'X * bhat = X'y => bhat = (X'X)^-1 * X'y ``` Each row of that system says the same thing the simple case said: the residual vector must have zero inner product with the corresponding column of `X`. Nothing about the derivation changes — only the bookkeeping. ## The degenerate case If every predictor value is identical, `Sxx = 0` and the slope formula divides by zero. That is not a numerical accident: with no variation in x there is no way to attribute variation in y to it, so the slope is simply not identified by the data. The intercept is still estimable — it is just `ybar`. ## Common derivation slips Forgetting the factor of `-1` (or `-x_i`) from the chain rule inverts a sign; writing the intercept as `ybar + b1hat*xbar` puts the line on the wrong side of the point of means; and solving only the slope equation while treating the intercept as an afterthought loses the constraint that pins the fit vertically.
- Why is the stationary point of the least-squares objective guaranteed to be a minimum?Because the objective is a convex quadratic in the coefficients. Its matrix of second derivatives is `2*X'X`, and `v'X'Xv` is a sum of squares for any direction `v`, so the surface curves upward everywhere — a bowl. A point where both partial derivatives vanish can only be the bottom of that bowl, never a maximum or a saddle.
- What happens to the slope formula if every predictor value in the sample is the same?The denominator `sum (x_i - xbar)^2` is zero and the slope is undefined. There is no statistical content to recover: with no variation in x, the data contain no information about how y responds to x. The intercept alone is still estimable and equals the sample mean of y.
- How does the derivation change when there are several predictors?It does not change in kind, only in count. You differentiate with respect to each coefficient and get one normal equation per coefficient, stacked as `X'X bhat = X'y`. Each equation states that the residual vector has zero inner product with one column of the predictor matrix, exactly as `sum x_i*e_i = 0` did in the simple case.
saying these in an interview costs you the question
- Writes the intercept as ybar plus slope times xbar
- Claims each observation contributes its own normal equation
- Drops the chain-rule factor and flips a sign in the derivative
- Solves only the slope equation and ignores the intercept condition
- Treats the normal equations as an approximation rather than an exact condition