skip to content

Normal Equations and Gauss-Markov

The linear model writes an outcome as a weighted sum of predictors plus error, and least squares solves the normal equations for those weights. Interviewers ask why squared error, not absolute.

on this pageshow

questions

5

What quantity does ordinary least squares minimise when fitting a linear model?

level: juniorimportance: must knowfreq 80%

answer

  1. vertical misses, not perpendicular ones
  2. signed residuals would cancel
  3. squares, summed over observations
  4. observed minus fitted, then squared
  5. each y splits into fitted plus residual

basics

~20 s

Ordinary least squares picks the coefficients that make the sum of squared residuals as small as possible. A residual is an observed value minus the value the fitted line predicts for it, so the misses are vertical.

solid answer

~50 s

The model is `y = b0 + b1*x1 + ... + e`, and OLS chooses the coefficient estimates that minimise `SSR = sum over i of (y_i - yhat_i)^2`, where `yhat_i` is the fitted value and `e_i = y_i - yhat_i` is the residual. Every observation splits exactly as `y_i = yhat_i + e_i`, so OLS is the split that pushes as much of the variation as possible into the fitted part. The distances are vertical: the model treats the predictors as given and puts the randomness in y, so a miss is measured in the units of y. Squaring makes the objective smooth, which is what produces a closed-form solution, and it penalises a miss of 4 sixteen times as heavily as a miss of 1 — one far-off point can visibly pull the line.

go deeper

for a junior

Be ready to state the objective in one line — smallest total squared vertical distance from the points to the line — and to compute a residual as observed minus fitted for a single point.

for a middle

Explain why squares rather than signed values or absolute values: signed residuals cancel, absolute values have no closed form, squares give a smooth convex objective and a solvable system of equations.

for a senior

Show that you know the cost of the squared objective in real data: quadratic penalties make a single extreme observation influential, so pair the fit with a look at which observations are driving it.

for a principal

Own the framing decision. Squared error encodes a specific loss — symmetric, quadratic, mean-seeking — and if the business cost of over- and under-prediction is asymmetric or heavy-tailed, argue for a different objective rather than defaulting to OLS.

## The model, and what fitting means A linear regression model says that an outcome `y_i` for observation `i` is a straight-line function of one or more predictors plus an unobservable disturbance: ``` y_i = b0 + b1*x_i1 + ... + bk*x_ik + e_i ``` Here `b0` is the intercept, `b1 ... bk` are the true (unknown, fixed) coefficients, and `e_i` is the **error** — the amount by which observation `i` sits off the true line. The error is never observable, because the true coefficients are never known. Fitting means choosing numerical **estimates** `b0hat ... bkhat` from the data. Those estimates give each observation a **fitted value** ``` yhat_i = b0hat + b1hat*x_i1 + ... + bkhat*x_ik ``` and a **residual** ``` e_i = y_i - yhat_i ``` The residual is computable; the error is not. Keeping the two apart is the single most common vocabulary slip in this area — a residual is a deviation from the *estimated* line, an error is a deviation from the *true* line. ## The objective Ordinary least squares is defined by its objective function, the residual sum of squares: ``` SSR(b0, ..., bk) = sum over i of (y_i - (b0 + b1*x_i1 + ... + bk*x_ik))^2 ``` OLS chooses the coefficient values that make this number smallest. Nothing else about OLS is a matter of choice: given the data and the chosen predictors, the estimates are whatever minimises this sum. Two details in that definition matter. First, the distances being summed are **vertical** — differences in y at a fixed x — not perpendicular distances from the point to the line. Minimising perpendicular distance is a different method (orthogonal or total least squares) that assumes both variables are measured with error and, unlike OLS, changes its answer if you rescale x from dollars to thousands of dollars. Second, the quantity minimised is the sum of *squares*, not of signed residuals: signed residuals cancel, so their sum can be zero for many badly-fitting lines and cannot single out a best one. ## The decomposition on a whiteboard-sized example The identity `y_i = yhat_i + e_i` holds for *any* line you draw, good or bad; OLS is the choice that makes the second piece as small as possible in squared total. Take five months of advertising spend and sales, both in thousands of dollars: ``` x (ad spend): 1 2 3 4 5 y (sales): 3 8 7 10 12 ``` The OLS fit for this data is `yhat = 2 + 2x`. That gives fitted values 4, 6, 8, 10, 12 and residuals ``` e: -1 +2 -1 0 0 ``` and `SSR = 1 + 4 + 1 + 0 + 0 = 6`. Check that no nearby line does better. Take `yhat = 2.5 + 1.9x`: fitted values 4.4, 6.3, 8.2, 10.1, 12.0, residuals -1.4, +1.7, -1.2, -0.1, 0.0, and `SSR = 1.96 + 2.89 + 1.44 + 0.01 + 0 = 6.30`. Larger, as it must be. Every observation still decomposes exactly — 8 is 6 plus 2 under the OLS fit, and 8 is 6.3 plus 1.7 under the alternative — but the OLS split leaves the smallest squared remainder. ## Why squares, and not something else The practical reason is smoothness. The squared objective is a differentiable, bowl-shaped (convex) function of the coefficients, so its minimum can be found by setting partial derivatives to zero, which yields a system of linear equations with a closed-form solution. Absolute deviations give a perfectly sensible alternative estimator (least absolute deviations, which is more resistant to outliers and behaves like a conditional median rather than a conditional mean), but the objective has a kink at every zero residual and no closed form. The deeper reason is that squares come with an optimality guarantee: under a standard set of assumptions about the errors, the least-squares estimator has the smallest variance of any linear unbiased estimator — the Gauss-Markov result. So squaring is not merely a trick for removing minus signs. The cost is sensitivity. Because the penalty grows quadratically, one observation far from the pattern contributes disproportionately and can tilt the fitted line noticeably, especially if it also sits at an extreme predictor value. ## What minimising SSR does not do Minimising squared residuals is an arithmetic operation on whatever data and whatever predictors you hand it. It does not check that the relationship is really linear, that the right variables are included, or that the observations are independent. A minimised SSR is the best fit *within the family of lines you asked for*, and nothing more.

  • Why does OLS measure the miss vertically rather than perpendicular to the line?
    Because the model treats the predictors as given and puts the randomness in y: `y = b0 + b1*x + e`. The residual is the part of y that the predictors do not account for, so it is measured in the units of y. Minimising perpendicular distance is orthogonal (total least squares) regression, which assumes both variables carry measurement error and gives a different answer if you rescale one axis.
  • What happens to the OLS fit if you add a constant to every y value?
    The slope is unchanged and the intercept shifts by exactly that constant. Slopes depend on how y varies around its own mean, and adding a constant leaves those deviations identical. The intercept absorbs the shift, because it is pinned by `b0hat = ybar - b1hat*xbar` and ybar moved by the constant.
  • Does a smaller residual sum of squares always mean a better model?
    No. SSR is measured on the data used to fit, and it can only go down when you add predictors, so it rewards complexity automatically. It also says nothing about whether the functional form is right — a badly curved relationship can still have a small SSR if the noise is small. Judge fit against held-out data or a penalised criterion, not by SSR alone.

Think of each residual as a fine for missing the target, charged as the square of the distance. Small misses are cheap and large misses are punishing, so the line settles where the total fine is lowest.

saying these in an interview costs you the question

  • Says OLS minimises the perpendicular distance from points to the line
  • Claims OLS minimises the sum of residuals rather than their squares
  • Uses residual and error interchangeably as if both were observable
  • Thinks squaring exists only to remove minus signs
  • Believes a low residual sum of squares proves the model is correct

context

open as a page

How do you derive the OLS slope and intercept from the normal equations?

level: middleimportance: must knowfreq 62%

basics

~20 s

Set both partial derivatives of the squared-residual sum to zero to get the two normal equations. Solving them gives slope = sum of (x - xbar)(y - ybar) over sum of (x - xbar) squared, and intercept = ybar minus slope times xbar.

open as a page

Under what assumptions is OLS the best linear unbiased estimator?

level: middleimportance: should knowfreq 55%

basics

~20 s

The Gauss-Markov conditions: linearity in the coefficients, errors with zero mean given the predictors, constant variance, and no correlation between them, plus non-redundant predictors. Then OLS has the smallest variance among linear unbiased estimators. Normality is not required.

open as a page

Why do OLS residuals sum to exactly zero when the model includes an intercept?

level: middleimportance: should knowfreq 45%

basics

~20 s

Because the intercept's normal equation is literally the condition that the residuals sum to zero. Fitting a constant term forces that, which also makes the fitted line pass through the point of sample means. Drop the intercept and the guarantee disappears.

open as a page

Why prefer the OLS slope over one computed from only the first and last observations?

level: seniorimportance: nice to knowfreq 18%

basics

~20 s

A slope from two endpoints is linear in the data and unbiased, so it is a fair competitor, but its sampling variance is larger. Gauss-Markov says OLS wins that class, so unbiasedness alone does not make an estimator good.

open as a page