What quantity does ordinary least squares minimise when fitting a linear model?
answer
- vertical misses, not perpendicular ones
- signed residuals would cancel
- squares, summed over observations
- observed minus fitted, then squared
- each y splits into fitted plus residual
basics
~20 sOrdinary least squares picks the coefficients that make the sum of squared residuals as small as possible. A residual is an observed value minus the value the fitted line predicts for it, so the misses are vertical.
solid answer
~50 sThe model is `y = b0 + b1*x1 + ... + e`, and OLS chooses the coefficient estimates that minimise `SSR = sum over i of (y_i - yhat_i)^2`, where `yhat_i` is the fitted value and `e_i = y_i - yhat_i` is the residual. Every observation splits exactly as `y_i = yhat_i + e_i`, so OLS is the split that pushes as much of the variation as possible into the fitted part. The distances are vertical: the model treats the predictors as given and puts the randomness in y, so a miss is measured in the units of y. Squaring makes the objective smooth, which is what produces a closed-form solution, and it penalises a miss of 4 sixteen times as heavily as a miss of 1 — one far-off point can visibly pull the line.
go deeper
Be ready to state the objective in one line — smallest total squared vertical distance from the points to the line — and to compute a residual as observed minus fitted for a single point.
Explain why squares rather than signed values or absolute values: signed residuals cancel, absolute values have no closed form, squares give a smooth convex objective and a solvable system of equations.
Show that you know the cost of the squared objective in real data: quadratic penalties make a single extreme observation influential, so pair the fit with a look at which observations are driving it.
Own the framing decision. Squared error encodes a specific loss — symmetric, quadratic, mean-seeking — and if the business cost of over- and under-prediction is asymmetric or heavy-tailed, argue for a different objective rather than defaulting to OLS.
## The model, and what fitting means A linear regression model says that an outcome `y_i` for observation `i` is a straight-line function of one or more predictors plus an unobservable disturbance: ``` y_i = b0 + b1*x_i1 + ... + bk*x_ik + e_i ``` Here `b0` is the intercept, `b1 ... bk` are the true (unknown, fixed) coefficients, and `e_i` is the **error** — the amount by which observation `i` sits off the true line. The error is never observable, because the true coefficients are never known. Fitting means choosing numerical **estimates** `b0hat ... bkhat` from the data. Those estimates give each observation a **fitted value** ``` yhat_i = b0hat + b1hat*x_i1 + ... + bkhat*x_ik ``` and a **residual** ``` e_i = y_i - yhat_i ``` The residual is computable; the error is not. Keeping the two apart is the single most common vocabulary slip in this area — a residual is a deviation from the *estimated* line, an error is a deviation from the *true* line. ## The objective Ordinary least squares is defined by its objective function, the residual sum of squares: ``` SSR(b0, ..., bk) = sum over i of (y_i - (b0 + b1*x_i1 + ... + bk*x_ik))^2 ``` OLS chooses the coefficient values that make this number smallest. Nothing else about OLS is a matter of choice: given the data and the chosen predictors, the estimates are whatever minimises this sum. Two details in that definition matter. First, the distances being summed are **vertical** — differences in y at a fixed x — not perpendicular distances from the point to the line. Minimising perpendicular distance is a different method (orthogonal or total least squares) that assumes both variables are measured with error and, unlike OLS, changes its answer if you rescale x from dollars to thousands of dollars. Second, the quantity minimised is the sum of *squares*, not of signed residuals: signed residuals cancel, so their sum can be zero for many badly-fitting lines and cannot single out a best one. ## The decomposition on a whiteboard-sized example The identity `y_i = yhat_i + e_i` holds for *any* line you draw, good or bad; OLS is the choice that makes the second piece as small as possible in squared total. Take five months of advertising spend and sales, both in thousands of dollars: ``` x (ad spend): 1 2 3 4 5 y (sales): 3 8 7 10 12 ``` The OLS fit for this data is `yhat = 2 + 2x`. That gives fitted values 4, 6, 8, 10, 12 and residuals ``` e: -1 +2 -1 0 0 ``` and `SSR = 1 + 4 + 1 + 0 + 0 = 6`. Check that no nearby line does better. Take `yhat = 2.5 + 1.9x`: fitted values 4.4, 6.3, 8.2, 10.1, 12.0, residuals -1.4, +1.7, -1.2, -0.1, 0.0, and `SSR = 1.96 + 2.89 + 1.44 + 0.01 + 0 = 6.30`. Larger, as it must be. Every observation still decomposes exactly — 8 is 6 plus 2 under the OLS fit, and 8 is 6.3 plus 1.7 under the alternative — but the OLS split leaves the smallest squared remainder. ## Why squares, and not something else The practical reason is smoothness. The squared objective is a differentiable, bowl-shaped (convex) function of the coefficients, so its minimum can be found by setting partial derivatives to zero, which yields a system of linear equations with a closed-form solution. Absolute deviations give a perfectly sensible alternative estimator (least absolute deviations, which is more resistant to outliers and behaves like a conditional median rather than a conditional mean), but the objective has a kink at every zero residual and no closed form. The deeper reason is that squares come with an optimality guarantee: under a standard set of assumptions about the errors, the least-squares estimator has the smallest variance of any linear unbiased estimator — the Gauss-Markov result. So squaring is not merely a trick for removing minus signs. The cost is sensitivity. Because the penalty grows quadratically, one observation far from the pattern contributes disproportionately and can tilt the fitted line noticeably, especially if it also sits at an extreme predictor value. ## What minimising SSR does not do Minimising squared residuals is an arithmetic operation on whatever data and whatever predictors you hand it. It does not check that the relationship is really linear, that the right variables are included, or that the observations are independent. A minimised SSR is the best fit *within the family of lines you asked for*, and nothing more.
- Why does OLS measure the miss vertically rather than perpendicular to the line?Because the model treats the predictors as given and puts the randomness in y: `y = b0 + b1*x + e`. The residual is the part of y that the predictors do not account for, so it is measured in the units of y. Minimising perpendicular distance is orthogonal (total least squares) regression, which assumes both variables carry measurement error and gives a different answer if you rescale one axis.
- What happens to the OLS fit if you add a constant to every y value?The slope is unchanged and the intercept shifts by exactly that constant. Slopes depend on how y varies around its own mean, and adding a constant leaves those deviations identical. The intercept absorbs the shift, because it is pinned by `b0hat = ybar - b1hat*xbar` and ybar moved by the constant.
- Does a smaller residual sum of squares always mean a better model?No. SSR is measured on the data used to fit, and it can only go down when you add predictors, so it rewards complexity automatically. It also says nothing about whether the functional form is right — a badly curved relationship can still have a small SSR if the noise is small. Judge fit against held-out data or a penalised criterion, not by SSR alone.
Think of each residual as a fine for missing the target, charged as the square of the distance. Small misses are cheap and large misses are punishing, so the line settles where the total fine is lowest.
saying these in an interview costs you the question
- Says OLS minimises the perpendicular distance from points to the line
- Claims OLS minimises the sum of residuals rather than their squares
- Uses residual and error interchangeably as if both were observable
- Thinks squaring exists only to remove minus signs
- Believes a low residual sum of squares proves the model is correct