skip to content

Closed-Form Least Squares

Squared error makes the fit solvable in one step by the normal equations, but the inversion is cubic in features and fails when they are collinear. Interviewers ask when to iterate instead.

on this pageshow

questions

3

In linear regression, why is the fit chosen to minimise squared errors rather than absolute errors?

level: juniorimportance: should knowfreq 58%

answer

  1. think about the derivative, not fairness
  2. absolute value has a kink at zero
  3. derivative of a square is linear
  4. residual of ten costs one hundred
  5. mean versus median

basics

~20 s

Squared error is smooth everywhere, so setting its derivative to zero gives one exact formula for the best-fitting weights. Absolute error has a kink at zero, needs iterative fitting, and squaring also makes large misses far more expensive.

solid answer

~50 s

Least squares minimises `sum (y_i - pred_i)^2` mainly for smoothness: the derivative of a squared term is linear in the weights, so setting the gradient to zero yields a system of linear equations with an exact solution. Absolute error has a kink at zero — its derivative is just the sign of the residual — so there is no linear system to solve and you fall back on iterative optimisation. The second consequence is influence: a residual of 10 costs 100 under squared error but only 10 under absolute error, so least squares pulls the fit hard towards far-away points, while an absolute-error fit is indifferent to how distant an outlier actually is. Squared error also targets the conditional mean; absolute error targets the conditional median. With heavy-tailed targets, that robustness can be worth giving up the closed form.

go deeper

for a junior

Be ready to state the objective in one line — the fit minimises the sum of squared residuals — and to say that squaring keeps errors positive and punishes big misses hardest.

for a middle

An interviewer expects the mechanics: the derivative of a squared term is linear in the weights, so the optimum becomes a solvable system of linear equations, while the derivative of an absolute value is only a sign and yields no such system.

for a senior

Show judgment about influence. Know when the outlier-chasing of squared error is wrong for the data in front of you, what you would switch to — absolute error, Huber, a log-transformed target — and what fitting cost that switch buys.

for a principal

Own the framing that the training objective encodes what the organisation considers an expensive mistake. Be able to argue when a quadratic penalty misstates real cost, and what it takes to change the objective consistently across a team's models.

## What "least squares" actually minimises A linear regression predicts the target as a weighted sum of the input features plus an intercept: `pred = w0 + w1*x1 + w2*x2 + ... + wd*xd`. Fitting the model means choosing the weights. To choose them you need a **training objective** — one number that says how bad a particular set of weights is on the training data. For ordinary linear regression that number is the mean squared error: `MSE(w) = (1/n) * sum_i (y_i - pred_i)^2` Here `y_i` is the true target for row `i`, `pred_i` is the model's output for that row, and the difference `y_i - pred_i` is the **residual**. The obvious alternative is mean absolute error: `MAE(w) = (1/n) * sum_i |y_i - pred_i|` Both are legitimate objectives and both are convex in the weights, so neither is "wrong". Squared error is the default for two concrete reasons — what squaring does to the mathematics of finding the optimum, and what it does to the influence of individual rows. ### Reason one: squaring is smooth, so the optimum is a linear equation The derivative of `(y - pred)^2` with respect to a weight is `-2 * x * (y - pred)`. Every term is **linear in the weights**, so the whole gradient is linear in the weights. Setting the gradient to zero therefore gives a system of linear equations — one equation per weight — and a system of linear equations can be solved exactly, in one shot, with standard linear algebra. That is the entire reason a closed-form fit exists at all. Absolute error has no such luck. The derivative of `|y - pred|` is `-x * sign(y - pred)`: a piecewise-constant quantity that jumps at zero and is undefined exactly there. Setting that to zero gives no linear system to solve. Fitting an absolute-error line means iterative optimisation (subgradient methods, iteratively reweighted least squares, or a linear program). It is perfectly doable; it just is not a formula. ### Reason two: squaring makes big misses expensive Under squared error a residual of 1 costs 1 and a residual of 10 costs 100 — a hundred times more. Under absolute error the same two residuals cost 1 and 10 — only ten times more. Squared error therefore buys down the *largest* residuals first, and the fitted line visibly leans towards far-away points. Whether that is a feature or a bug depends on the data: - **Feature** when the real-world cost of being wrong genuinely grows faster than linearly — a forecast that is off by an hour is far more than sixty times worse than one off by a minute. - **Bug** when the far-away points are data-entry mistakes, sensor dropouts or a heavy tail you do not care about. Then squared error lets a handful of bad rows drag the whole fit. Absolute error is robust in the precise sense that moving an already-distant point further away does not change the fit at all — only the *sign* of its residual enters the gradient, not the magnitude. ### Reason three: they estimate different things Take the simplest possible linear model, an intercept only, so the model predicts one constant `c` for every row. With targets `1, 2, 3, 100`: - Minimising squared error gives the **mean**, `c = 26.5`. - Minimising absolute error gives the **median**, `c = 2.5`. That generalises: a squared-error fit estimates the conditional **mean** of the target given the features; an absolute-error fit estimates the conditional **median**. Pick the objective that matches the quantity you actually want to predict. There is also a classical statistical justification — if the noise around the line is Gaussian with constant variance, minimising squared error is exactly maximum-likelihood estimation — but in day-to-day work the smoothness and the exact solve are the reasons that bite. ### What this means in practice - **The objective you optimise and the number you report to stakeholders need not match.** Training on squared error while reporting an absolute-error number is a common, sane combination. - **Units.** Squared error is in squared target units, which is why people take its square root before quoting it. Nothing about the fit changes; only the readability of the number does. - **Skewed targets.** Log-transforming a strictly positive, heavy-tailed target before fitting turns squared error in log space into a penalty on *relative* error, which often matches the intent better than raw squared error and keeps the closed form. - **When you really want robustness**, the standard escape hatches are an absolute-error objective, a Huber objective (quadratic near zero, linear in the tails, so gradients stay smooth but distant points stop dominating), or capping/winsorising extreme targets. All of the first two cost you the one-shot solve and put you back on iterative fitting — usually a fine trade. ### The short version to say out loud Squaring is chosen for smoothness first and influence second. Smoothness converts "find the minimum" into "solve a linear system", which is why an exact formula exists. Influence means outliers are counted quadratically, which is a deliberate and sometimes unwanted choice — and the moment you decide it is unwanted, you are back to iterating.

  • When would you deliberately train a linear model on absolute error instead?
    When the targets carry heavy tails or data-entry outliers you do not want the line chasing, and when the real cost of being wrong grows roughly linearly with the miss. You give up the one-shot solve and fit iteratively, and you are now estimating the conditional median rather than the conditional mean — which is often the more useful summary for skewed targets like price or duration.
  • What does a Huber objective buy you over either of those two?
    Huber is quadratic for small residuals and linear beyond a threshold, so gradients stay smooth near the optimum while distant points stop dominating. You get most of the robustness of absolute error without its non-differentiable kink. The price is a threshold parameter you have to pick, usually by cross-validation, and iterative fitting — there is no closed form.
  • How does the scale or skew of the target interact with this choice?
    Squared error lives in squared target units, so a strictly positive, heavily skewed target lets a few huge rows dominate the objective. Log-transforming the target first turns squared error in log space into a penalty on relative rather than absolute error, which often matches the intent better and keeps the exact solve. Remember predictions then come back on the log scale.

Squared error is a fine that quadruples when you are twice as far off, so the fit leans towards whoever is furthest away. Absolute error charges a flat rate per metre and stops caring how far the worst offender actually is.

saying these in an interview costs you the question

  • Says squaring exists only to make errors positive
  • Claims absolute error is non-convex or has no minimum
  • Calls least squares robust to outliers
  • Thinks squaring shrinks large errors
  • Cannot say the constant fit is mean versus median
  • Assumes absolute error simply cannot be optimised

context

open as a page

Why does a closed-form least-squares solve scale cubically in features but linearly in rows?

level: middleimportance: should knowfreq 45%

basics

~20 s

Rows are touched once, in a streaming pass building a p-by-p system. Solving that system costs about p^3 and stores p^2 numbers, so total cost is roughly n*p^2 + p^3: wide designs break the exact solve, not tall ones.

open as a page

What happens to a closed-form least-squares fit when two features are nearly identical?

level: seniorimportance: should knowfreq 58%

basics

~20 s

Near-duplicate columns leave the cross-product matrix almost singular, so the two coefficients come out enormous, opposite in sign and wildly unstable across refits. Predictions inside the training range stay fine; the individual coefficients are not interpretable.

open as a page