In linear regression, why is the fit chosen to minimise squared errors rather than absolute errors?
answer
- think about the derivative, not fairness
- absolute value has a kink at zero
- derivative of a square is linear
- residual of ten costs one hundred
- mean versus median
basics
~20 sSquared error is smooth everywhere, so setting its derivative to zero gives one exact formula for the best-fitting weights. Absolute error has a kink at zero, needs iterative fitting, and squaring also makes large misses far more expensive.
solid answer
~50 sLeast squares minimises `sum (y_i - pred_i)^2` mainly for smoothness: the derivative of a squared term is linear in the weights, so setting the gradient to zero yields a system of linear equations with an exact solution. Absolute error has a kink at zero — its derivative is just the sign of the residual — so there is no linear system to solve and you fall back on iterative optimisation. The second consequence is influence: a residual of 10 costs 100 under squared error but only 10 under absolute error, so least squares pulls the fit hard towards far-away points, while an absolute-error fit is indifferent to how distant an outlier actually is. Squared error also targets the conditional mean; absolute error targets the conditional median. With heavy-tailed targets, that robustness can be worth giving up the closed form.
go deeper
Be ready to state the objective in one line — the fit minimises the sum of squared residuals — and to say that squaring keeps errors positive and punishes big misses hardest.
An interviewer expects the mechanics: the derivative of a squared term is linear in the weights, so the optimum becomes a solvable system of linear equations, while the derivative of an absolute value is only a sign and yields no such system.
Show judgment about influence. Know when the outlier-chasing of squared error is wrong for the data in front of you, what you would switch to — absolute error, Huber, a log-transformed target — and what fitting cost that switch buys.
Own the framing that the training objective encodes what the organisation considers an expensive mistake. Be able to argue when a quadratic penalty misstates real cost, and what it takes to change the objective consistently across a team's models.
## What "least squares" actually minimises A linear regression predicts the target as a weighted sum of the input features plus an intercept: `pred = w0 + w1*x1 + w2*x2 + ... + wd*xd`. Fitting the model means choosing the weights. To choose them you need a **training objective** — one number that says how bad a particular set of weights is on the training data. For ordinary linear regression that number is the mean squared error: `MSE(w) = (1/n) * sum_i (y_i - pred_i)^2` Here `y_i` is the true target for row `i`, `pred_i` is the model's output for that row, and the difference `y_i - pred_i` is the **residual**. The obvious alternative is mean absolute error: `MAE(w) = (1/n) * sum_i |y_i - pred_i|` Both are legitimate objectives and both are convex in the weights, so neither is "wrong". Squared error is the default for two concrete reasons — what squaring does to the mathematics of finding the optimum, and what it does to the influence of individual rows. ### Reason one: squaring is smooth, so the optimum is a linear equation The derivative of `(y - pred)^2` with respect to a weight is `-2 * x * (y - pred)`. Every term is **linear in the weights**, so the whole gradient is linear in the weights. Setting the gradient to zero therefore gives a system of linear equations — one equation per weight — and a system of linear equations can be solved exactly, in one shot, with standard linear algebra. That is the entire reason a closed-form fit exists at all. Absolute error has no such luck. The derivative of `|y - pred|` is `-x * sign(y - pred)`: a piecewise-constant quantity that jumps at zero and is undefined exactly there. Setting that to zero gives no linear system to solve. Fitting an absolute-error line means iterative optimisation (subgradient methods, iteratively reweighted least squares, or a linear program). It is perfectly doable; it just is not a formula. ### Reason two: squaring makes big misses expensive Under squared error a residual of 1 costs 1 and a residual of 10 costs 100 — a hundred times more. Under absolute error the same two residuals cost 1 and 10 — only ten times more. Squared error therefore buys down the *largest* residuals first, and the fitted line visibly leans towards far-away points. Whether that is a feature or a bug depends on the data: - **Feature** when the real-world cost of being wrong genuinely grows faster than linearly — a forecast that is off by an hour is far more than sixty times worse than one off by a minute. - **Bug** when the far-away points are data-entry mistakes, sensor dropouts or a heavy tail you do not care about. Then squared error lets a handful of bad rows drag the whole fit. Absolute error is robust in the precise sense that moving an already-distant point further away does not change the fit at all — only the *sign* of its residual enters the gradient, not the magnitude. ### Reason three: they estimate different things Take the simplest possible linear model, an intercept only, so the model predicts one constant `c` for every row. With targets `1, 2, 3, 100`: - Minimising squared error gives the **mean**, `c = 26.5`. - Minimising absolute error gives the **median**, `c = 2.5`. That generalises: a squared-error fit estimates the conditional **mean** of the target given the features; an absolute-error fit estimates the conditional **median**. Pick the objective that matches the quantity you actually want to predict. There is also a classical statistical justification — if the noise around the line is Gaussian with constant variance, minimising squared error is exactly maximum-likelihood estimation — but in day-to-day work the smoothness and the exact solve are the reasons that bite. ### What this means in practice - **The objective you optimise and the number you report to stakeholders need not match.** Training on squared error while reporting an absolute-error number is a common, sane combination. - **Units.** Squared error is in squared target units, which is why people take its square root before quoting it. Nothing about the fit changes; only the readability of the number does. - **Skewed targets.** Log-transforming a strictly positive, heavy-tailed target before fitting turns squared error in log space into a penalty on *relative* error, which often matches the intent better than raw squared error and keeps the closed form. - **When you really want robustness**, the standard escape hatches are an absolute-error objective, a Huber objective (quadratic near zero, linear in the tails, so gradients stay smooth but distant points stop dominating), or capping/winsorising extreme targets. All of the first two cost you the one-shot solve and put you back on iterative fitting — usually a fine trade. ### The short version to say out loud Squaring is chosen for smoothness first and influence second. Smoothness converts "find the minimum" into "solve a linear system", which is why an exact formula exists. Influence means outliers are counted quadratically, which is a deliberate and sometimes unwanted choice — and the moment you decide it is unwanted, you are back to iterating.
- When would you deliberately train a linear model on absolute error instead?When the targets carry heavy tails or data-entry outliers you do not want the line chasing, and when the real cost of being wrong grows roughly linearly with the miss. You give up the one-shot solve and fit iteratively, and you are now estimating the conditional median rather than the conditional mean — which is often the more useful summary for skewed targets like price or duration.
- What does a Huber objective buy you over either of those two?Huber is quadratic for small residuals and linear beyond a threshold, so gradients stay smooth near the optimum while distant points stop dominating. You get most of the robustness of absolute error without its non-differentiable kink. The price is a threshold parameter you have to pick, usually by cross-validation, and iterative fitting — there is no closed form.
- How does the scale or skew of the target interact with this choice?Squared error lives in squared target units, so a strictly positive, heavily skewed target lets a few huge rows dominate the objective. Log-transforming the target first turns squared error in log space into a penalty on relative rather than absolute error, which often matches the intent better and keeps the exact solve. Remember predictions then come back on the log scale.
Squared error is a fine that quadruples when you are twice as far off, so the fit leans towards whoever is furthest away. Absolute error charges a flat rate per metre and stops caring how far the worst offender actually is.
saying these in an interview costs you the question
- Says squaring exists only to make errors positive
- Claims absolute error is non-convex or has no minimum
- Calls least squares robust to outliers
- Thinks squaring shrinks large errors
- Cannot say the constant fit is mean versus median
- Assumes absolute error simply cannot be optimised