Why does one extreme target value distort a squared-error regression fit more than an absolute-error fit?
answer
- look at the derivative, not the loss
- how hard does one row pull?
- squared error's pull scales with the miss
- absolute error gives every row one vote
- mean optimum versus median optimum
basics
~20 sSquared error penalises a residual by its square, so its gradient grows with the residual: a point ten times further off pulls ten times harder. Absolute error's gradient has fixed size, so extreme points get no extra vote.
solid answer
~50 sWhat matters during fitting is the derivative of the loss, not its value. For residual `r = y - prediction`, squared error contributes `r^2` whose derivative is proportional to `r`, so a row missed by 100 units pushes the parameters a hundred times harder than a row missed by 1. Absolute error contributes `|r|`, whose derivative is just the sign of the residual, so every row gets one equal-sized vote and only the count of rows above and below the fit matters. The consequence is not merely stability: the constant minimising mean squared error is the mean, and the constant minimising mean absolute error is the median, so a squared-error fit estimates the conditional mean and an absolute-error fit the conditional median. On a right-skewed target those are different heights, and which you want is a product decision.
go deeper
Recall that squared error squares the miss while absolute error does not, so a single badly missed row counts enormously more under squared error. Know that squared error is the default for linear regression.
Be ready to differentiate both losses aloud and show that the squared-error gradient scales with the residual while the absolute-error gradient is a constant sign. Then name the two optima: mean for squared error, median for absolute error.
Show the production call. On a skewed target, say which summary the business actually needs, and warn that swapping the loss moves every prediction, not just the extreme ones, so downstream totals shift the day you deploy it.
Own the tradeoff between a mean fit whose predictions add up into budgets and a median fit that reads better per case. Decide which the organisation forecasts against, and make that choice explicit rather than an accident of the default loss.
## The two losses For one observation with actual value `y` and prediction `p`, write the residual `r = y - p`. Training chooses the parameters that minimise the average per-row contribution: - squared error contributes `r^2` - absolute error contributes `|r|` Both are zero when the prediction is exact and both grow as the miss grows, so at a glance they look interchangeable. They are not. They produce different fitted lines on the same data, and they estimate different things about the target. ## The gradient, not the loss value, does the pulling An iterative fit moves the parameters against the derivative of the loss. That derivative is the pull each row exerts. - The derivative of `r^2` with respect to the prediction is `-2r`. Its magnitude grows **linearly with the residual**. - The derivative of `|r|` with respect to the prediction is `-sign(r)`. Its magnitude is **1 for every row** that is not exactly on the line. So under squared error a row missed by 100 units exerts a hundred times the pull of a row missed by 1. Ten ordinary rows pulling the other way cannot outvote it. Under absolute error each row exerts the same pull whatever its distance, so the fit responds to *how many* rows sit above and below it, not to *how far* above. Make it concrete with cloud-capacity provisioning. Peak CPU on an ordinary day sits between 40 and 60 percent of a node; a handful of incident days spike to several times normal. A squared-error fit will lift the whole curve to shave those few enormous residuals, because shaving one residual of 250 buys more loss reduction than fitting a hundred ordinary days better. Every quiet day is then over-provisioned. An absolute-error fit counts an incident day as one row above the line, exactly like a day two points above, and settles where half the days fall on each side. ## What each loss actually estimates This is the statement interviewers are listening for. - The constant `c` that minimises `E[(y - c)^2]` is the **mean** of `y`. - The constant `c` that minimises `E[|y - c|]` is the **median** of `y`. Conditioning on features carries the same result through the regression: a squared-error fit estimates the **conditional mean** of the target, an absolute-error fit estimates the **conditional median**. On a symmetric target distribution the two coincide and the choice is cosmetic. On a right-skewed one -- durations, spend, capacity, claim sizes -- the mean sits above the median, the two fits sit at visibly different heights, and neither is wrong. Which one you want is a product question, not a statistical one. If predictions are summed -- a million per-order predictions rolled into a budget -- the mean is the additive quantity and squared error is the honest loss, because medians do not add up to the median of the total. If a single prediction is shown to a human as "what to expect", the median is the better summary of a skewed outcome. ## What you pay for absolute error 1. **No closed form.** A linear model under squared error has an exact algebraic solution. Under absolute error it is fitted iteratively. 2. **Non-differentiable at zero.** The derivative jumps from `-1` to `+1` as the residual crosses zero. Subgradient methods handle this, but progress near the optimum is slower, and with certain data the minimiser is not unique -- a whole interval of lines can tie. 3. **Statistically less efficient under clean Gaussian noise.** If the errors really are normal with no contamination, squared error extracts more information per observation and yields a tighter estimate. Robustness is insurance, and insurance has a premium. ## What it does not protect against Absolute error buys robustness to an outlying **target**. It buys very little against an outlying **feature** value. A row sitting far out in feature space -- a high-leverage point -- can still swing the line, because the fit can satisfy that isolated row almost exactly while paying nearly nothing elsewhere. Robust losses and leverage diagnostics solve different problems, and a candidate who claims absolute error makes a regression "outlier-proof" is overselling it. ## The middle ground Because the two extremes each give something up, the usual production answer is a hybrid: a loss that behaves quadratically for small residuals -- keeping the smooth, efficient behaviour where the data is well-behaved -- and linearly once a residual passes a chosen threshold, capping how hard any single row can pull. That is Huber loss, and the threshold is its one knob. One framing point worth keeping straight: changing the loss and changing the data are different levers. Switching to absolute error keeps every row in the training set and changes only how much each row is allowed to influence the parameters.
- If absolute error is more robust, why is squared error still the default for linear regression?It is smooth and differentiable everywhere, it has an exact closed-form solution for a linear model, and it is statistically more efficient when the noise really is Gaussian. It also estimates the conditional mean, which is the quantity that adds up correctly when predictions are summed into totals or budgets.
- What does an absolute-error fit predict when the target distribution is strongly right-skewed?The conditional median, which sits below the conditional mean on a right-skewed target. Per-row it reads as the typical case, but summing those predictions across many rows will understate the total, so it is the wrong basis for a budget even while being the better single-case answer.
- Does absolute error protect against an extreme feature value the same way it protects against an extreme target value?No. It caps the influence of a row with an unusual target, but a high-leverage row far out in feature space can still swing the line, because the fit satisfies that isolated point at almost no cost elsewhere. Leverage needs its own diagnostics, not a different loss.
- Why is an absolute-error fit harder to optimise near the solution?Its derivative is the sign of the residual, which jumps discontinuously from minus one to plus one at zero, so there is no gradient information telling the optimiser how close it is. Steps do not shrink naturally near the optimum, and in some datasets a whole interval of lines ties for best.
Squared error is a vote weighted by how loudly you shout; absolute error is one person, one vote. A single very loud voice can carry the room under the first rule and cannot under the second.
saying these in an interview costs you the question
- Says absolute error removes the outlying rows from the data
- Claims the two losses fit the same line, just with different reported numbers
- Thinks squaring is only there to make residuals positive
- Says a robust loss also fixes high-leverage points in the features
- Cannot say which loss targets the mean and which the median