skip to content

Why does RMSE react far more strongly than MAE to a single large prediction error?

level: juniorimportance: must knowfreq 85%

answer

  1. squaring is not the same as averaging
  2. one big miss dominates the sum
  3. both live in the target's units
  4. mean of squares versus mean of sizes
  5. RMSE never falls below MAE

basics

~20 s

RMSE squares every error before averaging, so one huge miss contributes out of all proportion. MAE averages the absolute errors, giving each mistake weight in line with its size. RMSE tracks the worst errors, MAE the typical one.

solid answer

~50 s

Both are averages of the per-row error, but RMSE passes each error through a square first: `RMSE = sqrt(mean((y - pred)^2))` against `MAE = mean(|y - pred|)`. Squaring is convex, so doubling an error quadruples its contribution and a single extreme row can dominate the sum. Take 1,000 house-price predictions, 999 of them off by $20,000 and one mansion off by $2,000,000: MAE moves from $20,000 to about $21,980, while RMSE jumps from $20,000 to about $66,300. Both are in the target's own units, dollars, which is why RMSE and not MSE is the reportable number - MSE would be in squared dollars. RMSE is never below MAE on the same errors; the gap measures how uneven the errors are. Pick RMSE when a few big misses are disproportionately costly, MAE when you want the typical error.

code

python · 15 lines
python
import math

def mae(errors):
    return sum(abs(e) for e in errors) / len(errors)

def rmse(errors):
    return math.sqrt(sum(e * e for e in errors) / len(errors))

clean = [20_000] * 1000
with_mansion = [20_000] * 999 + [2_000_000]

print(round(mae(clean)), round(rmse(clean)))
print(round(mae(with_mansion)), round(rmse(with_mansion)))
# 20000 20000
# 21980 66329

go deeper

for a junior

Be ready to write both formulas from memory and say which one squares the errors first. Knowing that RMSE and MAE are both in the target's own units, while MSE is in squared units, carries most of the answer at this level.

for a middle

Explain why squaring makes one large error dominate and back it with numbers: a single extreme row can triple RMSE while moving MAE by a tenth. Expect to be asked which one you would report and to justify it.

for a senior

Connect the metric to the cost of being wrong in the product, and show that you inspect the largest residuals before letting RMSE decide anything, because one mistyped record can drive the whole number.

for a principal

Own which single figure steers the organisation. An RMSE headline points the team at rare catastrophic misses; an MAE headline points it at the typical case. That choice quietly determines what gets built next.

## The two definitions For `n` held-out rows with actual `y_i` and prediction `p_i`, write the error `e_i = y_i - p_i`. - `MAE = (1/n) * sum(|e_i|)` - the mean absolute error. - `MSE = (1/n) * sum(e_i^2)` - the mean squared error. - `RMSE = sqrt(MSE)` - the root mean squared error. MAE and RMSE are both expressed in the target's own units. If you are predicting house prices in dollars, both come out in dollars and a stakeholder can read them directly: "we are typically about $22,000 out." MSE comes out in squared dollars, which means nothing to anybody, and that is the only reason the square root is taken. ## Why the square changes the ranking of errors Absolute value is linear: an error of $200,000 counts exactly ten times an error of $20,000. Squaring is convex: the same $200,000 error counts one hundred times the $20,000 one before the averaging. The consequence is that RMSE is driven by the tail of the error distribution while MAE is driven by its middle. A concrete version, using a house-price model evaluated on 1,000 held-out sales. Suppose 999 of them are missed by $20,000 each and one $9M mansion is missed by $2,000,000. - Without the mansion, MAE = RMSE = $20,000, because every error is the same size. - With it, MAE = (999 * 20,000 + 2,000,000) / 1,000 = about $21,980 - a 10% move. - With it, MSE = (999 * 20,000^2 + 2,000,000^2) / 1,000 = about 4.40e9, so RMSE = about $66,300 - more than triple. One row in a thousand more than tripled the headline number. That is not a bug in RMSE; it is what RMSE is for. The question is whether it is what you want. ## RMSE is never smaller than MAE For any fixed set of errors, `RMSE >= MAE`, with equality exactly when every error has the same magnitude. So the ratio `RMSE / MAE` is a cheap diagnostic of error spread: near 1 means the model is uniformly mediocre, well above 1 means most predictions are decent and a minority are badly wrong. Two models can share an MAE and differ in RMSE - the higher-RMSE one concentrates its error in fewer, larger misses. ## Choosing between them The honest way to choose is by the cost of being wrong, not by which number looks better. - **Cost grows faster than the error.** A capacity or supply model where one enormous miss means an outage, a stockout or an emergency purchase. Big errors hurt superlinearly, so RMSE is the aligned metric. - **Cost is roughly proportional to the error.** A price estimate where each dollar off costs about the same as the last, or a dashboard reporting how far off a typical listing is. MAE is the aligned metric and is far easier to explain: half of the description of MAE is its own definition. - **You do not yet know.** Report both, plus the units. The pair tells the reader more than either alone, and the gap between them flags a skewed error distribution before anyone asks. One caution before you let RMSE decide anything: check that the rows blowing it up are real. A single mistyped price with an extra zero, a currency mix-up or a duplicated record can dominate an RMSE while being an artefact of the data rather than a property of the model. MAE is comparatively unbothered by one bad row, which is exactly why disagreement between the two metrics is worth a look at the largest residuals. ## Common misreadings - **"Lower RMSE means a better model."** Lower RMSE means fewer or smaller large errors on this test set. If the test set happens to contain one extreme row, RMSE largely measures how that one row went, and re-splitting the data can reorder your model ranking. - **"RMSE and MAE always agree."** They frequently disagree on model selection, and the disagreement is informative: the RMSE winner spreads its error more evenly, the MAE winner is better on the bulk of rows. - **"MSE is fine to report."** It is fine to optimise - it is smooth, differentiable and has the same minimiser as RMSE, because the square root is monotone. It is not fine to put in front of a stakeholder in squared units. - **"RMSE is scale-free."** Neither metric is. Both scale with the target, so an RMSE of 50,000 is meaningless without knowing whether houses here cost 200,000 or 20,000,000.

  • When would you deliberately headline RMSE even though a handful of outliers dominate it?
    When the cost of an error grows faster than the error itself. A capacity, load or inventory model where one enormous miss causes an outage or an emergency purchase is exactly that case: the rare catastrophic row is the thing you are trying to prevent, and MAE hides it behind nine hundred comfortable rows.
  • Two candidate models have the same MAE but clearly different RMSE. What does that tell you?
    Their errors have the same average size but different shapes. The higher-RMSE model concentrates its error in a few large misses while the other spreads it evenly, since RMSE equals MAE only when every error has the same magnitude. Choose by whether many small errors or a few large ones cost more.
  • If RMSE is the readable number, why is MSE still the usual training objective?
    MSE is smooth and differentiable everywhere, which makes the optimisation well behaved, and the square root is monotone, so whatever minimises MSE also minimises RMSE. The convention is to optimise MSE and report RMSE, because only the latter is in units a reader can act on.

Absolute error is a flat fine per mile per hour over the speed limit. Squared error is a fine that grows with the square of how far over you were: a dozen minor infractions cost less than one spectacular one.

saying these in an interview costs you the question

  • Claims RMSE and MAE always rank two models the same way
  • Reports MSE in dollars rather than squared dollars
  • Says RMSE is just MAE with an extra step
  • Believes RMSE can come out below MAE on the same errors
  • Treats a lower RMSE as automatically the better business outcome

context