skip to content

Your model is tuned to minimise MAE and its predicted totals fall short of actual totals on a right-skewed target. Why?

level: seniorimportance: should knowfreq 45%

answer

  1. each metric prefers a different constant
  2. counting points versus balancing distances
  3. skew separates the two centres
  4. median lies below the mean here
  5. predictions land on the conditional median

basics

~10 s

Minimising absolute error drives predictions toward the conditional median, and on a right-skewed target the median sits below the mean. Median-like predictions therefore sum low, even though each individual row looks accurate.

solid answer

~50 s

The two metrics have different optimal constants. If you had to predict one number for a whole column, the value minimising squared error is its mean, and the value minimising absolute error is its median. That carries over conditionally: a squared-error objective estimates the conditional mean, an absolute-error objective the conditional median. On a right-skewed target - revenue, claim size, session length - the median is below the mean, so every prediction is pulled under the average and the sum of predictions lands under the sum of actuals. Nothing is broken; the metric you optimised simply does not promise unbiased aggregates. Confirm it by comparing the shortfall to the target's mean-minus-median gap, then decide: keep MAE if you care about the typical row and correct or separately model the total, or move to squared error if the aggregate is what the business consumes.

go deeper

for a junior

Learn the pair of facts this rests on: the mean is the constant that minimises squared error and the median is the constant that minimises absolute error. Noticing that those are different numbers on a skewed column is the whole insight.

for a middle

Explain why the absolute-error optimum balances the count of points on each side while the squared-error optimum balances their distances. That difference is precisely what makes median-like predictions add up short.

for a senior

Show that you check aggregates and not only row-level error before shipping. Be ready to describe how you would separate a metric-induced shortfall from drift, and what you would change if a finance process consumes the total.

for a principal

Decide whether the organisation is served by unbiased aggregates or by robust per-row estimates - they are different products with different objectives. Own that call explicitly instead of letting a default metric make it silently.

## The fact underneath the symptom Start with the simplest possible model: predict one constant `c` for every row. - Minimising `mean((y_i - c)^2)` over `c` gives `c = mean(y)`. - Minimising `mean(|y_i - c|)` over `c` gives `c = median(y)`. The reason is what each derivative is balancing. For squared error, each point pulls on `c` in proportion to how far away it is, so the optimum is the balance point of the distances - the mean. For absolute error, each point pulls with the same unit force regardless of distance, and only its side matters, so the optimum is where the *counts* on either side balance - the median. A single enormous value moves the balance point of distances a long way and moves the headcount split not at all. This generalises from a constant to a model with features. A model fitted or tuned to minimise squared error is estimating the conditional mean `E[y | x]`; one fitted or tuned to minimise absolute error is estimating the conditional median. Two different targets, and they only coincide when the conditional distribution is symmetric. ## Why that makes the totals fall short Right-skewed targets - order value, claim amount, time on page, repair cost - have a long upper tail, and for such distributions the median lies below the mean. If every prediction is an estimate of a conditional median, then every prediction sits below the corresponding conditional mean, and summing thousands of them accumulates the gap into a visible shortfall against the actual total. The magnitude is predictable: roughly the number of rows multiplied by the average mean-minus-median gap of the conditional distributions. That predictability is your diagnostic. A metric-induced shortfall is stable and proportional - roughly the same percentage this month as last, and present on the training data as well as on the held-out data. A data problem behaves differently: leakage inflates apparent accuracy rather than biasing the sum in one direction, and drift shows as a shortfall that grows over time rather than one that sits still. ## Is it a defect? Only relative to what the predictions are used for. - **Row-level decisions.** "How long will this repair take?", "what should we quote this customer?" - a median-like answer is often exactly right, and is deliberately robust: the rare enormous job does not drag every other quote upward. - **Aggregate decisions.** Revenue forecasting, capacity planning, budget allocation, anything summed. Here you need predictions whose expectation matches the target's expectation, and only a mean-seeking objective gives you that. A model that is beautifully accurate per row and 8% light in total will be rejected by finance, correctly. The trap is shipping a model chosen on one criterion into a use case that needs the other. Absolute error looks like the safe, robust, grown-up choice - it is the one that resists outliers - and that very robustness is what breaks the totals. ## What to do about it 1. **Measure both.** Alongside your error metric, report the ratio of the sum of predictions to the sum of actuals on held-out data. It is one number and it catches this class of bug instantly. 2. **Match objective to use.** If the consumer sums the output, tune on squared error and accept the greater sensitivity to extreme rows. If the consumer reads rows one at a time, absolute error is the better fit. 3. **Correct explicitly if you must keep both.** A single multiplicative factor, estimated on held-out data, can reconcile the total while the per-row predictions stay median-like. Say so in the documentation: the corrected numbers are no longer the absolute-error optimum, and the factor needs re-estimating when the target distribution shifts. 4. **Do not fix it by adding features.** More features reduce the error but do not change which functional of the conditional distribution the objective is aiming at. The shortfall will shrink only in so far as the residual skew shrinks. ## The shape of the answer in an interview Name the mechanism first - absolute error targets the median, squared error the mean - then the skew, then the aggregation. Interviewers are checking whether you know that a metric is a *choice of estimand*, not just a scoreboard. The candidate who reaches for missing features or a bigger model has missed the point entirely.

  • How would you confirm this is the metric's doing rather than drift or a data bug?
    Check the same sum-of-predictions over sum-of-actuals ratio on the training data. A metric-induced shortfall appears there too, is stable across periods, and is close in size to the target's mean-minus-median gap. Drift produces a shortfall that grows over time, and leakage inflates accuracy rather than biasing the total in one direction.
  • The business wants accurate totals and a robust per-row number. What do you ship?
    Two outputs with two honest labels. Keep the absolute-error model for row-level answers, and either fit a squared-error model for the aggregate or apply a single multiplicative correction estimated on held-out data. Document that the corrected figures are no longer per-row optimal and that the factor must be re-estimated when the target distribution moves.
  • Would this shortfall also appear on a symmetric target?
    No. The median and mean coincide for a symmetric conditional distribution, so a median-seeking model has no systematic pull below the average and the totals reconcile. The gap is created by skew, which is why it shows up on money, durations and counts rather than on, say, standardised test scores.

The median is the point where equal numbers of people stand on each side of the room; the mean is where the floor would balance on a pivot. One billionaire walking in moves the pivot a long way and the headcount split hardly at all.

saying these in an interview costs you the question

  • Says absolute and squared error pick the same optimal constant
  • Blames the shortfall on missing features or an undertrained model
  • Thinks the median sits above the mean under right skew
  • Assumes any per-row accurate model must also sum correctly
  • Proposes adding a constant to every prediction without measuring the gap

context