Why is MAPE a poor scoring metric for a call-centre wait-time model where some waits are under a minute?
answer
- look hard at the denominator
- tiny actuals, enormous percentages
- the penalty is not symmetric
- under-prediction is capped at 100%
- the bounded variant leans the other way
basics
~10 sMAPE divides each error by the actual value, so a 30-second miss on a one-minute wait scores 50% and near-zero waits explode. It is also asymmetric: over-prediction is unbounded, under-prediction capped at 100%.
solid answer
~50 sMAPE is `mean(|actual - pred| / |actual|)`, usually times 100, so the actual value sits in the denominator. Short waits make that denominator tiny: predicting 90 seconds when the true wait was 30 gives a 200% error, and a wait recorded as zero makes the term undefined outright. A handful of such rows can dominate the average and turn the score into a measure of how many quick calls you had, not how good the model is. The second defect is asymmetry. For a positive actual, an over-prediction can produce an unbounded percentage error, while the worst possible under-prediction - predicting zero - is exactly 100%. Optimising or selecting on MAPE therefore drifts toward forecasts that are too low, which for a wait-time promise is the damaging direction. Report MAE in minutes instead, and if a relative figure is required use a weighted form: total absolute error divided by total actual minutes.
go deeper
Remember what MAPE divides by. Each error is scaled by that row's actual value, so tiny actuals produce enormous percentages and an actual of zero makes the term undefined altogether.
Give both defects with numbers: the exploding denominator on short waits, and the 100% cap on under-prediction against no cap on over-prediction. Be able to state what sMAPE changes and what it does not.
Show you would catch the resulting bias in production - a model selected on MAPE quietly under-forecasts - and name the metric and unit you would put on the dashboard instead, with a reason a stakeholder would accept.
Own the reporting standard. Percentage metrics spread because they look comparable across teams and products, so argue where a weighted or bounded variant is tolerable and where absolute units must be mandatory.
## What MAPE computes `MAPE = (100/n) * sum(|actual_i - pred_i| / |actual_i|)`. Each row's error is rescaled by that row's actual value before averaging, which is what makes the result a percentage and what makes it comparable across targets of different sizes. That is the attraction: a single number that a business audience reads without knowing the units. Everything that goes wrong with it flows from the same denominator. ## Defect one: the denominator Wait times are not bounded away from zero. Plenty of calls are answered in ten or twenty seconds; some are logged as zero. - Actual 30 seconds, predicted 90 seconds: the absolute error is a minute, which nobody would call a bad prediction, but the percentage error is 60/30 = 200%. - Actual 5 seconds, predicted 65 seconds: the same one-minute miss now scores 1,200%. - Actual 0: the term is undefined, and dropping those rows silently changes what the metric measures. Averaged over the test set, a few dozen quick calls can outweigh thousands of ordinary ones. Two models can then be ranked by their behaviour on the shortest waits, which is usually the least important segment of the problem. Worse, the score moves when the *mix* of calls changes even if the model does not: a busy month with fewer instant answers will show a better MAPE from the same predictions. ## Defect two: asymmetry Fix an actual value of 100 seconds and vary the prediction. Assuming predictions cannot be negative: - Predict 0 - the most extreme possible under-prediction - and the absolute percentage error is 100%. - Predict 200 and it is 100%. Predict 400 and it is 300%. Predict 1,000 and it is 900%. Under-prediction is capped at 100%; over-prediction has no ceiling. Any procedure that selects, tunes or thresholds on MAPE therefore prefers models that sit low. On a wait-time model this is the worst possible bias: the system will systematically promise callers shorter waits than they experience, generating exactly the complaints the model was meant to prevent. Note the phrasing carefully - MAPE penalises *over*-prediction more heavily, which means it *favours* under-prediction. Candidates routinely get that direction backwards. ## Does sMAPE fix it? The symmetric mean absolute percentage error puts the average of the two values in the denominator: `sMAPE = (100/n) * sum( |actual - pred| / ((|actual| + |pred|) / 2) )`. That bounds each term at 200% and removes the divide-by-zero whenever either value is positive, so the metric no longer explodes on short waits. But the name oversells it. With actual 100 and prediction 150, the term is 50/125 = 40%. With actual 100 and prediction 50 - the same absolute miss in the other direction - it is 50/75 = 67%. sMAPE punishes under-prediction more, so it has simply reversed MAPE's bias rather than removed it. It is also still undefined when actual and prediction are both zero, and a bounded value of 200% is not an intuitive thing to explain to a stakeholder. ## What to report instead For this problem, the honest reporting set is short. - **MAE in minutes.** The stakeholder acts in minutes and staffs in minutes; give them the number in that unit. "We are typically 1.4 minutes out" is directly usable in a way that "MAPE is 63%" is not. - **A weighted absolute percentage error, if a relative number is demanded.** Total absolute error divided by total actual time, `sum(|actual - pred|) / sum(actual)`. Because the denominator is aggregated once rather than per row, near-zero waits cannot blow it up, and the result is dominated by the long waits that actually cost money. - **RMSE in minutes** if the long waits are the expensive failures and you want a metric that leans on them. ## The general rule A per-row relative metric is only safe when the target is bounded well away from zero and roughly on one scale. Prices of houses, yes. Wait times, call volumes, counts, anything that can legitimately be near zero, no. When someone asks for "error as a percentage" because it sounds comparable, the correct response is to ask what decision the number feeds, then choose between an absolute unit and an aggregate-weighted percentage.
- Which direction does MAPE bias model selection, and why does that matter here?Toward under-prediction. Over-shooting can produce arbitrarily large percentage errors while under-shooting tops out at 100%, so the metric quietly prefers low forecasts. A call centre steered by MAPE will promise shorter waits than callers actually get, which is the direction that generates complaints and abandoned calls.
- Does sMAPE solve MAPE's problems?Only partly. Averaging actual and prediction in the denominator bounds each term at 200% and removes the blow-up on tiny actuals, but it reverses the asymmetry - under-prediction now scores worse than the same over-prediction - and it is still undefined when both values are zero. It swaps one bias for another.
- A stakeholder insists on a percentage. What do you give them?A weighted absolute percentage error: total absolute error divided by total actual minutes across the test set. Aggregating the denominator once means no single short call can distort it, and the figure is driven by the long waits that carry the operational cost. Quote MAE in minutes beside it.
Scoring an error as a percentage of the actual is like judging lateness against the length of the errand: five minutes late on a two-minute walk to the postbox reads as a catastrophe, while the same five minutes on an eight-hour flight reads as perfect.
saying these in an interview costs you the question
- Claims MAPE treats over and under prediction symmetrically
- Says MAPE is fine as long as no actual is exactly zero
- Believes an absolute percentage error cannot exceed 100%
- Takes sMAPE at its name and calls it genuinely symmetric
- Silently drops the rows where the actual is zero