skip to content

Forecast Evaluation

Scoring a forecast honestly: splits that respect the arrow of time, rolling-origin backtests, and error measures like MASE that survive zeros and scale changes. Shuffled folds cheat here.

on this pageshow

explore

questions

10

Why does MAPE break down on intermittent demand series with many zero-sales days?

level: juniorimportance: must knowfreq 76%

answer

  1. look at what sits in the denominator
  2. some periods have zero demand
  3. one direction of error has no ceiling
  4. worst under-forecast costs exactly 100%

basics

~20 s

MAPE divides each absolute error by the actual value, so a zero-sales day makes that term undefined and a near-zero day makes it explode. On intermittent demand the average is dominated by a handful of tiny denominators.

solid answer

~50 s

MAPE is the mean over held-out periods of `|actual - forecast| / |actual|`, usually reported as a percentage. Spare-part or slow-moving SKU demand is intermittent: most days are zero and the rest are small integers. A zero actual makes the term undefined outright, and an actual of 1 against a forecast of 3 contributes 200%, so the mean is driven by the smallest actuals rather than by the periods that matter commercially. MAPE is also asymmetric: because the actual sits in the denominator, an under-forecast can never cost more than 100% per period, while an over-forecast is unbounded. Optimising MAPE therefore quietly rewards forecasts that sit too low. On series with zeros I would score with an error scaled by a naive benchmark, or with a volume-weighted absolute error such as `sum|actual - forecast| / sum(actual)`, and report the baseline alongside it.

go deeper

for a junior

Be ready to write the formula and say out loud that the actual is the denominator, so a zero actual makes the term undefined and a tiny actual makes it huge.

for a middle

Explain the asymmetry mechanically: an under-forecast caps at 100% per period while an over-forecast is unbounded, and name one defined alternative for series with zeros.

for a senior

Show the operating consequence. A MAPE-tuned forecast biases low, which turns into stockouts, and the common fixes each redefine what is being reported — say which one is in use.

for a principal

Own the reporting standard. Decide whether percentage-shaped metrics belong in a scorecard at all when the portfolio contains intermittent series, and make the choice explicit rather than per-team.

## What MAPE is Mean absolute percentage error is defined over a held-out window of n periods as ``` MAPE = (100 / n) * sum_t |y_t - f_t| / |y_t| ``` where `y_t` is the actual for period t and `f_t` the forecast for it. Each period contributes an absolute percentage error (APE) — the error expressed as a fraction of that period's actual — and MAPE is the plain average of those fractions. Its appeal is real: percentages are unit-free, so a business audience can compare a forecast of pallets against a forecast of euros, and "we are 12% off" needs no explanation. That appeal is also why it is over-used on series where it is not defined. ## Why intermittent demand breaks it Intermittent demand means the series is zero most of the time and small when it is not — the classic case is spare parts, where a given part sells on a few days a year. Three things go wrong. **Undefined terms.** If `y_t = 0`, the ratio `|y_t - f_t| / 0` is undefined. Any non-zero forecast gives a division by zero; even a perfect forecast of 0 gives 0/0. Tooling usually papers over this by dropping those periods or substituting a small constant, and both choices change what is being measured. Dropping zero days removes exactly the days the forecast is most often wrong about, so the reported number describes a subset chosen by the outcome. **Explosion near zero.** Even without exact zeros, small actuals dominate. An actual of 1 with a forecast of 3 contributes 200%. An actual of 400 with a forecast of 380 contributes 5%. Averaged together the day with one unit of demand outweighs the day with four hundred, although the second carries eighty times the commercial error in units. MAPE has no upper bound, so a single low-volume period can move the headline number by tens of percentage points. **Asymmetry.** Because the denominator is the actual and not the forecast, the two directions of error are not treated alike. For a positive actual, the worst possible under-forecast is `f_t = 0`, which contributes exactly 100%. An over-forecast has no ceiling: forecasting 3 when the actual is 1 costs 200%, forecasting 10 costs 900%. A model tuned to minimise MAPE will therefore drift low, which on a spare-part inventory means systematic stockouts. This is a genuine bias in the metric, not a modelling artefact. ## What people reach for instead, and what it fixes **sMAPE** replaces the denominator with the average of actual and forecast, `2|y_t - f_t| / (|y_t| + |f_t|)`. That bounds each term (at 200% in this parameterisation) and removes the division-by-zero when only the actual is zero, but it is still undefined when both actual and forecast are zero, and despite the name it is not symmetric either — it still treats equal-sized over- and under-forecasts differently. It is a patch, not a fix. **Scaling by a benchmark's error.** Dividing the mean absolute error by the mean absolute error of a naive rule computed on the training data gives a unit-free number that stays finite as long as the series is not perfectly repetitive. Zeros in the data are harmless because the zeros never enter a denominator on their own; only the benchmark's average error does. **Volume-weighted absolute error.** Summing errors and actuals separately, `sum_t |y_t - f_t| / sum_t y_t`, gives a percentage-shaped number defined whenever total demand over the window is positive. It weights each period by its size, which is usually what the business meant when it asked for "percentage accuracy". **Errors in units.** For a single series read by people who know the units, mean absolute error in pieces per day is often more honest than any ratio, precisely because it does not pretend to be comparable across SKUs. ## How to talk about it in an interview The strong answer is not "MAPE is bad". It is: MAPE is a ratio whose denominator is the actual, so its behaviour is governed entirely by how small actuals can get. On a smooth, strictly positive, high-volume series it is a reasonable and very legible metric. On a series that touches zero it is undefined, unbounded and directionally biased, and the usual mitigations — dropping zero periods, adding an epsilon to the denominator, switching to sMAPE — each silently redefine the quantity being reported. Say which mitigation you chose and what it changed, or pick a metric that is defined on the data you actually have. One more practical habit: whichever error measure you report, report a naive benchmark's score on the same held-out window next to it. A MAPE of 30% means nothing until you know the benchmark scored 28% or 90%.

  • Does sMAPE actually fix the zero problem?
    Only partly. sMAPE uses `2|y - f| / (|y| + |f|)`, so a zero actual with a positive forecast is defined and each term is bounded. But it is still undefined when actual and forecast are both zero, which happens constantly on intermittent series, and it is not symmetric despite the name: equal-magnitude over- and under-forecasts get different penalties. It changes the failure mode rather than removing it.
  • Why does minimising MAPE push forecasts downward?
    For a positive actual, the largest possible under-forecast error is 100%, reached when the forecast is zero. Over-forecasting has no ceiling — twice the actual costs 100%, ten times costs 900%. The penalty surface is steeper above the actual than below it, so the MAPE-optimal forecast sits below the middle of the predictive distribution. On inventory that means chronic stockouts.
  • Someone adds 1 to every actual so MAPE is always defined. What do you say?
    That it silently changes the metric. Adding a constant to the denominator makes low-volume periods look accurate by construction: an actual of 0 forecast as 2 now scores 200% instead of undefined, and an actual of 1 becomes far more forgiving. The number is no longer a percentage of anything, and it is not comparable to a MAPE computed anywhere else. If you must do it, name the constant in the report.

It is like grading drivers on percentage of speed limit exceeded, then including a road where the limit is zero. The rule is fine on a motorway and meaningless in the car park.

saying these in an interview costs you the question

  • Calls MAPE scale-free and therefore always safe to use
  • Claims adding a small epsilon to the denominator fixes zeros
  • Says MAPE treats over- and under-forecasting symmetrically
  • Thinks an absolute percentage error cannot exceed 100%
  • Drops zero-demand periods without saying so in the report

context

open as a page

Why is a shuffled random train/test split invalid for evaluating a daily demand forecast?

level: juniorimportance: must knowfreq 78%

basics

~20 s

A shuffled split trains on future days and tests on past ones, so the model interpolates between neighbouring dates instead of forecasting. Because adjacent days are highly correlated, the score looks excellent and says nothing about future performance.

open as a page

How does MASE scale forecast errors, and what does a MASE above 1 mean?

level: middleimportance: must knowfreq 58%

basics

~20 s

MASE divides the forecast's mean absolute error by the mean absolute error of a one-step naive rule computed on the training data. The result is unit-free, and a value above 1 means the model averaged larger errors than that naive benchmark.

open as a page

In rolling-origin backtesting, how do expanding and sliding training windows differ?

level: middleimportance: must knowfreq 66%

basics

~20 s

An expanding window keeps its start fixed so the training set grows at every origin; a sliding window keeps a fixed length and drops the oldest data. Expanding uses more history, sliding adapts faster after a regime change.

open as a page

An 80% prediction interval covers only 55% of held-out actuals. How do you diagnose it?

level: seniorimportance: should knowfreq 44%

basics

~20 s

Empirical coverage far below the nominal level means the intervals are too narrow. Recompute coverage per forecast horizon, check whether the interval width grows with horizon, and confirm the shortfall is larger than the sampling noise in the count of covered points.

open as a page

In a 14-day-ahead backtest, why put a 14-day gap between training and test data?

level: seniorimportance: should knowfreq 44%

basics

~20 s

Because a row dated within 14 days of the forecast origin has a target that lands at or after the origin, so it was not yet observed there. Dropping those rows keeps training to outcomes genuinely known at prediction time.

open as a page

Is a last-year holdout a valid backtest if you tuned the model on the full history?

level: seniorimportance: should knowfreq 50%

basics

~20 s

No. If the last year influenced which model or hyperparameters you picked, its error is an optimistic in-sample number for that choice, not an independent estimate. Selection must happen on origins entirely before the holdout begins.

open as a page

How do you choose the headline forecast accuracy metric for a portfolio of thousands of series?

level: principalimportance: should knowfreq 36%

basics

~20 s

Pick a scale-free measure so series of different sizes can be aggregated, weight by business value rather than series count, and publish a naive baseline's score on the same window so the number reads as skill, not difficulty.

open as a page

Why does minimising pinball loss at the 0.9 quantile give a higher forecast than minimising MAE?

level: seniorimportance: nice to knowfreq 27%

basics

~20 s

Pinball loss weights error directions unequally: at the 0.9 level, falling short costs nine times as much per unit as overshooting. Its minimiser is the 0.9 quantile, while absolute error is minimised by the median.

open as a page

How should a rolling-origin backtest reflect how often the model is refit in production?

level: principalimportance: nice to knowfreq 30%

basics

~20 s

The backtest should refit on the same schedule the deployed system uses. Refitting at every origin while production retrains quarterly reports the accuracy of a model far fresher than the one that will actually be serving forecasts.

open as a page