skip to content

How do you choose the headline forecast accuracy metric for a portfolio of thousands of series?

level: principalimportance: should knowfreq 36%

answer

  1. the number becomes the incentive
  2. scale-free before it can be averaged
  3. window difficulty versus model quality
  4. always score the simple rules alongside
  5. one mean cannot describe thousands of series

basics

~20 s

Pick a scale-free measure so series of different sizes can be aggregated, weight by business value rather than series count, and publish a naive baseline's score on the same window so the number reads as skill, not difficulty.

solid answer

~50 s

A portfolio headline number has to survive three problems: series of wildly different scale, series where the metric is undefined, and an audience that will act on a single figure. Averaging per-series percentage errors fails all three, because low-volume series produce enormous percentages and dominate. I would report an error scaled by a naive benchmark, such as mean absolute scaled error, or a volume-weighted absolute error `sum|actual - forecast| / sum(actual)` when the audience needs a percentage. Whichever is chosen, publish the seasonal-naive, naive and drift baselines scored on the same window and horizons, and express the model as a skill score `1 - error_model / error_baseline`. A headline mean also hides the tail, so pair it with the share of series where the model loses to the baseline — that number, not the mean, tells you where to invest. And weight by revenue or volume, because a thousand dormant SKUs should not outvote the twenty that fund the business.

go deeper

for a junior

Be ready to say why a single accuracy percentage across many series can mislead, and that a simple benchmark score is needed to know whether a number is good.

for a middle

Explain the mechanics of the two aggregation choices: an average of per-series scores versus a volume-weighted total, and which part of the catalogue each one lets dominate.

for a senior

Show that you would score naive, seasonal-naive and drift baselines on the same window and horizons, report skill, and break out the share of series losing to the baseline.

for a principal

Own the metric as an incentive. Argue which number goes on the scorecard, what it would take to game it, how it aligns with the actual decision cost, and why it must stay frozen long enough for its trend to be readable.

## The question behind the question At portfolio scale the metric stops being a statistical choice and becomes an incentive system. Whatever appears on the weekly slide is what teams will optimise, what budget conversations will hinge on, and what a model refresh will be judged by six months from now. So the design criteria are not only mathematical. ## Criterion one: aggregable across scale Series in a real catalogue differ by orders of magnitude. Any metric averaged across them must be dimensionless, and it must be dimensionless in a way that does not systematically favour one size of series. Averaging per-series percentage error fails badly here. Percentage errors are largest exactly where the actuals are smallest, so the mean across a catalogue is dominated by dormant and low-volume series, and the number moves whenever the long tail is added or pruned. Errors scaled by a benchmark's error avoid this, because the scaling constant is per-series and in the same units as that series. The alternative that keeps percentage legibility is a volume-weighted absolute error: sum the absolute errors across all series and periods, sum the actuals across the same cells, and divide. It is defined whenever total demand is positive, it weights each cell by its size, and it is usually what a stakeholder meant when they asked for percentage accuracy. Its weakness is the mirror image: it is dominated by the largest series, so a small number of high-volume products can hide broad failure in the tail. Neither is correct in the abstract. The right move is to state which weighting the headline uses, and to publish the other one beside it. ## Criterion two: a baseline on the same window An accuracy number alone measures the difficulty of the series as much as the quality of the model. A portfolio of stable weekly series will post a flattering number under any method; a portfolio of promotion-driven, intermittent series will look terrible under a very good one. Comparing this quarter's figure to last quarter's is only meaningful if the mix of series and the volatility of the window are unchanged, which they never are. The fix is to score simple benchmarks on exactly the same held-out window and horizons, and report the model relative to them. Three are worth carrying: - **Naive**: every future period equals the last observed value. - **Seasonal naive**: every future period equals the value one seasonal period earlier — last Tuesday for daily data with a weekly cycle, last December for monthly data with a yearly one. - **Drift**: the last observed value plus the average per-period change over the training sample, extrapolated across the horizon; equivalently, extending the straight line through the first and last training points. A skill score, `1 - error_model / error_baseline`, then reads directly: positive means better than the baseline, 0.2 means the error is 20% smaller, negative means the baseline won. Skill absorbs the difficulty of the window, because a hard quarter inflates both terms. It is common and healthy for the seasonal-naive baseline to beat a carefully tuned model on a meaningful share of series. That is a finding, not an embarrassment. It identifies the series where the machinery is not paying for itself, and routing those to the baseline is often the cheapest accuracy win available. ## Criterion three: the distribution, not just the centre One number cannot describe thousands of series. The mean is dragged by outliers, the median hides them. The two summaries worth putting next to the headline are the share of series where the model loses to its baseline, and a small set of percentiles of the per-series score. Those answer the questions a lead actually has: is this broadly working, and where is it not? Segmenting by lifecycle stage, volume band and intermittency usually explains most of the spread, and each segment often deserves a different treatment rather than a better global model. ## Criterion four: alignment with the decision A symmetric error metric assumes being short and being long cost the same. If the downstream decision is inventory or capacity, they do not, and a portfolio scored on symmetric error will be optimised toward a target the business never wanted. Where the asymmetry is real, the scorecard should carry a quantile-aware score at the operating level, alongside the symmetric one. And if the forecast is consumed as a distribution, calibration and interval width belong on the same scorecard as the point error. ## Criterion five: it must be hard to game and stable over time Any single number becomes a target. Ask, for each candidate, what the cheapest way to improve it without improving decisions would be — dropping hard series from the denominator, forecasting low where the metric is asymmetric, quietly shortening the horizon. Then close those doors: fix the series set, fix the horizons, and define the window rule once so the series and the metric are not both moving. Also keep the definition frozen for long enough that a trend means something. A metric changed every two quarters produces a chart nobody can read, which is worse than an imperfect metric held constant. ## The shape of a good answer Name one headline number, say why it is scale-free, pair it with a same-window baseline and a skill score, add the share of series losing to the baseline, weight by business value, and state which decisions the metric is and is not aligned with. The judgment being tested is whether you understand that the metric will be optimised, and whether you have chosen one where doing so is good for the business.

  • The seasonal-naive baseline beats the tuned model on a third of the series. What do you do?
    Treat it as a routing decision, not a failure. Identify what those series have in common — usually short history, intermittency or heavy promotion effects — and serve them from the baseline while the model keeps the series where it earns skill. Then report the share explicitly each period, because it is the cleanest measure of whether the modelling investment is still paying for itself.
  • Stakeholders insist on a percentage. What do you give them?
    A volume-weighted absolute error: total absolute error divided by total actuals over the same cells. It is percentage-shaped, defined whenever total demand is positive, and immune to the near-zero blowup that ruins an average of per-series percentages. State plainly that it is weighted by volume, so improvement on the largest series moves it most, and publish the tail separately.
  • How do you stop a headline accuracy metric from being gamed?
    Fix the things a team could otherwise adjust: the series set, the horizons scored, and the definition of the held-out window. Ask of every candidate metric what the cheapest improvement that helps nobody would be, and close that path. Then pair the headline with a baseline skill score, since improvements that come from an easier window raise the baseline too and cancel out.
  • Why weight by revenue rather than treating every series equally?
    Because an unweighted mean gives a dormant SKU the same vote as one carrying a large share of the business, and dormant series are both numerous and hard to forecast. Equal weighting therefore reports mostly on the part of the catalogue nobody acts on. Weighting by revenue or volume aligns the number with consequence, with the tail reported separately so it is not simply lost.

saying these in an interview costs you the question

  • Averages per-series percentage errors across the whole catalogue
  • Reports an accuracy number with no baseline on the same window
  • Compares this quarter's number to last quarter's ignoring mix changes
  • Uses a mean across series and never inspects the spread
  • Treats a baseline beating the model as a result to hide

context