Your team wants to replace exponential smoothing baselines with deep learning across 5,000 SKUs. How do you decide?
answer
- the baseline is the thing to beat
- per-segment, not one global verdict
- competition evidence favours simple defaults
- cross-learning needs many related series
- count the operational cost, not just accuracy
basics
~20 sKeep exponential smoothing as the benchmark the new method must beat on the same series. Complex forecasters tend to win only with many related series, useful covariates and enough history per series; otherwise damped-trend smoothing is very hard to beat.
solid answer
~50 sFrame it as a challenge, not a migration: the smoothing baseline stays in place and the new method has to demonstrate an advantage on the same series before anything is switched off. The published forecasting competitions found damped-trend exponential smoothing among the most accurate methods across large heterogeneous collections, and the winning entry in the most recent large competition combined exponential smoothing with a neural component rather than discarding smoothing. The conditions under which a learned model genuinely wins are specific: many related series that share structure, useful covariates such as price and promotion, and enough history per series. A 5,000-SKU catalogue usually contains a long tail of short, intermittent series where those conditions fail. So the answer is normally a portfolio — a learned global model where it earns its place, smoothing everywhere else, per-series selection between them, and the operational costs of the complex option counted honestly.
go deeper
Be ready to say why a simple baseline exists at all: it is the number the fancier method has to beat, and without it nobody can tell whether the fancier method helped.
Expect to name the conditions under which each family does better — covariates and many related series favour a learned model, short or intermittent series favour smoothing — rather than picking a side outright.
Show you would run the comparison fairly: tune both sides, report per segment rather than pooled, and check that the challenger's inputs are genuinely available at forecast time.
Own the portfolio decision and its economics. Decide which catalogue segments justify the complexity, itemise retraining, monitoring and ownership costs, keep the baseline as the operational fallback, and state in advance the condition that reverses the decision.
## Reframe the question 'Should we replace smoothing with deep learning' is the wrong shape. Forecasting 5,000 SKUs is not one problem, it is 5,000 problems with wildly different data. The decision that matters is per-series or per-segment, and the baseline's role is to be the thing every candidate must beat. A lead is expected to hold that line calmly and without being anti-modern about it. ## What the evidence actually says The large public forecasting competitions are the only broad, honest comparison anyone has. Two findings from them are directly relevant. The M3 competition compared a wide range of methods over a large and varied collection of series. Damped-trend exponential smoothing came out among the most accurate entries, and a headline conclusion of that work was that statistical sophistication did not systematically translate into forecast accuracy. That result is why damped-trend smoothing became the standard automatic baseline rather than a footnote. The later, much larger M4 competition did see machine learning matter — but the winning entry combined exponential smoothing with a neural component, and strong entries came from combinations of methods rather than from a single learned model that discarded the statistical toolkit. The lesson is not 'neural networks lose'. It is that they earn their keep by adding to a well-specified statistical structure and by borrowing strength across series, not by replacing the structure. ## When a learned global model genuinely wins Be specific about the conditions, because they are checkable against your own catalogue: - **Many related series with shared structure.** A single global model trained across SKUs can learn a promotion response or a launch curve from the SKUs that have one and apply it to the SKUs that do not. This is real and is the main mechanism behind learned forecasters winning. - **Covariates that matter.** Price, promotion calendars, stock-outs, holidays, weather. Smoothing models have no mechanism for these at all, and this is the honest weak point of the baseline. - **Enough history per series, or enough series.** Cross-learning can compensate for short individual series, but only if the population is large and genuinely homogeneous in the relevant way. - **Stable regime.** A learned model trained on last year's demand generation mechanics is worth less after a channel or pricing overhaul. ## Where the baseline keeps winning - **Short series.** A SKU launched four months ago cannot support anything much, and smoothing degrades gracefully to a level-only form. - **Intermittent demand.** Long runs of zeros break most smooth-forecast assumptions; neither family handles them well, but the failure is cheaper to diagnose in the simple one. - **The long tail.** In most catalogues a small fraction of SKUs carries most of the volume. The tail is where a learned model's marginal accuracy gain is smallest and its opacity costs the most. - **Explainability under challenge.** When a planner asks why the forecast moved, a level and a slope answer the question. That matters more than teams expect during a stock-out post-mortem. ## Costs the proposal usually omits - **Retraining and monitoring.** A global learned model needs a training schedule, feature pipelines that must not leak future information, and monitoring for drift. Smoothing needs neither. - **Failure modes.** A smoothing model fails visibly and locally, on one series. A global model fails silently and everywhere at once, which is a materially worse operational profile for a planning system. - **People.** Someone must own the model in year three. A baseline that any analyst can read is cheap insurance against that person leaving. - **Latency and cost per refresh** at 5,000 series, multiplied by however often the plan is refreshed. ## The decision I would make Run the comparison rather than arguing about it. Keep the smoothing baseline as the incumbent, apply the same evaluation protocol to both, and require the challenger to show a material, consistent gain — not an average gain driven by a handful of high-volume series. Segment the result: the head of the catalogue, where covariates exist and history is long, is where a learned model most often earns its place. The tail usually does not justify it. Ship a portfolio, with per-series selection between the candidates, and keep the baseline running underneath as the fallback when the complex model is unavailable or its inputs are late. Then state the reversal condition explicitly: if the challenger stops beating the baseline on a segment, that segment goes back. A benchmark you are willing to lose to is the only kind worth keeping. ## Common mistakes Treating the choice as a technology preference rather than an empirical question; evaluating on the head of the catalogue and generalising to the tail; comparing a tuned new model against an untuned baseline; and forgetting that the baseline is also the fallback that keeps the plan running when the sophisticated pipeline breaks.
- What conditions make a learned global model genuinely worth the complexity?Many related series that share structure so the model can borrow strength across them, covariates such as price and promotion that smoothing cannot use at all, enough history or enough series to train on, and a stable regime. Absent those, the extra machinery buys little and costs a training pipeline, drift monitoring and an owner.
- How would you avoid an unfair comparison in favour of the new method?Tune both sides, not just the challenger, and apply the identical evaluation protocol to each. Report results per segment rather than one pooled average, since a gain concentrated in a few high-volume series can hide losses across the tail. And make sure no covariate available to the challenger encodes information unavailable at forecast time.
- Why keep the smoothing model running even after the new one wins?It is the fallback. A global learned model depends on feature pipelines and a training schedule, and when an input is late or a pipeline breaks the planning process still needs numbers. A per-series smoothing model has almost no dependencies and degrades gracefully, so it costs little to keep and prevents an outage from becoming a planning failure.
- How do you set the bar for what counts as a win?Define it before running the comparison: a material improvement, consistent across segments and across repeated evaluation periods, large enough to justify the operational cost you have already itemised. Also state the reversal condition, so a segment where the challenger stops winning goes back to the baseline without a fresh argument.
The baseline is the incumbent in an election, not a placeholder. The challenger does not win by being newer; it wins by carrying the vote in each district it claims.
saying these in an interview costs you the question
- Treats the choice as a technology preference, not an empirical test
- Compares a tuned new model against an untuned baseline
- Generalises head-of-catalogue results to the long tail
- Ignores retraining, monitoring and ownership costs
- Switches the baseline off instead of keeping it as fallback