Would you train one global model across thousands of store-item series, or one model per series?
answer
- borrow strength or stay independent
- what happens to a brand-new item
- one artifact versus thousands to operate
- big series dominate an unweighted loss
- normalise by each series' own level
basics
~20 sDefault to one global model trained on the pooled rows of all series, with a series identifier and static attributes as columns. It shares structure across series, covers brand-new ones, and leaves one artifact to operate. Reserve per-series models for high-value, atypical series.
solid answer
~50 sA global model stacks every series into one training table, one row per series-date, and adds identifiers and static attributes as columns. Cross-learning is the point: a new item with three weeks of history inherits the weekly and holiday shape learned from thousands of comparable items, which a per-series model cannot. The operational argument is just as strong — one artifact to train, version, monitor and roll back instead of thousands. The real difficulties are scale and dominance: high-volume series dominate an unweighted loss, so normalise each series by its own trailing level or model a relative target, and check error by series segment rather than only in aggregate. Keep a cheap per-series seasonal-naive baseline as a floor, and carve out dedicated models only where a series is both high-value and structurally unlike the rest. Deciding factors are series count, history length, and fleet heterogeneity.
go deeper
Know the distinction: a global model is trained on all series pooled together with an identifier column, while a local model is fit to one series' own history alone.
Explain how the pooled table is built, why lags and windows are computed within each series, and why a short or brand-new series is exactly where pooling helps most.
Show the practical hardening: per-series normalisation, segment-level error reporting, a naive floor kept live, and evidence rather than assertion when promoting a series to its own model.
Own the architecture decision and its consequences — retraining budget, on-call surface, release process for a model whose failure is fleet-wide — and be able to justify the split between global coverage and a small set of bespoke exceptions.
## The two architectures **Local, or per-series.** Every series gets its own fitted model. With 30,000 store-item combinations that is 30,000 parameter sets. Each is estimated only from its own history and knows nothing about any other series. **Global, or cross-learning.** One model is fit on the pooled rows of every series. The table is one row per series per timestamp, features are the usual lags, rolling aggregates and calendar terms computed *within* each series, plus identifiers and static attributes — item, category, store, region, price tier — that let the model condition on which series a row belongs to. The global architecture is now the default in large-scale demand forecasting, and the empirical evidence from open forecasting competitions on retail data of this shape supports it. But it is a genuine tradeoff, and the interviewer is testing whether the candidate can argue both directions. ## Why global usually wins **Cross-learning.** Most of what matters in a store-item series — weekend shape, holiday response, the ramp around a promotion, the sag after a price rise — is shared structure that appears in thousands of series. A local model rediscovers it independently from short, noisy history each time. A global model estimates it once from a vastly larger sample and applies it everywhere. **Short and cold-start series.** A series with six weeks of data cannot support a per-series seasonal model at all. Pooled, it borrows the shape from similar items and needs only a level anchor. New product launches are the extreme case: with no history, only static attributes and neighbours can carry a forecast, which is structurally impossible locally. **Operational surface.** One model is one training job, one artifact, one monitoring dashboard, one rollback. Thousands of models mean per-series training failures, per-series drift, per-series staleness, and a support burden that grows with the catalogue rather than with the team. **Capacity where it belongs.** A single large model with identifiers and attributes can allocate capacity to the series that need it, rather than spending equal parameters on the busiest and the sleepiest item in the catalogue. ## What goes wrong with global models **Volume dominance.** Under an unweighted absolute-error loss, the aggregate error is driven by the highest-volume series, and the model will happily trade away accuracy on the long tail. The standard remedies are to normalise each series by a statistic of its own history preceding the origin — divide by its trailing median or mean level, or predict a ratio to a naive baseline — and to weight the loss deliberately according to what the business actually values, which is rarely units. **Heterogeneity.** Pooling assumes the series are similar enough that a shared function is a good approximation. When the fleet contains genuinely different processes — intermittent spare parts alongside fast-moving groceries — one function is a compromise. Segmenting into a handful of global models by behaviour class is usually better than either extreme. **Blast radius.** A bad global release degrades every forecast at once. Per-series models fail independently, which is worse in expectation and better in the tail. The mitigation is staged rollout, a per-series champion-challenger comparison, and a cheap seasonal-naive baseline kept live as a floor so any catastrophic regression is detectable per series. **Explainability.** "Why is store 42's forecast down?" is harder to answer from a shared function than from a model fit to store 42 alone. This matters more than it should politically, and is worth planning for with per-series diagnostics. ## Where local still earns its place A handful of series that are both high-value and structurally unusual — a flagship location, a series with a contractual supply schedule, one product with a bespoke promotional calendar — can justify dedicated models. So can a small fleet: with fifty long, well-behaved series there is little cross-learning to gain and local models are simple and interpretable. Local classical models also remain excellent baselines, and the best practical answer is usually a global model that must beat a per-series naive or seasonal-naive benchmark on every segment before it ships. ## How to decide, and what to say Frame the answer around three measurable facts. *How many series, and how long is the typical history?* Many short series push hard towards global; few long series weaken the case. *How heterogeneous is the fleet?* High heterogeneity argues for segmented global models rather than one. *What does the business actually pay for?* If value concentrates in a hundred series out of thirty thousand, spend the attention there and let the tail be served cheaply. Then state the plan rather than the verdict: start global with per-series normalisation and a naive floor, evaluate by segment rather than in aggregate, and promote specific series to dedicated models only where the evidence and the value justify the extra operational surface. A principal-level answer also owns the consequences — the on-call surface, the retraining budget, the release process for a model whose failure is fleet-wide — because that, not the accuracy delta, is what the organisation lives with.
- How do you stop high-volume series from dominating a global model?Change the target scale or the loss, not just the features. Normalise each series by a statistic of its own history before the origin — its trailing median level — or predict a ratio to a naive baseline so every series contributes on a comparable scale. Then weight the loss to reflect what the business values, and evaluate by segment so long-tail degradation cannot hide inside an aggregate number.
- What can a global model do for a brand-new item that a per-series model cannot?Forecast at all. With no history there is nothing for a local model to fit, whereas a global model conditions on static attributes — category, price tier, store, pack size — and applies the shape learned from comparable items, needing only a level assumption. That cold-start capability is often the single strongest argument for pooling in a retail setting.
- What is the main operational risk of consolidating thousands of models into one?The blast radius. A bad release degrades every forecast simultaneously, where independent models fail one at a time. Mitigate with staged rollout, champion-challenger comparison against the previous version, and a cheap per-series seasonal-naive baseline kept live as a floor so a fleet-wide regression is detected per series rather than in an aggregate metric.
- When would you still fit a dedicated model for a single series?When it is both high-value and structurally unlike the fleet — a flagship site, a product with a contractual supply schedule, a series with its own promotional calendar. The test is whether the extra operational surface is repaid by measured accuracy on something the business actually cares about, evaluated against the global model on that series alone.
saying these in an interview costs you the question
- Assumes global always beats per-series regardless of fleet
- Ignores that volume dominates an unweighted pooled loss
- Trains globally without normalising each series' scale
- Judges the model only on aggregate error across all series
- Overlooks the fleet-wide blast radius of one bad release
- Claims per-series models can cover a cold-start item
- Skips a cheap per-series baseline as a floor