Would you fit one global model across 4,000 SKUs in 60 stores, or one model per series?
answer
- count the series before choosing
- most histories are short or sparse
- identity becomes a feature, not a model
- new products have no history to fit
- pooled loss favours the big sellers
basics
~20 sAlmost always one global model. Pooling 240,000 store-SKU series into a single training table, with store and product identity as features, shares patterns across short and sparse histories, handles new products, and leaves one artifact to operate instead of 240,000.
solid answer
~50 s4,000 SKUs across 60 stores is 240,000 series, and most of them are short, intermittent, or newly launched — nowhere near enough history each to fit a decent individual model. A global model trains on all of them stacked into one table, with lag and rolling features computed within each series and identity carried as features: store, product category, price tier, promotion flags. It borrows strength across series, forecasts a SKU launched last week from its attributes, and gives one model to train, monitor and explain. The costs are real: a single squared-error objective is dominated by the highest-volume series, so I would normalise each series by its own recent level or predict a ratio; and a global fit serves the average series well and the unusual ones less well. The pragmatic answer is usually a global model plus a handful of bespoke models for the few series where the money actually is.
go deeper
Know that many related series can be trained as one table with the series identity supplied as a feature, and that a brand-new product has no history of its own to learn from.
Explain the mechanics: rows stacked across series, lags computed within each series, identity and attributes as columns, and why pooling reduces variance for series with short histories.
Show that you would confront scale heterogeneity head-on with per-series normalisation, respect series boundaries in feature construction, and check performance separately on the long tail rather than in aggregate.
Own the allocation call: where the revenue sits, how much catalogue churn there is, what the team can actually operate, and whether segmentation into a few global models beats both extremes.
## The two designs **Local, or per-series:** every store-SKU combination gets its own model, fitted on its own history alone. This is the classical forecasting posture, and it is the right one when you have a handful of long, well-behaved series. **Global, or cross-learning:** all series are stacked into one table — each row still one forecast origin of one series — and a single model is fitted across the lot. Series identity does not disappear; it enters as features: store id, product id or its embedding-free encodings, category, price band, region, plus attributes of the pairing. With 4,000 SKUs in 60 stores, local means up to 240,000 fitted models and global means one. ## Why global usually wins at this scale **Data per series is the binding constraint.** A two-year daily history is 730 points. Many SKUs will have far less: launched recently, discontinued, stocked in only some stores, or selling in ones and twos with long runs of zeros. A model fitted on 730 noisy points cannot learn a holiday effect that occurs twice in the window. Pooled across 240,000 series, that same holiday appears hundreds of thousands of times. **Structure is shared.** Weekday shape, holiday lifts, promotional response and price sensitivity behave similarly within a category. A local model must rediscover each pattern from one series' noise; a global model estimates it once from all of them, which is a variance reduction, not a shortcut. **Cold start becomes tractable.** A SKU launched last week has no history to fit a local model on at all, and the fallback is a manual rule. A global model forecasts it from its attributes — category, price tier, store — plus whatever days exist. In a catalogue with constant churn, this alone often decides the design. **Operations.** One model is one thing to train, version, monitor and explain. 240,000 models mean 240,000 potential failures, and no one will look at them individually. The engineering burden, not accuracy, is frequently the deciding argument in practice. ## What global costs you **Scale heterogeneity.** A flagship SKU sells thousands of units a day; a slow one sells two. Under a squared-error objective the loss is dominated almost entirely by the high-volume series, and the model will happily be terrible on the long tail. Remedies: divide each series by its own recent mean and model the normalised value; predict a ratio to a baseline level; or use a loss whose errors are relative rather than absolute. Whichever you choose, it must be applied consistently at training and at prediction, and the inverse transform applied on the way out. **The average-series problem.** One functional form is fitted to a very heterogeneous population, so genuinely unusual series — a highly promoted flagship, a seasonal-only item — get the behaviour of the crowd. The mitigation is segmentation: fit a handful of global models over meaningful groups (fast movers versus intermittent items, food versus general merchandise) rather than one or 240,000. **Identity handling.** Series identity as a raw high-cardinality label invites the model to memorise individual series, which fails exactly where you need cold start to work. Prefer attributes that generalise — category, price band, store format, region — and treat the identity itself as one signal among many. **Feature construction must respect series boundaries.** Every lag and rolling window has to be computed within a series. A shift applied across the stacked table pulls one SKU's history into another's row, and at this scale nobody will notice by eye. ## Framing the decision as a lead The questions I would actually ask: 1. **Where is the money?** If a few hundred series carry most of the revenue, a global model for the tail plus bespoke attention for the head is a better allocation than either extreme. 2. **How much churn is in the catalogue?** High launch and discontinuation rates push hard toward global, because cold start dominates. 3. **What can the team operate?** A design that needs 240,000 models monitored is a design that will not be monitored. 4. **How heterogeneous are the series really?** If they are genuinely different processes — a warehouse's throughput and a store's footfall — pooling them buys nothing and segmentation is the honest answer. 5. **What does the tail cost when it is wrong?** If long-tail errors translate directly into stockouts, the loss weighting is a business decision, not a modelling detail. ## The answer that lands Don't present it as an ideological choice. State that global is the default at this scale for data, cold-start and operational reasons; name scale heterogeneity as the concrete thing that breaks it and normalisation as the fix; and reserve per-series models for the small set of high-value series where bespoke treatment and human oversight actually pay for themselves.
- What breaks when one global model pools series whose volumes differ by three orders of magnitude?A squared-error objective is dominated by the high-volume series, so the fit optimises them and neglects the long tail. The standard fixes are normalising each series by its own recent level and modelling the scaled value, predicting a ratio rather than a level, or choosing a loss whose errors are relative — always with the inverse transform applied on output.
- How does a SKU launched last week get a forecast under each approach?Per-series has essentially nothing to fit on and falls back to a manual rule or a category average. A global model forecasts it from its attributes — category, price tier, store, promotion status — combined with the few days that exist, because it has already learned how products like it behave. This cold-start advantage is often the deciding argument.
- When would you still keep per-series models?When there are few series, each with long history; when the series are genuinely different generating processes that share no structure; or for a small set of high-value series where a bespoke model plus human oversight pays for itself. In practice the answer is often a global model for the tail and bespoke treatment for the head.
- How would you narrow the gap between one global model and 240,000 local ones?Segment. Fit a handful of global models over groups that actually behave alike — fast movers versus intermittent items, perishables versus general merchandise — so each fit sees a more homogeneous population while still pooling thousands of series. It keeps most of the data-sharing benefit and most of the operational simplicity.
Fitting one model per series is teaching 240,000 students in separate rooms from one page of notes each; the global model teaches one class from the whole library and tells each student which desk they sit at.
saying these in an interview costs you the question
- Assumes one model per series is always more accurate
- Pools series of wildly different scale without normalising
- Computes lags across the stacked table ignoring series boundaries
- Ignores the cost of operating 240,000 models
- Treats series identity as a memorisable label, breaking cold start