skip to content

How should you scale 3,000 SKU series whose volumes span four orders of magnitude for one shared forecaster?

level: seniorimportance: should knowfreq 50%

answer

  1. one set of weights, four orders of magnitude
  2. per series, never pooled
  3. the small ones flatten to a constant
  4. hand the size back as a feature
  5. fit on train, invert before reporting

basics

~20 s

Scale each series by its own statistics, not the pooled distribution. A global scaler squashes a two-unit-a-day SKU toward a constant while a 20,000-unit SKU dominates the loss. Fit each scale on training data only, and invert it on forecasts.

solid answer

~50 s

One set of recurrent weights sees every series, so inputs must be brought onto a comparable scale — but the scale is computed **per series**, never pooled. With a pooled mean and standard deviation, a SKU selling two units a day maps to a nearly constant value, its variation is numerically invisible, and a squared-error loss is dominated by the largest series; the model learns to predict a flat line for the tail. The usual scheme is to divide each series by its own training-period mean, optionally after a `log1p` transform for heavy skew, and to pass the log of that scale in as a static feature so the model can still condition on SKU size. Statistics come from the training portion only, and forecasts are inverted back into units before anyone reads them. New SKUs borrow a category-level scale.

go deeper

for a junior

Know that inputs of wildly different magnitudes have to be brought onto a comparable range before a neural network sees them, and that the transform must be undone before the forecast is used.

for a middle

Explain what a pooled scaler does to a low-volume series, name a per-series scheme such as mean scaling or log-then-centre, and describe how the scale is inverted.

for a senior

Show that scales are fitted on training periods only, that cold-start series borrow a group scale, and that series size is handed back to the model as a static feature.

for a principal

Own the fact that per-series scaling equalises the portfolio's loss weighting, and decide deliberately whether accuracy should be long-tail-fair or revenue-weighted.

## Why scaling is not optional here A global forecaster shares one set of weights across thousands of series. The recurrent cell has one input weight matrix, one set of saturating non-linearities and one learning rate, and it sees a value of 2 from one SKU and 20,000 from another. Without scaling, the large series drive the activations into saturation, the gradient magnitudes across the batch differ by orders of magnitude, and training is unstable or simply ignores the small series. ## Why a pooled scaler is the wrong fix The reflex is to fit one standardiser on all the data. Consider what that does with a pooled mean around, say, 400 units and a pooled standard deviation of several thousand. The 20,000-unit SKU occupies a wide range of scaled values. The two-unit SKU maps to something like -0.13 every single day; its entire dynamic range — the difference between a good day and a bad day — is a rounding error in scaled space, smaller than the noise the weight initialisation introduces. Under a squared-error loss the model gets almost no gradient from it, so the cheapest policy is to output the same near-constant for it forever. You will see this as a portfolio where the top hundred SKUs are forecast well and the long tail is a flat line. ## Per-series scaling schemes **Mean scaling.** Divide each series (and its targets) by that series' training-period mean, usually `1 + mean` to survive zeros. Every series then averages around 1 regardless of size. This is the workhorse for demand data: it is robust, cheap, and inverts by a single multiplication. **Per-series standardisation.** Subtract that series' own training mean and divide by its own training standard deviation. Works well for roughly symmetric series; less well for intermittent demand, where the standard deviation is dominated by rare spikes. **Log then centre.** For volumes spanning four orders of magnitude, `log1p` compresses the range before any centring and turns multiplicative variation into additive variation. Remember that a model trained on log targets and inverted naively predicts something closer to a median than a mean — decide whether that is what the business wants. **Window-relative (instance) scaling.** Divide each input window *and its target* by a statistic of that window — its mean, or its last value. Now the model predicts a ratio to the current level rather than an absolute magnitude, which handles trending series and level shifts, and lets forecasts exceed anything seen in absolute training units. You multiply the level back in at inference. ## Give the model back what scaling removed Scaling deliberately erases size, but size is informative: a 20,000-unit SKU is smoother, more predictable and differently seasonal than a two-unit one. Feed the log of the series scale back in as a static feature alongside category, price band and store. The model then learns size-dependent behaviour explicitly instead of having to infer it from magnitudes it can no longer see. ## Leakage, cold start and inversion - **Fit on training data only.** Computing a series' mean over its whole history, evaluation period included, leaks future level into every training row and flatters the backtest. Compute per series on the training window, store the scales, apply them unchanged afterwards, and recompute them at each scheduled refresh. - **Cold start.** A SKU with three weeks of history has no stable scale. Borrow one: the median scale of its category, price band or store cluster, and pass the group identity as a feature. Estimating a scale from a handful of noisy points is worse than inheriting a stable one. - **Invert before you report.** Errors measured in scaled units are meaningless to the business and not comparable across series. Multiply the forecast back into units before any accuracy or inventory calculation. ## The consequence nobody mentions Per-series scaling makes every series contribute roughly equally to the loss. That is usually what you want for long-tail accuracy, but it is a business decision disguised as preprocessing: a tiny SKU now counts as much as a flagship. If the objective is revenue-weighted error, re-weight the loss explicitly by volume or margin rather than letting the scaler choose the priority for you. Scaling is where a multi-series forecaster silently decides what it cares about.

  • Once every series is on a comparable scale, what has that done to the loss across the portfolio?
    It weights all 3,000 series roughly equally, so a tiny SKU counts as much as a flagship. That is often right for long-tail accuracy, but if the business measures revenue-weighted error you must re-weight the loss explicitly by volume or margin. Scaling quietly sets the portfolio's priorities, so make it a deliberate decision rather than a side effect.
  • How do you scale a SKU that launched three weeks ago?
    Borrow a scale rather than estimate one. Use the median scale of its category, price band or store cluster, and pass the group identity and the log scale in as static features so the model knows what it is looking at. Recompute from its own history at a later refresh, once there is enough of it to be stable.
  • A series trends steadily upward and the forecast plateaus at the top of its training range. Which scaling change helps?
    Scale relative to the recent window instead of the whole training history: divide the input window and its target by the window's own mean or last value, so the model predicts a ratio to the current level. Multiply the level back at inference. Level becomes context you supply rather than a magnitude the network has to have memorised.

saying these in an interview costs you the question

  • Fits one standard scaler on all series pooled together
  • Computes scaling statistics over the full history including the test period
  • Reports accuracy in scaled units and never inverts
  • Assumes per-series scaling means one model per series
  • Estimates a scale for a brand-new SKU from a few noisy days

context