With 120 daily observations from a single ATM, would you ship a recurrent forecaster or a classical model?
answer
- count the cycles, not the rows
- thousands of weights, hundreds of targets
- the baseline is last Tuesday
- tiny holdout, coin-flip comparison
- depth needs a portfolio, not a series
basics
~20 sShip the classical model. One short series yields roughly a hundred overlapping windows and about seventeen weekly cycles, far too little to fit thousands of recurrent weights. Deep forecasters earn their keep by pooling many related series, not on one short one.
solid answer
~40 sThe classical model, and I would say so before seeing any results. A 120-point daily series with a day-of-week cycle gives about seventeen repetitions of the pattern and, once windowed, roughly a hundred heavily overlapping examples. A single gated recurrent cell with a 32-unit hidden state already carries over four thousand weights — capacity orders of magnitude beyond the data. A seasonal statistical model fits a handful of parameters, gives usable intervals and needs no tuning budget, and seasonal-naive is the baseline both must beat. Second problem: with a hundred points every backtest fold is tiny, so an architecture comparison is mostly noise and a tuning search overfits the evaluation. What changes my answer is scale in the cross-section — four thousand ATMs with 120 days each, pooled into one model.
go deeper
Be able to say that a neural sequence model needs far more data than a short series provides, and that a simple seasonal baseline is the first thing you fit.
Compare the parameter count of a gated recurrent cell against the number of usable training windows, and explain why regularisation does not close that gap.
Show that you would not trust a model comparison on a 20-point holdout, and give the concrete conditions, above all a large pool of related series, under which you would revisit the choice.
Own the platform decision: a layered stack from naive baseline to classical default to pooled deep model, with an explicit revisit trigger and the maintenance cost of each layer priced in.
## Frame it as a data-budget question The instinct to reach for a sequence model comes from the sequence in the name. The right first question is how much data the model actually gets to learn from, and for a single 120-point daily series the answer is brutal. - **Cycles observed.** With a day-of-week pattern, 120 days is about seventeen repetitions. Everything the model can learn about weekly structure it must learn from seventeen noisy examples of it. - **Training examples.** Window it with a 14-day lookback and a 7-day horizon and you get roughly a hundred windows, each overlapping its neighbours almost completely. The effective sample size is nearer seventeen than a hundred. - **Parameters.** A vanilla recurrent cell with hidden size 32 on a single input channel carries 32 by 32 recurrent weights plus input weights and biases, about 1,100 parameters. A gated cell has four such parameter blocks, so over 4,000 — before the output layer, before a second layer. You are fitting thousands of free parameters to a few hundred target values. No amount of dropout or weight decay closes that gap. Regularisation narrows the hypothesis space; it does not manufacture information. ## What the classical model brings A seasonal statistical model — exponential smoothing with a weekly seasonal component, or a seasonal ARIMA fit — carries on the order of ten parameters. Its structure encodes the assumption that the future resembles a level, a trend and a repeating weekly pattern, which is exactly the prior a short series cannot afford to learn from scratch. It produces prediction intervals essentially for free, it fits in milliseconds so refitting nightly is trivial, and there is no tuning search to overfit. And before either: **seasonal-naive** — this Tuesday equals last Tuesday. On a strongly weekly ATM series that baseline is surprisingly hard to beat, and any model that does not beat it convincingly across several rolling origins should not be deployed. Reporting a neural model that loses to seasonal-naive is a bad day; not having run the baseline at all is worse. ## The evaluation trap With 120 points, whatever you hold out is small. A single 20-point holdout has an error estimate so noisy that the ranking between two models is close to a coin flip, and comparing five architectures on it guarantees the winner is the luckiest, not the best. Two consequences follow. First, prefer models whose selection requires few decisions — a classical family with an order-selection criterion makes far fewer choices than an architecture search. Second, be honest that on this data you cannot reliably detect a small accuracy improvement at all, which is itself an argument for the simpler, cheaper option. ## What would flip the decision - **Many related series.** This is the big one. Four thousand ATMs with 120 days each is half a million observations of the same underlying behaviour. A single global recurrent model pooled across them can learn payday effects, holiday effects and day-of-week shape once and share them, conditioning on per-ATM features for location and size. Cross-series learning, not sequence length, is what deep forecasters are actually for. The unit of data is the portfolio. - **Rich covariates and non-linear interactions.** If withdrawals depend jointly on paydays, holidays, weather and local events in ways that interact, a neural model with many input channels has something to offer that a small linear-in-structure model does not — but only once there is enough data to estimate those interactions. - **Long or multi-quantile horizons across many series.** One global model producing a full horizon and several quantiles for thousands of series at once can be cheaper to operate than thousands of individually fitted classical models. ## The organisational angle A neural forecaster for one series is a permanent operating cost — training infrastructure, a serving path, monitoring, someone who understands it — with no accuracy upside. The durable answer is a layered platform: seasonal-naive as the floor, a classical per-series model as the default that covers the short and sparse tail, and a pooled deep model introduced only where the portfolio is large enough to earn it, always benchmarked against the two below it on the same rolling origins. State the revisit trigger explicitly when you make the call — more series, new covariates, a longer horizon — so the decision is a dated engineering judgment rather than a veto on a technology.
- You now own 4,000 ATMs with the same 120 days each. Does the answer change?Yes. Pooling turns 120 points into roughly half a million, and one global recurrent model can learn the day-of-week and payday structure shared across machines while conditioning on per-ATM features. Cross-series learning is what deep forecasters are for. Keep the classical per-series model as the baseline the global model has to beat on the same rolling origins.
- How would you justify this to a stakeholder who wants the neural model?With numbers and a review date. Backtest seasonal-naive, the classical model and the recurrent model on the same origins, and put the operating cost of each next to its accuracy. Frame it as accuracy per unit of maintenance, and name the trigger that would reopen the decision, such as more series or new covariates, so it reads as a judgment rather than a veto.
- What would make even the classical model the wrong tool here?Withdrawals driven mainly by things the history cannot show: a nearby branch closing, a changed payday calendar, the machine being relocated. When level shifts have known external causes, a model with explicit covariates, or a simple rule plus human review, beats any purely historical extrapolator, neural or classical.
saying these in an interview costs you the question
- Reaches for the neural model because it is the newer method
- Believes dropout or weight decay can substitute for data
- Never fits the seasonal-naive baseline
- Compares architectures on one tiny holdout and declares a winner
- Assumes pooling unrelated series always helps