skip to content

Time-Series Analysis

Data ordered in time breaks the i.i.d. assumption: trend and seasonality, stationarity, autocorrelation, ARIMA and Holt-Winters forecasts. Interviewers use it to find who peeks ahead.

on this pageshow

explore

questions

page 1 of 2

Why does MAPE break down on intermittent demand series with many zero-sales days?

level: juniorimportance: must knowfreq 76%

answer

  1. look at what sits in the denominator
  2. some periods have zero demand
  3. one direction of error has no ceiling
  4. worst under-forecast costs exactly 100%

basics

~20 s

MAPE divides each absolute error by the actual value, so a zero-sales day makes that term undefined and a near-zero day makes it explode. On intermittent demand the average is dominated by a handful of tiny denominators.

solid answer

~50 s

MAPE is the mean over held-out periods of `|actual - forecast| / |actual|`, usually reported as a percentage. Spare-part or slow-moving SKU demand is intermittent: most days are zero and the rest are small integers. A zero actual makes the term undefined outright, and an actual of 1 against a forecast of 3 contributes 200%, so the mean is driven by the smallest actuals rather than by the periods that matter commercially. MAPE is also asymmetric: because the actual sits in the denominator, an under-forecast can never cost more than 100% per period, while an over-forecast is unbounded. Optimising MAPE therefore quietly rewards forecasts that sit too low. On series with zeros I would score with an error scaled by a naive benchmark, or with a volume-weighted absolute error such as `sum|actual - forecast| / sum(actual)`, and report the baseline alongside it.

go deeper

for a junior

Be ready to write the formula and say out loud that the actual is the denominator, so a zero actual makes the term undefined and a tiny actual makes it huge.

for a middle

Explain the asymmetry mechanically: an under-forecast caps at 100% per period while an over-forecast is unbounded, and name one defined alternative for series with zeros.

for a senior

Show the operating consequence. A MAPE-tuned forecast biases low, which turns into stockouts, and the common fixes each redefine what is being reported — say which one is in use.

for a principal

Own the reporting standard. Decide whether percentage-shaped metrics belong in a scorecard at all when the portfolio contains intermittent series, and make the choice explicit rather than per-team.

## What MAPE is Mean absolute percentage error is defined over a held-out window of n periods as ``` MAPE = (100 / n) * sum_t |y_t - f_t| / |y_t| ``` where `y_t` is the actual for period t and `f_t` the forecast for it. Each period contributes an absolute percentage error (APE) — the error expressed as a fraction of that period's actual — and MAPE is the plain average of those fractions. Its appeal is real: percentages are unit-free, so a business audience can compare a forecast of pallets against a forecast of euros, and "we are 12% off" needs no explanation. That appeal is also why it is over-used on series where it is not defined. ## Why intermittent demand breaks it Intermittent demand means the series is zero most of the time and small when it is not — the classic case is spare parts, where a given part sells on a few days a year. Three things go wrong. **Undefined terms.** If `y_t = 0`, the ratio `|y_t - f_t| / 0` is undefined. Any non-zero forecast gives a division by zero; even a perfect forecast of 0 gives 0/0. Tooling usually papers over this by dropping those periods or substituting a small constant, and both choices change what is being measured. Dropping zero days removes exactly the days the forecast is most often wrong about, so the reported number describes a subset chosen by the outcome. **Explosion near zero.** Even without exact zeros, small actuals dominate. An actual of 1 with a forecast of 3 contributes 200%. An actual of 400 with a forecast of 380 contributes 5%. Averaged together the day with one unit of demand outweighs the day with four hundred, although the second carries eighty times the commercial error in units. MAPE has no upper bound, so a single low-volume period can move the headline number by tens of percentage points. **Asymmetry.** Because the denominator is the actual and not the forecast, the two directions of error are not treated alike. For a positive actual, the worst possible under-forecast is `f_t = 0`, which contributes exactly 100%. An over-forecast has no ceiling: forecasting 3 when the actual is 1 costs 200%, forecasting 10 costs 900%. A model tuned to minimise MAPE will therefore drift low, which on a spare-part inventory means systematic stockouts. This is a genuine bias in the metric, not a modelling artefact. ## What people reach for instead, and what it fixes **sMAPE** replaces the denominator with the average of actual and forecast, `2|y_t - f_t| / (|y_t| + |f_t|)`. That bounds each term (at 200% in this parameterisation) and removes the division-by-zero when only the actual is zero, but it is still undefined when both actual and forecast are zero, and despite the name it is not symmetric either — it still treats equal-sized over- and under-forecasts differently. It is a patch, not a fix. **Scaling by a benchmark's error.** Dividing the mean absolute error by the mean absolute error of a naive rule computed on the training data gives a unit-free number that stays finite as long as the series is not perfectly repetitive. Zeros in the data are harmless because the zeros never enter a denominator on their own; only the benchmark's average error does. **Volume-weighted absolute error.** Summing errors and actuals separately, `sum_t |y_t - f_t| / sum_t y_t`, gives a percentage-shaped number defined whenever total demand over the window is positive. It weights each period by its size, which is usually what the business meant when it asked for "percentage accuracy". **Errors in units.** For a single series read by people who know the units, mean absolute error in pieces per day is often more honest than any ratio, precisely because it does not pretend to be comparable across SKUs. ## How to talk about it in an interview The strong answer is not "MAPE is bad". It is: MAPE is a ratio whose denominator is the actual, so its behaviour is governed entirely by how small actuals can get. On a smooth, strictly positive, high-volume series it is a reasonable and very legible metric. On a series that touches zero it is undefined, unbounded and directionally biased, and the usual mitigations — dropping zero periods, adding an epsilon to the denominator, switching to sMAPE — each silently redefine the quantity being reported. Say which mitigation you chose and what it changed, or pick a metric that is defined on the data you actually have. One more practical habit: whichever error measure you report, report a naive benchmark's score on the same held-out window next to it. A MAPE of 30% means nothing until you know the benchmark scored 28% or 90%.

  • Does sMAPE actually fix the zero problem?
    Only partly. sMAPE uses `2|y - f| / (|y| + |f|)`, so a zero actual with a positive forecast is defined and each term is bounded. But it is still undefined when actual and forecast are both zero, which happens constantly on intermittent series, and it is not symmetric despite the name: equal-magnitude over- and under-forecasts get different penalties. It changes the failure mode rather than removing it.
  • Why does minimising MAPE push forecasts downward?
    For a positive actual, the largest possible under-forecast error is 100%, reached when the forecast is zero. Over-forecasting has no ceiling — twice the actual costs 100%, ten times costs 900%. The penalty surface is steeper above the actual than below it, so the MAPE-optimal forecast sits below the middle of the predictive distribution. On inventory that means chronic stockouts.
  • Someone adds 1 to every actual so MAPE is always defined. What do you say?
    That it silently changes the metric. Adding a constant to the denominator makes low-volume periods look accurate by construction: an actual of 0 forecast as 2 now scores 200% instead of undefined, and an actual of 1 becomes far more forgiving. The number is no longer a percentage of anything, and it is not comparable to a MAPE computed anywhere else. If you must do it, name the constant in the report.

It is like grading drivers on percentage of speed limit exceeded, then including a road where the limit is zero. The rule is fine on a motorway and meaningless in the car park.

saying these in an interview costs you the question

  • Calls MAPE scale-free and therefore always safe to use
  • Claims adding a small epsilon to the denominator fixes zeros
  • Says MAPE treats over- and under-forecasting symmetrically
  • Thinks an absolute percentage error cannot exceed 100%
  • Drops zero-demand periods without saying so in the report

context

open as a page

In a time series, what is the difference between a point outlier and a level shift?

level: juniorimportance: must knowfreq 70%

basics

~20 s

A point outlier is one observation far from expectation, after which the series returns to its old level. A level shift moves the series to a new baseline that persists. The difference decides whether you drop a point or re-baseline.

open as a page

Why is a shuffled random train/test split invalid for evaluating a daily demand forecast?

level: juniorimportance: must knowfreq 78%

basics

~20 s

A shuffled split trains on future days and tests on past ones, so the model interpolates between neighbouring dates instead of forecasting. Because adjacent days are highly correlated, the score looks excellent and says nothing about future performance.

open as a page

In a classical time-series decomposition, what do the trend, seasonal and remainder components each represent?

level: juniorimportance: must knowfreq 76%

basics

~20 s

Decomposition splits a series into three parts: trend-cycle, the slow movement of the level; seasonal, the pattern that repeats at a fixed known period such as 12 months; and remainder, the variation left after removing both.

open as a page

In simple exponential smoothing, what does the smoothing parameter alpha control?

level: juniorimportance: must knowfreq 76%

basics

~20 s

Alpha sets how much weight the update puts on the newest observation versus the accumulated past. Alpha near 1 tracks recent data and reacts fast; alpha near 0 averages over a long history and smooths noise.

open as a page

In a daily demand forecasting model, what is a lag-7 feature and why include it?

level: juniorimportance: must knowfreq 74%

basics

~20 s

A lag-7 feature is the target value from seven days earlier, placed on today's row as an input column. It hands the model last week's same-weekday demand, capturing weekly seasonality that a lag-1 feature alone misses.

open as a page

What does it mean for a time series to be weakly stationary?

level: juniorimportance: must knowfreq 82%

basics

~20 s

Weak stationarity means the mean is constant over time, the variance is constant and finite, and the covariance between two observations depends only on the gap between them, not on where in the series they sit.

open as a page

How does MASE scale forecast errors, and what does a MASE above 1 mean?

level: middleimportance: must knowfreq 58%

basics

~20 s

MASE divides the forecast's mean absolute error by the mean absolute error of a one-step naive rule computed on the training data. The result is unit-free, and a value above 1 means the model averaged larger errors than that naive benchmark.

open as a page

How do you score anomalies in a strongly seasonal daily metric without flagging every Monday?

level: middleimportance: must knowfreq 65%

basics

~20 s

Score the residual against a seasonal baseline instead of the raw value: compare each day with the same weekday one cycle earlier, then flag unusually large residuals. A Monday is only anomalous relative to other Mondays.

open as a page

In an ARIMA(p,d,q) model, what do the p, d, and q orders each control?

level: middleimportance: must knowfreq 82%

basics

~20 s

In ARIMA(p,d,q), p is how many of the series' own past values the model regresses on, d is how many times the series is differenced before fitting, and q is how many past forecast errors enter the equation.

open as a page

What is the difference between the ACF and the PACF of a time series?

level: middleimportance: must knowfreq 78%

basics

~20 s

The ACF measures the correlation between a series and its own values k steps earlier, including everything routed through the intermediate lags. The PACF measures that same lag-k correlation after the effect of lags 1 through k-1 has been removed.

open as a page

In rolling-origin backtesting, how do expanding and sliding training windows differ?

level: middleimportance: must knowfreq 66%

basics

~20 s

An expanding window keeps its start fixed so the training set grows at every origin; a sliding window keeps a fixed length and drops the oldest data. Expanding uses more history, sliding adapts faster after a regime change.

open as a page

Why can regressing one random walk on an unrelated random walk give a high R-squared?

level: middleimportance: must knowfreq 58%

basics

~20 s

A random walk never returns to a mean, so over any window it looks trended and least squares lines up the two drifts. The residuals stay non-stationary, so standard errors are far too small and the fit statistics are meaningless.

open as a page

How do you decide between an additive and a multiplicative time-series decomposition?

level: middleimportance: must knowfreq 68%

basics

~10 s

Look at whether the seasonal swing grows with the level. Constant-size swings mean additive; swings that widen as the level rises mean multiplicative. Taking logs turns a multiplicative series into an additive one.

open as a page

In Holt-Winters, when do you choose multiplicative rather than additive seasonality?

level: middleimportance: must knowfreq 62%

basics

~20 s

Choose multiplicative seasonality when the size of the seasonal swing grows with the level of the series, and additive when the swing stays roughly the same absolute size no matter how high the level is. Multiplicative needs strictly positive data.

open as a page

Why does first differencing turn a random walk into a stationary series?

level: middleimportance: must knowfreq 68%

basics

~20 s

A random walk accumulates every past shock, so its variance grows with time and its covariance depends on the date rather than the lag. Differencing once strips the accumulation away and leaves the shock itself, which is stationary noise.

open as a page

Your forecast model's residuals fail a Ljung-Box test at lag 10 — what does that mean?

level: seniorimportance: must knowfreq 58%

basics

~20 s

It means the residuals still carry autocorrelation somewhere in the first ten lags, so the model has left predictable structure behind. The Ljung-Box null is that all those residual autocorrelations are zero; a small p-value rejects it.

open as a page

How does a centred rolling mean feature leak future values into a training row?

level: seniorimportance: must knowfreq 62%

basics

~20 s

A centred window straddles the row: a 7-day centred mean at day t averages t-3 through t+3, so three future values enter that row. The model then trains on inputs that will not exist at prediction time.

open as a page

What do the dashed confidence bands on an ACF correlogram represent?

level: juniorimportance: should knowfreq 52%

basics

~20 s

They mark how large a sample autocorrelation can get by chance if the series were white noise. The band is plus or minus 1.96 over the square root of n, so a bar inside it is indistinguishable from noise.

open as a page

What does a cross-correlation peak at lag -14 between daily ad spend and sign-ups tell you?

level: juniorimportance: should knowfreq 46%

basics

~20 s

The two series align best at a two-week offset. Under the usual convention, where the lag shifts ad spend, spend from 14 days earlier tracks today's sign-ups, so spend leads. Confirm the sign convention before claiming a direction.

open as a page

What do the seasonal orders (P,D,Q)[m] add to a SARIMA model on monthly data?

level: middleimportance: should knowfreq 57%

basics

~20 s

The seasonal orders repeat the autoregressive, differencing and moving-average structure at multiples of the season length m. On monthly data with m = 12, P uses lag 12, D differences values a year apart, and Q carries last year's shock.

open as a page

What does an ACF with large spikes at lags 12 and 24 tell you about monthly data?

level: middleimportance: should knowfreq 45%

basics

~20 s

Spikes at lags 12 and 24 on monthly data mean each month resembles the same month one and two years earlier: an annual cycle of period 12. A seasonal period shows up at its multiples.

open as a page

What does it mean to say oil prices Granger-cause airline stock returns?

level: middleimportance: should knowfreq 54%

basics

~20 s

Past oil prices improve the forecast of airline returns beyond what the returns' own history explains. It is predictive precedence, established by testing whether lagged oil terms jointly add explanatory power, not evidence of a causal mechanism.

open as a page

How does a centred moving average estimate the trend-cycle of a seasonal series?

level: middleimportance: should knowfreq 54%

basics

~20 s

Average each point over a window exactly one seasonal period long, so every season appears once and the seasonal effect cancels, leaving the local level. Even periods need a two-step centred average, and the first and last half-window get no trend.

open as a page

Why add a 28-day rolling mean and rolling standard deviation alongside raw lag features?

level: middleimportance: should knowfreq 60%

basics

~20 s

Rolling aggregates compress recent history into stable columns. A 28-day trailing mean gives the current level with daily noise averaged out; the rolling standard deviation gives recent volatility. A single lag carries one noisy day and can express neither.

open as a page

In the augmented Dickey-Fuller test, what does failing to reject the null mean?

level: middleimportance: should knowfreq 60%

basics

~20 s

The augmented Dickey-Fuller null is that the series has a unit root, so failing to reject only means a unit root could not be ruled out. That is weak evidence of non-stationarity, never proof of it.

open as a page

When does a monthly series need seasonal differencing at lag 12 rather than a first difference?

level: middleimportance: should knowfreq 42%

basics

~20 s

Seasonal differencing is needed when a repeating annual pattern, not a drifting level, is what moves the mean. Subtracting the value twelve months earlier removes a stable yearly pattern that a lag-1 difference leaves intact.

open as a page

An 80% prediction interval covers only 55% of held-out actuals. How do you diagnose it?

level: seniorimportance: should knowfreq 44%

basics

~20 s

Empirical coverage far below the nominal level means the intervals are too narrow. Recompute coverage per forecast horizon, check whether the interval width grows with horizon, and confirm the shortfall is larger than the sampling noise in the count of covered points.

open as a page

A metric shows a permanent level shift the day of a tracking-SDK release. How do you test whether the cause is instrumentation, not users?

level: seniorimportance: should knowfreq 50%

basics

~20 s

Test whether the drop appears in an independent measurement. Instrumentation breaks confine themselves to one SDK version, platform or event pipeline; a real behaviour change also shows in server-side or revenue counts the SDK never touched.

open as a page

When a SARIMAX model includes exogenous regressors, what must you supply to forecast?

level: seniorimportance: should knowfreq 43%

basics

~20 s

Future values of every exogenous regressor, for every period you forecast. A SARIMAX forecast is conditional on those inputs, so calendar or planned values are safe while values you must forecast yourself add error the intervals never show.

open as a page

showing 1–30 of 53