skip to content

Simple exponential smoothing is equivalent to which ARIMA model?

level: middleimportance: nice to knowfreq 26%

answer

  1. difference once, one error term
  2. no autoregressive part at all
  3. the mapping is one minus the other
  4. write it in error-correction form first
  5. the sign convention decides the formula

basics

~20 s

Simple exponential smoothing is equivalent to ARIMA(0,1,1): difference the series once and model the result with a single moving-average term. Writing that as y_t - y_{t-1} = e_t - theta * e_{t-1}, the smoothing weight is alpha = 1 - theta.

solid answer

~50 s

Start from the smoothing recursion in error-correction form. If `f_t` is the forecast made for time t and `e_t = y_t - f_t` is the one-step error, then `f_{t+1} = f_t + alpha * e_t`, and since `f_t = y_t - e_t` this rearranges to `y_t - y_{t-1} = e_t - (1 - alpha) * e_{t-1}`. That is exactly an ARIMA(0,1,1): one difference, no autoregressive terms, one moving-average term with `theta = 1 - alpha`. The equivalence explains where the method's flat forecast comes from — a differenced series with a short-memory error term has no level to revert to — and it supplies proper prediction intervals for what is otherwise a bare recursion. The correspondence extends: Holt's linear trend matches ARIMA(0,2,2) and the damped-trend form matches ARIMA(1,1,2). It does not extend everywhere, though: models with multiplicative errors or multiplicative seasonality have no ARIMA counterpart, and stationary ARIMA models have no smoothing counterpart.

go deeper

for a junior

Be able to name the model — ARIMA(0,1,1) — and say that it means one difference and one error term, with no autoregressive part. The parameter mapping can wait.

for a middle

Expect to derive it: rewrite the smoothing update in error-correction form, substitute, and land on a differenced series with one lagged error. State the mapping and name the sign convention you used.

for a senior

Show where the correspondence stops. Multiplicative-error and multiplicative-seasonal forms have no counterpart, and stationary models have no smoothing counterpart, so neither family contains the other.

for a principal

Own the practical point: the equivalence means a team choosing between the two families is often choosing a parameterisation and a tooling habit, not a modelling philosophy. Steer that debate toward what each form makes easy to automate and explain.

## The two families Exponential smoothing and ARIMA are usually taught as separate traditions: one a recursive updating rule invented for practical inventory forecasting, the other a class of linear stochastic processes derived from stationarity theory. They overlap. The linear, additive-error smoothing models each correspond to a specific ARIMA model, and the simplest case is the cleanest. ## Deriving the equivalence Simple exponential smoothing forecasts one step ahead as ``` f_{t+1} = alpha * y_t + (1 - alpha) * f_t ``` Define the one-step error `e_t = y_t - f_t`, so `f_t = y_t - e_t`. Substituting: ``` f_{t+1} = alpha * y_t + (1 - alpha) * (y_t - e_t) = y_t - (1 - alpha) * e_t ``` Shifting the index back one period gives `f_t = y_{t-1} - (1 - alpha) * e_{t-1}`. Now write the observation as forecast plus error: ``` y_t = f_t + e_t = y_{t-1} - (1 - alpha) * e_{t-1} + e_t ``` and rearrange: ``` y_t - y_{t-1} = e_t - (1 - alpha) * e_{t-1} ``` The left side is the first difference of the series. The right side is a linear combination of the current error and one lagged error. That is the definition of an integrated model of order one with a single moving-average term — ARIMA(0,1,1) — written in the convention `(1 - B) y_t = e_t - theta * e_{t-1}`, with ``` theta = 1 - alpha, equivalently alpha = 1 - theta ``` Be explicit about the sign convention when you state this, because the opposite convention writes the moving-average term with a plus sign and reports the mapping as `theta = alpha - 1`. Both are the same model; only the bookkeeping differs. An interviewer will accept either if you say which one you are using. ## What the parameter range means A moving-average model is invertible — expressible as a convergent infinite autoregression — when `|theta| < 1`. Under `alpha = 1 - theta`, that maps to `0 < alpha < 2`. The conventional smoothing range `0 < alpha < 1` therefore sits strictly inside the invertible region, which is why smoothing practitioners rarely bump into invertibility as a constraint. The state-space treatment of the model admits the wider range, though alphas above 1 are unusual in applied work. The two boundaries are instructive. At alpha = 1, theta = 0 and the model reduces to `y_t - y_{t-1} = e_t`, a pure random walk whose optimal forecast is the last observation — exactly the naive forecast that simple exponential smoothing collapses to at alpha = 1. At alpha near 0, theta near 1, the differenced series is close to non-invertible, which is the signature of a series that was over-differenced: it did not need differencing because its level barely moves. ## The other view: geometric weights The same model has a third representation. Unrolling the recursion gives ``` f_{t+1} = alpha*y_t + alpha(1-alpha)*y_{t-1} + alpha(1-alpha)^2*y_{t-2} + ... + (1-alpha)^{t}*f_1 ``` So the forecast is a weighted average of all past observations with weights decaying geometrically at rate `1 - alpha`. Three descriptions — a recursive update, an infinite geometric weighting, and a differenced single-moving-average model — describe one object. Being able to move between them is what the question is really testing. ## How far the correspondence goes Among the linear additive-error smoothing models: - Simple exponential smoothing corresponds to ARIMA(0,1,1). - Holt's linear trend corresponds to ARIMA(0,2,2). - The additive damped-trend method corresponds to ARIMA(1,1,2). Seasonal additive smoothing models also have integrated counterparts, but with constraints imposed on the parameters, so the smoothing model is a restricted special case rather than a free re-parameterisation. The correspondence is not a bijection between the families: - Smoothing models with **multiplicative errors** or **multiplicative seasonality** are non-linear and have no ARIMA equivalent at all. - Conversely, **stationary** ARIMA models — anything with d = 0, such as a pure autoregression around a fixed mean — have no exponential smoothing counterpart, because smoothing models are built around a level that moves rather than a mean that pulls the series back. So neither family contains the other; they intersect in the linear, additive-error, non-stationary middle. ## Why it matters in an interview Three reasons. It shows the smoothing recursion is not an ad-hoc heuristic but the optimal forecast of a well-defined stochastic process, which is what licenses likelihood estimation and prediction intervals. It explains the flat forecast function without hand-waving. And it is a useful sanity check when two colleagues fit different families to the same series and get suspiciously similar numbers — sometimes they have fitted the same model twice. ## Common mistakes Saying the equivalent model is autoregressive of order one (it is not — the level is not mean-reverting), forgetting the differencing, reporting the sign of theta without naming the convention, or claiming that every exponential smoothing model has an ARIMA twin.

  • Do all exponential smoothing models have an ARIMA equivalent?
    No. Only the linear, additive-error ones do. Models with multiplicative errors or multiplicative seasonality are non-linear and have no ARIMA counterpart. The relationship also fails the other way: stationary ARIMA models with no differencing have no smoothing counterpart, because smoothing is built around a moving level rather than a fixed mean.
  • What does the invertibility condition on the moving-average term imply for alpha?
    Invertibility requires the moving-average parameter to satisfy `|theta| < 1`. Under `alpha = 1 - theta` that corresponds to `0 < alpha < 2`, so the conventional smoothing range from 0 to 1 is comfortably inside it. That is why invertibility is rarely a binding concern when fitting simple exponential smoothing.
  • Why should you name the sign convention when stating the parameter mapping?
    Because two conventions are in common use for writing the moving-average term, one with a minus sign and one with a plus. They describe the same model, but the reported mapping flips between `alpha = 1 - theta` and `theta = alpha - 1`. Stating the equation you are using removes the apparent contradiction.

saying these in an interview costs you the question

  • Says the equivalent model is a first-order autoregression
  • Omits the differencing and names a stationary model
  • Claims every smoothing model has an ARIMA twin
  • Thinks the moving-average term averages past observations
  • Quotes a sign for theta without naming the convention

context