In simple exponential smoothing, what does the smoothing parameter alpha control?
answer
- one number, the weight on the newest point
- geometric decay over all past data
- high alpha chases, low alpha smooths
- mean age is (1 - alpha) / alpha
- forecast is flat at the last level
basics
~20 sAlpha sets how much weight the update puts on the newest observation versus the accumulated past. Alpha near 1 tracks recent data and reacts fast; alpha near 0 averages over a long history and smooths noise.
solid answer
~50 sSimple exponential smoothing keeps a single state, the level, and updates it each period as `level_t = alpha * y_t + (1 - alpha) * level_{t-1}`, with the forecast for every future horizon equal to the current level. Alpha, between 0 and 1, is the weight on the newest observation. Unrolling the recursion shows the weight on the observation j periods back is `alpha * (1 - alpha)^j`, so every past point still counts but its influence decays geometrically. The mean age of the data in that weighted average is `(1 - alpha) / alpha`. On a noisy weekly demand series, alpha = 0.05 gives an effective memory of about 19 weeks and rides straight through the wiggles but is slow to notice a genuine step change; alpha = 0.9 reacts to a step almost immediately but inherits nearly all the week-to-week noise. Alpha is normally estimated by minimising the sum of squared one-step-ahead errors, not chosen by eye.
go deeper
Be ready to write the level update, say which direction alpha moves responsiveness, and state that the forecast is flat at every horizon. Knowing that alpha lives between 0 and 1 is the minimum.
Expect to unroll the recursion to the geometric weights alpha times (1 - alpha)^j, quantify effective memory with the mean age (1 - alpha) / alpha, and explain that alpha is fitted by minimising squared one-step errors.
Show you read a fitted alpha as a diagnostic: a value near 1 says random-walk behaviour, near 0 says a stable level, and a boundary value plus structured residuals says the model form is wrong for the series.
Own the operating consequence: a responsive alpha churns downstream plans and widens long-horizon intervals, while a sluggish one delays detection of real demand shifts. Decide which cost the business would rather carry, and set the policy explicitly.
## The method in one recursion Simple exponential smoothing (SES) is the forecasting method for a series with no trend and no seasonal pattern — a quantity that wanders around a slowly-moving level. It carries exactly one piece of state, the **level**, and updates it every period: ``` level_t = alpha * y_t + (1 - alpha) * level_{t-1} ``` Here `y_t` is the observation at time t and `alpha` is the **smoothing parameter**, a number between 0 and 1. The forecast made at time t for any horizon h is flat: ``` forecast_{t+h} = level_t for every h = 1, 2, 3, ... ``` An equivalent and very useful way to write the update is in error-correction form. If `e_t = y_t - level_{t-1}` is the one-step-ahead forecast error, then `level_t = level_{t-1} + alpha * e_t`: alpha is the fraction of each surprise that you fold permanently into your view of the level. ## What alpha actually does: geometric weights Substituting the recursion into itself repeatedly gives ``` level_t = alpha*y_t + alpha(1-alpha)*y_{t-1} + alpha(1-alpha)^2*y_{t-2} + ... + (1-alpha)^t * level_0 ``` The weight on the observation j periods ago is `alpha * (1 - alpha)^j`. Two consequences follow, and both are commonly probed in interviews. First, **nothing is ever discarded**. Unlike a moving average of the last N points, which drops the (N+1)-th point entirely, SES keeps every observation with a small but non-zero weight. That is why the method is sometimes described as an infinite weighted average of the past. Second, **alpha sets the effective memory**. Two summary numbers make this concrete: - Mean age of the data in the weighted average: `(1 - alpha) / alpha`. - Half-life of the weights: `ln(0.5) / ln(1 - alpha)` periods. For alpha = 0.05 the mean age is 19 periods and the weights halve about every 13.5 periods. For alpha = 0.9 the mean age is roughly 0.11 periods and the weights are effectively gone after one or two points. Low alpha means a long memory; high alpha means a short one. The single most common error is to invert this — to say that a small alpha means the model ignores the past, when a small alpha is precisely what makes the past dominate. ## The tradeoff on a real series Take a noisy weekly demand series for one product. Week-to-week variation is a mixture of two things: transient noise (a promotion at a competitor, a delivery landing on Monday instead of Sunday) and genuine, persistent shifts in the level (a price cut, a new distribution channel). With alpha = 0.05 the model treats almost every movement as noise. The fitted line is smooth and the forecast is stable, which is exactly what a downstream planner wants — but when demand genuinely steps up, the model takes many weeks to catch up and forecasts too low the whole time. With alpha = 0.9 the model treats almost every movement as signal. It catches the step change within a week or two, but its forecast series is nearly as jumpy as the raw data, so ordinary one-step errors are large and the plan churns. That is the whole bias-versus-responsiveness tradeoff, controlled by one number. ## How alpha is chosen Alpha (and the initial level `level_0`) are estimated from the data, typically by minimising the sum of squared one-step-ahead errors over the training sample, or equivalently by maximising the likelihood of the corresponding state-space model. Interviewers like to hear that this is a fitted quantity, not a hand-set knob, and that the fitted value is itself a diagnostic: - alpha close to 1 says almost every movement persists — the series behaves like a random walk, and the naive forecast (last observed value) is nearly as good. - alpha close to 0 says almost every movement is noise around a stable level, so a long-run average is nearly as good. - alpha pinned to a boundary, or residuals that still show structure, usually means the series has a trend or a seasonal pattern that a level-only model cannot represent, and you need a richer form. ## Uncertainty The point forecast is flat, but the uncertainty is not. In the additive-error state-space form of SES, the variance of the h-step-ahead forecast is `sigma^2 * (1 + (h - 1) * alpha^2)`, where `sigma^2` is the one-step error variance. Intervals therefore widen with the horizon, and they widen faster when alpha is large — a responsive model is also a more uncertain one far out. A flat point forecast surrounded by widening intervals is the correct picture of SES, and candidates who describe the forecast as a single line with constant uncertainty have missed half the model. ## Common mistakes Calling alpha a significance level, confusing high alpha with more smoothing, believing that only the last few points get weight, or asserting that SES can handle trend. It cannot: with a trending series the flat forecast is biased low (or high) at every horizon, and that is the cue to move to a trend form of the method.
- How is alpha usually chosen in practice?It is estimated, not set by hand: pick alpha and the initial level to minimise the sum of squared one-step-ahead errors on the training sample, or equivalently to maximise the likelihood of the state-space form. Hand-setting a round number like 0.3 is only defensible as a deliberate stability constraint, and you should say so.
- What does the simple exponential smoothing forecast look like five steps ahead?Identical to one step ahead. With no trend and no seasonal state there is nothing to extrapolate, so every horizon returns the final level and the point forecast is a flat line. The prediction intervals still widen with the horizon, so the picture is a flat line inside a spreading band.
- What happens at the extremes, alpha = 1 and alpha = 0?At alpha = 1 the level equals the last observation, so the method collapses to the naive forecast. At alpha = 0 the level never updates and the forecast stays at the initial level forever. Fitted values close to either extreme are a signal to check whether a different model form fits the series better.
Alpha is how stubborn you are. A stubborn forecaster (small alpha) barely moves an opinion after one surprising week; an impressionable one (large alpha) rewrites it on the spot.
saying these in an interview costs you the question
- Calls alpha a significance level or a p-value
- Says a small alpha means old data is ignored
- Claims alpha = 0.9 produces the smoother forecast
- Believes only the last few observations get any weight
- Thinks simple exponential smoothing extrapolates a trend