skip to content

In an ARIMA(p,d,q) model, what do the p, d, and q orders each control?

level: middleimportance: must knowfreq 82%

answer

  1. three orders, three separate mechanisms
  2. one uses the series' own past values
  3. one uses the model's own past errors
  4. the middle order changes the data, not the parameters

basics

~20 s

In ARIMA(p,d,q), p is how many of the series' own past values the model regresses on, d is how many times the series is differenced before fitting, and q is how many past forecast errors enter the equation.

solid answer

~40 s

The three orders name three separate mechanisms. `d` is the integrated part: the series is differenced `d` times, so with `d = 1` the model is fitted to `y_t - y_{t-1}` and forecasts are cumulated back to the original scale. `p` is the autoregressive order: the current (differenced) value is regressed on its own `p` most recent values, with coefficients `phi_1 ... phi_p`. `q` is the moving-average order: the current value also depends on the `q` most recent one-step forecast errors, with coefficients `theta_1 ... theta_q`. Those errors are the model's own unobserved shocks, not a smoothing window over past observations — the name misleads people constantly. So ARIMA(2,1,0) regresses first differences on their previous two values, while ARIMA(0,1,1) drives first differences purely off the last shock.

go deeper

for a junior

Be ready to say what the three letters stand for and to name ARIMA(0,1,0) as a random walk. Knowing that d means differencing and p means lagged values already clears the screening bar.

for a middle

Explain the actual equations: AR regresses on lagged values, MA on lagged errors, and the errors are unobserved. Be able to count parameters for a given order and say what the differencing costs.

for a senior

Show that you connect orders to forecast behaviour — which order flattens the forecast, which produces a straight line, and why intervals keep widening once d is at least one. Interviewers probe that link.

for a principal

Own the parsimony argument: extra AR and MA terms buy little beyond the first few horizons while destabilising estimates, so defend a small order as a deliberate choice rather than a limitation.

## The three orders are three different mechanisms ARIMA(p, d, q) is a model for one time series expressed in terms of its own history. The three integers do not measure the same thing on different scales — each switches on a distinct piece of machinery. Writing `B` for the backshift operator (`B y_t = y_{t-1}`) makes the whole model compact. ### d — the integrated part `d` is the number of times the series is differenced before an ARMA structure is fitted. First differencing replaces `y_t` with `z_t = y_t - y_{t-1}`; second differencing differences that result again, `z_t = (y_t - y_{t-1}) - (y_{t-1} - y_{t-2})`. In backshift form the differenced series is `(1 - B)^d y_t`. With `d = 0` the model works on the levels themselves. Two consequences follow. Each difference consumes one usable observation, so `d = 1` on 120 points leaves 119 rows to fit. And whatever the model forecasts is on the differenced scale, so forecasts must be cumulated — integrated — back to the original units. That summation is what the "I" in ARIMA names. ### p — the autoregressive part An AR(p) component regresses the current value of the differenced series on its own `p` immediately preceding values: `z_t = c + phi_1 z_{t-1} + phi_2 z_{t-2} + ... + phi_p z_{t-p} + e_t` Each `phi` is an estimated coefficient and `e_t` is a white-noise shock. This is ordinary regression where the predictors happen to be lags of the response. `p` counts predictors, so it also counts parameters. ### q — the moving-average part An MA(q) component makes the current value a weighted combination of the most recent `q` shocks: `z_t = c + e_t + theta_1 e_{t-1} + ... + theta_q e_{t-q}` The crucial point is that the `e` terms are not columns in your data. They are the model's own one-step-ahead errors, recovered recursively during estimation. An MA term therefore says "a surprise last period still moves this period by a fixed fraction, and then it is gone." It is emphatically not a moving-average smoother over past observations — that shared name is the single most common source of confusion in the topic. ### The combined model `(1 - phi_1 B - ... - phi_p B^p) (1 - B)^d y_t = c + (1 + theta_1 B + ... + theta_q B^q) e_t` AR polynomial on the left acting on the differenced series, MA polynomial on the right acting on the shocks. ### Special cases worth recognising instantly - ARIMA(0,0,0): white noise around a constant mean. - ARIMA(1,0,0): a first-order autoregression on the levels. - ARIMA(0,1,0): a random walk. Every forecast equals the last observation, flat forever. Add a constant and you get a random walk with drift, whose forecast is a straight line. - ARIMA(0,1,1): first differences driven by the previous shock — a very common, very robust two-parameter workhorse for series with a wandering level. - ARIMA(1,1,1): one lag of the differenced level and one lag of the shock. ### How the orders shape a forecast `d` decides the long-run shape. With `d = 0` and a constant, forecasts converge to the estimated mean. With `d = 1` and no constant, they flatten out at a level near the end of the series. With `d = 1` and a constant, or `d = 2` without one, they follow a straight line. `p` and `q` shape only the first few steps: an MA(q) term contributes nothing beyond horizon `q`, and AR effects decay geometrically toward the long-run shape. Prediction intervals behave similarly. For `d = 0` the interval width converges to a finite band around the mean; for `d >= 1` it keeps widening with the horizon, because the shocks accumulate rather than wash out. ### Counting parameters An ARIMA(p,d,q) estimates `p + q` coefficients, plus the innovation variance and optionally a constant. `d` costs no parameter — it costs data. That asymmetry matters when you compare candidate orders: raising `p` or `q` buys flexibility at a fixed price in parameters, while raising `d` changes the very quantity being modelled. ### Where people go wrong Saying `d` counts lagged predictors; describing MA as averaging past observations; assuming higher `p` and `q` always forecast better (extra terms usually widen intervals and destabilise estimates); and forgetting to reverse the differencing so a forecast comes back in percent-change units instead of the original scale.

  • What does an ARIMA(0,1,0) with no constant forecast?
    A flat line at the last observed value. With p = q = 0 there is nothing to model in the first differences, so the best forecast of every future difference is zero and the level stays where it ended. That is the random walk, and it is the baseline any richer order has to beat.
  • How many parameters does an ARIMA(2,1,1) estimate?
    Three coefficients: two autoregressive and one moving-average, plus the innovation variance and optionally a constant. The differencing order costs no parameter at all — it costs one observation and changes which series is being modelled, which is why it is not interchangeable with adding an AR lag.
  • Why is calling q a 'moving average' order misleading?
    Because it does not average past observations. An MA(q) term is a weighted sum of the model's last q one-step forecast errors, which are unobserved quantities backed out during estimation. A smoothing window over past values would be a completely different, and far less useful, model.

An AR term remembers where the series was; an MA term remembers how wrong the model was. The differencing order decides which quantity — the level or its change — is being remembered at all.

saying these in an interview costs you the question

  • Says d counts the number of lagged predictors
  • Describes MA terms as averaging past observations
  • Thinks larger p and q always improve forecasts
  • Forgets to reverse differencing before reporting a forecast
  • Treats p and d as interchangeable ways to add memory

context