In an ARIMA model, what do the stationarity and invertibility conditions require of its roots?
answer
- the polynomials have roots; location decides
- outside the circle, in the backshift convention
- for first-order models it collapses to a magnitude
- two MA coefficients share one autocorrelation function
basics
~20 sBoth require roots of a lag polynomial to lie outside the unit circle: the autoregressive polynomial's roots for stationarity, the moving-average polynomial's for invertibility. For a first-order model that reduces to the coefficient having absolute value below one.
solid answer
~40 sWrite the model as `phi(B) z_t = theta(B) e_t`, where `z` is the differenced series and `B` is the backshift operator. Stationarity of the autoregressive part requires every root of `phi(B) = 0` to lie strictly outside the unit circle in the complex plane; for AR(1) that is `|phi| < 1`. Invertibility requires the same of `theta(B) = 0`; for MA(1) that is `|theta| < 1`. A root on the circle means an unremoved unit root — forecasts never revert — and inside it means explosive growth. A moving-average root on the circle makes the model non-invertible: two coefficient values, `theta` and `1/theta`, imply identical autocorrelations, so invertibility is what picks the unique, forecastable one. Nearly shared roots between the two polynomials are a separate warning: the factors cancel, and the model is over-parameterised.
go deeper
Know the first-order versions: an AR(1) coefficient must be below one in absolute value to be stationary, and an MA(1) coefficient below one to be invertible. The general polynomial statement can wait.
State the condition in terms of roots of the lag polynomials and name your convention. Be able to describe the three regimes of an AR(1) coefficient and what each does to long-horizon forecasts.
Show you check fitted roots as a diagnostic: a boundary moving-average estimate and near-common autoregressive and moving-average factors are both signals to change the specification rather than accept the fit.
Own the standard for what a fitted model must satisfy before it goes anywhere near a decision, and be able to explain to a non-specialist why a model that fits well can still be unusable.
## Roots of the lag polynomials An ARMA model on a series `z` (already differenced, if it needed to be) can be written with the backshift operator `B`, where `B z_t = z_{t-1}`: `phi(B) z_t = theta(B) e_t` with `phi(B) = 1 - phi_1 B - ... - phi_p B^p` and `theta(B)` the analogous polynomial built from the moving-average coefficients. Both are ordinary polynomials in a complex variable, so both have roots, and the location of those roots is what the stationarity and invertibility conditions constrain. In this convention — polynomials in `B`, not in an eigenvalue — the requirement is that **all roots lie strictly outside the unit circle**, meaning modulus greater than one. Some texts factor the model differently and state the mirror-image condition on the reciprocal roots, which must then lie inside. Both say the same thing; state which convention you are using and the interviewer will follow, but mixing them mid-sentence is a genuine error. ### Stationarity, from the AR side For AR(1), `phi(B) = 1 - phi B` has the single root `B = 1/phi`, which sits outside the unit circle exactly when `|phi| < 1`. That is the familiar condition, and it maps onto three regimes: - `|phi| < 1`: shocks decay geometrically, the series reverts toward its mean, the variance is finite and constant, and long-horizon forecasts converge to that mean. - `|phi| = 1`: a unit root. The series is a random walk; shocks are permanent, the variance grows with time, and there is no mean to revert to. - `|phi| > 1`: explosive. Each shock is amplified, and forecasts diverge — essentially never what a real series does. When an estimation returns an autoregressive root sitting essentially on the circle, it is telling you that the series at the differencing level you chose is still not stationary. The model is straining to represent a random-walk-like level with a bounded parameter. ### Invertibility, from the MA side Invertibility is less intuitive but has a clean motivation. An invertible MA process can be rewritten as an infinite autoregression: the current value expressed in terms of past **observations** with geometrically declining weights. That matters because forecasting works from observed history — the shocks `e_t` are not data, they are reconstructed recursively, and that recursion only converges if the moving-average roots lie outside the unit circle. The second motivation is identification. For MA(1), the coefficients `theta` and `1/theta` generate exactly the same autocorrelation function, so the two are indistinguishable from the second-moment structure of the data. Requiring `|theta| < 1` picks one of the pair, and it is the one with the convergent autoregressive representation. Without the restriction the parameter simply is not identified. A moving-average estimate that lands exactly at the boundary — modulus one — deserves attention rather than a shrug. It means the fitted process cannot be inverted, the likelihood surface is flat against a wall at that point, and the reported standard error is not trustworthy. The usual reading is that the model has been asked to undo a transformation that was applied one time too many, and the specification should be reconsidered rather than accepted. ### Near-common roots: the over-parameterisation signal A distinct pathology appears when the autoregressive and moving-average polynomials share, or nearly share, a root. Because the model is `phi(B) z_t = theta(B) e_t`, a common factor cancels from both sides, and what looked like an ARMA(1,1) is really white noise, or an ARMA(2,2) is really an ARMA(1,1). Estimation on real data will not cancel it exactly; instead you see the symptoms: - coefficients with standard errors as large as, or larger than, the estimates themselves; - estimates that swing wildly when a few observations are added or removed; - an optimiser that converges slowly, to different points from different starting values, because the likelihood is nearly flat along the direction where the factors trade off; - roots of the two polynomials that, when you compute them, sit almost on top of each other. The fix is to reduce the order. The data is telling you the extra pair of terms is buying nothing, and a leaner model will produce steadier coefficients and better-behaved forecasts. ### Why this matters in practice The conditions are not academic hygiene. Stationarity of the autoregressive part is what makes long-horizon forecasts and their intervals meaningful; invertibility is what makes the forecast recursion converge from observed data; and the absence of near-common roots is what makes the estimated parameters mean anything at all. Checking the fitted roots takes a moment and catches a whole family of silent failures that a fit statistic will happily hide.
- What does an estimated autoregressive root sitting essentially on the unit circle tell you?That the series, at the differencing level you chose, still behaves like a random walk. The autoregressive term is being pushed to the boundary trying to represent a permanent-shock level with a bounded coefficient. The specification needs reconsidering — the model as fitted has no mean to revert to and its long-horizon intervals are not credible.
- Why does invertibility matter if two MA coefficients give the same autocorrelations?That is exactly why it matters. Since theta and 1/theta are indistinguishable from the autocorrelation structure, the parameter is unidentified without a rule; invertibility picks the one below one in modulus. It is also the one with a convergent infinite-autoregression form, which is what lets forecasts be computed from observed values.
- How do you recognise near-cancelling AR and MA factors in a fitted model?Compute the roots of both polynomials and look for a near-match. The accompanying symptoms are standard errors as large as the estimates, coefficients that move sharply with small data changes, and an optimiser that lands in different places from different starting values. The remedy is a lower order, not a better optimiser.
saying these in an interview costs you the question
- Mixes the outside-the-circle and inside-the-circle conventions
- Calls an autoregressive coefficient above one merely strongly trending
- Accepts a boundary moving-average estimate without comment
- Keeps a high order despite exploded standard errors
- Thinks invertibility affects estimation but never forecasts