Your forecast model's residuals fail a Ljung-Box test at lag 10 — what does that mean?
answer
- it pools many lags into one verdict
- the null is the good outcome
- small p-value means structure remains
- chi-squared, with parameters subtracted
- residual ACF tells you which lag
basics
~20 sIt means the residuals still carry autocorrelation somewhere in the first ten lags, so the model has left predictable structure behind. The Ljung-Box null is that all those residual autocorrelations are zero; a small p-value rejects it.
solid answer
~50 sLjung-Box is a portmanteau test: it pools the first h residual autocorrelations into one statistic, `Q = n(n+2) * sum_{k=1..h} r_k^2 / (n - k)`, compared against a chi-squared distribution. The null is that the residuals are uncorrelated up to lag h — that the model extracted all the linear lag structure and left white noise behind. A small p-value rejects that, so failing at lag 10 says predictable signal remains and the forecasts are leaving accuracy on the table. My next step is diagnostic: plot the residual ACF and see *where* the correlation sits. A lag-1 spike points at short-memory structure; a spike at the seasonal lag says the seasonal handling is wrong. One caveat — on fitted residuals the degrees of freedom must be reduced by the number of estimated lag parameters, or the p-value is optimistic.
go deeper
Know that this is a test on model residuals, that the null is no autocorrelation up to the chosen lag, and that a small p-value means the model left structure behind. Getting the null direction right is the whole ask here.
Explain the mechanics: pooling several lags into one chi-squared statistic, why individual correlogram bars invite a multiplicity problem, and why degrees of freedom shrink when the test runs on fitted rather than raw residuals.
Show what you do after a rejection — locate the offending lag on the residual ACF and map lag 1, seasonal lags and a broad positive band to different causes. Name what the test cannot see: changing variance, fat tails, bias, nonlinear dependence.
Own residual checking as a standard rather than a one-off. Decide what evidence gates a model into production, resist teams tuning the lag window until a test passes, and weigh a statistically detectable residual correlation against whether fixing it changes any decision.
## What the test does When a forecasting model has done its job, the residuals — actual minus fitted, at each time point — should look like white noise: no remaining correlation with their own past, because any such correlation is predictable signal the model failed to use. Eyeballing the residual correlogram lag by lag works, but it invites the multiplicity problem: with twenty lags plotted, one or two chance crossings of the significance bands are expected under the null, so a per-lag reading gives no clean verdict. A **portmanteau test** solves this by pooling. The Ljung-Box statistic is ``` Q = n(n + 2) * sum over k = 1..h of [ r_k^2 / (n - k) ] ``` where `n` is the number of observations, `r_k` is the sample autocorrelation of the residuals at lag k, and `h` is the number of lags tested. Each squared autocorrelation is weighted by `1/(n-k)`, which upweights longer lags to compensate for their being estimated from fewer overlapping pairs — that small-sample correction is what distinguishes Ljung-Box from the earlier Box-Pierce statistic, which simply summed `n * r_k^2`. Under the null, `Q` follows approximately a chi-squared distribution. Since `Q` grows as the residual autocorrelations get larger in either direction, the test is one-sided in the statistic: **large Q, small p-value, reject**. ## The null, stated precisely The null hypothesis is: *the autocorrelations of the residual series at lags 1 through h are all zero.* It is a joint statement about the whole block of lags, not about any single one. So: - **Small p-value (reject).** At least one of those lags carries real autocorrelation. The model has missed structure. The test does not say which lag — that is what the residual ACF plot is for. - **Large p-value (fail to reject).** No detectable linear autocorrelation in the first h lags. This is *not* proof the model is correct. It is the absence of one specific kind of evidence against it, and on a short series the test has little power to find modest autocorrelation at all. The standard misstatement to avoid: a large p-value does not "accept" the null or prove the residuals are white noise. It says the data give no strong reason to reject that description. ## Degrees of freedom This is where candidates most often slip. Applied to a *raw* series, `Q` has h degrees of freedom. Applied to the **residuals of a fitted model**, the residual autocorrelations have already been shrunk by the fitting procedure — the estimation deliberately drove some of them toward zero — so `Q` is smaller than it would be for a genuinely independent series. The correction is to subtract the number of estimated lag parameters from the degrees of freedom. Using the uncorrected h on fitted residuals produces p-values that are too large, biasing you toward concluding the model is fine when it is not. ## Choosing h There is no single right answer, but two conventions dominate. For non-seasonal data, a moderate h in the region of 10 is common. For seasonal data with period m, testing to about `2m` is the usual choice, so that at least the first two seasonal lags fall inside the window — a model that mishandles seasonality will show its failure at lag m, and a test truncated at lag 5 on monthly data would never see it. The tradeoff: too small an h can miss structure that lives at longer lags; too large an h dilutes real autocorrelation at a few lags across many null ones, costing power. Do not tune h until the test passes — that is a form of selective reporting. ## What the test does not cover Ljung-Box tests **linear autocorrelation of the residual level**, and nothing else. Residuals can pass it comfortably and still be badly behaved: - **Changing variance.** Calm periods followed by volatile ones leave the residual level uncorrelated while the residual *magnitudes* cluster. A common probe is to run the same test on squared residuals, where volatility clustering shows up as autocorrelation. - **Non-normality.** Fat tails or skew affect interval forecasts and are invisible to this test. - **Nonlinear dependence.** A relationship the residuals carry in a non-linear form contributes little linear autocorrelation. - **Bias.** A residual mean far from zero is a systematic offset the test says nothing about. ## Durbin-Watson, the narrower relative Durbin-Watson is the other statistic you will be asked about here. It is computed on residuals as ``` d = sum (e_t - e_{t-1})^2 / sum e_t^2 ``` and ranges from 0 to 4, with the approximation `d ≈ 2(1 - r_1)`. A value **near 2 is the no-serial-correlation reading**; values near 0 indicate strong *positive* first-order autocorrelation, and values near 4 strong *negative* first-order autocorrelation. Its limitation is right there in the algebra: it looks only at **lag 1**. Residuals with a clean seasonal spike at lag 12 and nothing at lag 1 will produce a Durbin-Watson comfortably near 2 while being obviously autocorrelated. It is also not valid in the usual form when a lagged value of the dependent variable appears among the regressors, a situation that arises constantly in time-series work. That is precisely why a portmanteau test over a block of lags is the better default for residual checking. ## What to do after a rejection Rejection is a prompt, not a prescription. Plot the residual ACF and PACF and locate the offending lags. A lag-1 spike says short-memory structure remains. A spike at the seasonal period says the seasonal component is mishandled. Large positive autocorrelation across many short lags usually means the series was never adequately leveled and the model is fitting a moving mean. A single marginal crossing in a long lag window, with a p-value hovering near the threshold, may be the multiplicity problem reasserting itself and deserves less alarm than a clear, interpretable spike.
- How do you choose the number of lags h for a Ljung-Box test?For non-seasonal data a moderate window around 10 lags is conventional. For seasonal data with period m, testing to about 2m is standard so the first two seasonal lags fall inside the window — otherwise a model that mishandles the season can pass. Too few lags misses long-range structure; too many dilute real autocorrelation across null lags and cost power. Never tune h until the test passes.
- How does Durbin-Watson relate to this, and when is it not enough?Durbin-Watson is a first-order-only statistic, roughly `2(1 - r_1)`, ranging from 0 to 4 with 2 meaning no lag-1 serial correlation, near 0 strong positive and near 4 strong negative. It is blind past lag 1, so residuals with a clean seasonal spike at lag 12 pass it while being plainly autocorrelated. A portmanteau test over a block of lags is the better default.
- The residuals pass Ljung-Box. Is the model adequate?Not established — only one failure mode has been ruled out. Failing to reject is not proof of white noise, and on short series the test has weak power. The residuals can still show changing variance, fat tails, systematic bias or nonlinear dependence, none of which this test sees. Running the same test on squared residuals is a cheap check for volatility clustering.
- Why does Ljung-Box weight each lag by 1/(n-k)?Because the autocorrelation at lag k is computed from only `n - k` overlapping pairs, so longer lags are estimated less precisely. The weight compensates, giving the statistic a better chi-squared approximation in finite samples. Without it — as in the earlier Box-Pierce form, which simply sums `n * r_k^2` — the test is noticeably less reliable on short series.
The correlogram is like inspecting twenty windows one at a time and arguing about each; the portmanteau test is one alarm wired to all twenty at once. It tells you somebody got in, not which window they used.
saying these in an interview costs you the question
- Saying the null hypothesis is that autocorrelation exists
- Reading a large p-value as proof the residuals are white noise
- Forgetting to subtract estimated parameters from the degrees of freedom
- Trusting Durbin-Watson near 2 when seasonal lags are untested
- Increasing or shrinking the lag window until the test passes
- Assuming passing means the residuals have constant variance