skip to content

What do the dashed confidence bands on an ACF correlogram represent?

level: juniorimportance: should knowfreq 52%

answer

  1. they encode a null, not a truth
  2. the null is white noise
  3. width shrinks as the sample grows
  4. 1.96 times a standard error
  5. one bar in twenty crosses by chance

basics

~20 s

They mark how large a sample autocorrelation can get by chance if the series were white noise. The band is plus or minus 1.96 over the square root of n, so a bar inside it is indistinguishable from noise.

solid answer

~50 s

The bands are a 95% acceptance region under the null hypothesis that the series is white noise — no autocorrelation at any lag. Under that null the standard error of a sample autocorrelation is roughly `1/sqrt(n)`, so the band sits at `+/- 1.96/sqrt(n)`. With `n = 100` observations that is about `+/- 0.196`. A bar poking outside the band is evidence of real structure at that lag; a bar inside it is consistent with noise. The critical caveat is multiplicity: the band is a per-lag 95% test, so on a genuinely random series roughly 1 bar in 20 will cross it. Seeing one lonely significant spike at lag 17 among 40 plotted lags is exactly what chance produces, and treating it as a finding is the standard beginner error. What counts as evidence is a spike at an interpretable lag, or a run of spikes forming a pattern.

go deeper

for a junior

Know that the dashed lines mark the range a sample autocorrelation can reach by chance under white noise, and that bars inside them are not evidence of structure. Recalling the plus-or-minus 1.96 over root-n form is a strong answer here.

for a middle

Explain where the formula comes from: the standard error of a sample autocorrelation under the white-noise null is about one over root n, and 1.96 is the normal critical value. Be ready for the one-in-twenty multiplicity point.

for a senior

Show that you do not read single bars. Talk about crossings at interpretable lags versus arbitrary ones, about pooling many lags into one portmanteau test, and about the gap between a bar clearing the band and an autocorrelation large enough to change a forecast.

for a principal

Own how correlogram evidence should feed decisions across a team. Set the norm that a lag effect is acted on when it is interpretable and reproduces out of sample, not when a bar clears a dashed line on one plot.

## What the bands are testing A correlogram plots the sample autocorrelation `r_k` against lag k. Every one of those values is an *estimate* computed from a finite sample, so every one of them carries sampling error. Even if the underlying process is pure noise with no memory whatsoever, the estimates will not come out at exactly zero — they will scatter around zero. The bands quantify that scatter. The null hypothesis behind them is **white noise**: the observations are uncorrelated at every lag, with constant variance. Under that hypothesis, for reasonably large n, each sample autocorrelation is approximately normally distributed with mean 0 and standard error `1/sqrt(n)`. Multiplying by the 1.96 critical value from the standard normal distribution gives the familiar 95% band: ``` band = +/- 1.96 / sqrt(n) ``` Some n values worked through: | n | band | |---|---| | 50 | about +/- 0.277 | | 100 | about +/- 0.196 | | 400 | about +/- 0.098 | | 1000 | about +/- 0.062 | Two consequences fall straight out of the formula. First, the band **narrows with the square root of the sample size** — quadrupling the data halves the band. Second, on a short series the band is wide, so genuinely real but modest autocorrelation may sit inside it and go undetected. Absence of a crossing on 40 observations is weak evidence of absence. ## The multiplicity trap Each band is a separate 5%-level test at a single lag. A correlogram typically shows 20, 30 or 40 lags. If the series really is white noise, the expected number of bars crossing the band is 5% of however many you plot — about one bar in twenty, about two in forty. Those crossings are guaranteed by the arithmetic, not by the data. So the sensible reading rule is not "any crossing means autocorrelation". It is: - **A spike at an interpretable lag counts.** A crossing at lag 1, or at the seasonal period, is a specific claim you can check against what you know about how the data were generated. - **A pattern counts.** Several consecutive bars outside the band, or a decaying series of them, is not something noise produces by accident. - **One isolated crossing at an arbitrary lag, with everything else inside, is noise** until something else corroborates it. - **The magnitude of the excess matters.** A bar just kissing the band and a bar at three times the band height are very different pieces of evidence. If you want a single verdict over many lags at once rather than eyeballing a plot, a portmanteau test that pools the first h autocorrelations into one statistic is the right instrument; that is what the Ljung-Box test does. ## A refinement on the band width The flat `+/- 1.96/sqrt(n)` band assumes white noise at *every* lag, which is the right assumption when the correlogram is being used to test that null. Some plotting conventions instead widen the band as the lag increases, using Bartlett's formula: when testing whether the process cuts off after lag q, the standard error of `r_k` for `k > q` accounts for the nonzero autocorrelations at lags 1 through q, and is larger than `1/sqrt(n)`. That produces a funnel-shaped band rather than two flat lines. Both conventions exist; know which one you are looking at, because a bar can be significant against the flat band and not against the widening one. The flat band is the more common default and the one interview answers usually mean. ## Assumptions worth stating The approximation relies on n being reasonably large — the normal approximation to the distribution of `r_k` is poor on very short series — and on the estimate being computed from a stationary series. On a trending series the bands are not meaningful in the intended sense, because the null being tested is already obviously false: the plot will show large positive bars decaying slowly across most of the lag range, and the correct conclusion is "there is a trend here", not "there is structure at lag 13". Finally, the bands say nothing about *practical* importance. On a very long series a lag-1 autocorrelation of 0.04 will comfortably clear the band and still be far too small to change any forecast worth making. Statistical significance on a correlogram is a statement about whether the value is distinguishable from zero, not about whether it is big enough to act on.

  • You plot 40 lags and exactly two bars cross the band. How do you read that?
    As nothing, on its own. Each band is a 5% test at a single lag, so about two crossings out of forty are what white noise produces by chance. I would look at *where* they fall: crossings at lag 1 or at the seasonal period are interpretable and worth pursuing; two isolated bars at arbitrary lags with everything else inside are noise until something corroborates them.
  • Why does the band get narrower as the series gets longer?
    Because the standard error of a sample autocorrelation under the white-noise null is about `1/sqrt(n)`, and the band is 1.96 of those. More observations means each autocorrelation is estimated more precisely, so a smaller deviation from zero is already distinguishable from chance. Quadrupling the sample size halves the band width.
  • A bar clears the band on a series of 5000 points but the value is only 0.05. Does that matter?
    It is statistically distinguishable from zero and practically negligible. With n = 5000 the band sits near 0.028, so 0.05 clears it easily, yet an autocorrelation that small explains a fraction of a percent of the variance and will not move a forecast. Significance on a correlogram is about detectability, not about effect size.

saying these in an interview costs you the question

  • Treating any single crossing as proof of autocorrelation
  • Saying the bands are a 95% interval for the true autocorrelation
  • Forgetting the bands assume a white-noise null
  • Ignoring that band width depends on the sample size
  • Reading bands off a strongly trending, non-stationary series
  • Confusing statistical significance with a forecast-relevant effect

context