skip to content

In the augmented Dickey-Fuller test, what does failing to reject the null mean?

level: middleimportance: should knowfreq 60%

answer

  1. which side is non-stationarity on
  2. the default is a unit root
  3. absence of evidence, not evidence of absence
  4. regress the change on the previous level
  5. critical values sit left of the t-table

basics

~20 s

The augmented Dickey-Fuller null is that the series has a unit root, so failing to reject only means a unit root could not be ruled out. That is weak evidence of non-stationarity, never proof of it.

solid answer

~50 s

The augmented Dickey-Fuller test puts non-stationarity in the null: `H0` is that the series contains a unit root, and the alternative is stationarity, around a constant or a deterministic trend depending on which terms you include. It is run as a regression of the change on the previous level, `dX_t = alpha + beta*t + gamma*X_{t-1} + (lagged changes) + e_t`, testing `gamma = 0`; the lagged changes are the *augmentation*, present to soak up serial correlation. The test is left-tailed and its statistic is compared against Dickey-Fuller critical values, not an ordinary t-table, because under a unit root the estimator has a non-standard limiting distribution. Failing to reject is therefore the usual asymmetry of testing: absence of evidence. Power is weak against a coefficient near one, so on a short series a non-rejection says very little.

go deeper

for a junior

Be ready to say which hypothesis is the null — the unit root — and to phrase the outcome in the right direction rather than claiming the test proved non-stationarity.

for a middle

Expect to explain the underlying regression of the change on the previous level, what the lagged difference terms are there for, and why the critical values are not ordinary t values.

for a senior

Show judgment about power and specification: state which deterministic terms you ran with, treat a non-rejection on a short series as weak, and pair the test with other evidence rather than deciding on a p-value alone.

for a principal

Own the standard the team applies — what evidence licenses differencing a production series, and how that decision gets documented so different analysts do not each pick a different transformation of the same data.

## What the test is actually testing The augmented Dickey-Fuller test asks whether a series contains a unit root. The crucial design choice is which side of the hypothesis pair non-stationarity sits on: - **Null hypothesis:** the series has a unit root, and is therefore non-stationary. - **Alternative hypothesis:** the series is stationary — around a constant, or around a deterministic trend, depending on the specification. That orientation drives every interpretation. Rejecting is affirmative evidence for stationarity. Failing to reject is not affirmative evidence for a unit root; it is the ordinary asymmetry of significance testing, where a non-rejection means the data did not supply enough evidence to overturn the assumed default. ## The regression behind it Start from `X_t = phi * X_{t-1} + e_t` and subtract `X_{t-1}` from both sides: `dX_t = (phi - 1) * X_{t-1} + e_t = gamma * X_{t-1} + e_t` A unit root, `phi = 1`, is exactly `gamma = 0`. So the test regresses the period-to-period change on the previous level and asks whether that coefficient is zero. The full specification adds deterministic terms and lagged changes: `dX_t = alpha + beta*t + gamma*X_{t-1} + d_1*dX_{t-1} + ... + d_p*dX_{t-p} + e_t` The intuition for `gamma` is mean reversion. If `gamma` is meaningfully negative, a high level today pulls the next change downward — the series is pulled back toward a centre. If `gamma` is zero, the level carries no information about the next change, which is the random-walk case. ## What the augmentation adds The original Dickey-Fuller test assumes the errors are white noise. Real series are rarely that clean: even after accounting for the level, the changes are usually autocorrelated. The **augmented** version adds `p` lagged difference terms whose only job is to absorb that serial correlation so the residuals behave and the statistic follows its intended distribution. Those lag coefficients are nuisance parameters — you never interpret them. Choosing `p` matters: too few leaves autocorrelation in the errors and distorts the test size; too many wastes degrees of freedom and lowers power. ## Why the critical values are special Under a unit root the regressor `X_{t-1}` is itself non-stationary, and the usual asymptotics do not apply. The estimator of `gamma` has a non-standard limiting distribution — the Dickey-Fuller distribution — that is shifted left relative to a normal. Its critical values are more negative than ordinary t critical values, and they differ depending on whether you included a constant, a constant and a trend, or neither. Comparing an ADF statistic against a standard t-table is a classic error and rejects far too often. The test is one-sided in the negative direction, because only a negative `gamma` corresponds to mean reversion. ## Choosing the deterministic terms The specification is a modelling decision, not a default. Including no constant assumes the series has mean zero. Including a constant allows a non-zero level. Including a constant and a trend makes the alternative *trend stationarity* — stationary fluctuation around a sloping line. Omitting a trend when the data plainly have one biases the test toward not rejecting, because the unmodelled slope looks like persistence. Adding an unnecessary trend costs power. Always state which version was run alongside the result. ## Where power fails The honest weakness of this test is power. A stationary series with a coefficient of 0.95 is hard to distinguish from one with a coefficient of exactly 1.00 unless the sample is long: the two produce very similar-looking paths over a few hundred points. On short samples the test therefore fails to reject nearly everything, and a non-rejection carries very little information. This is not a flaw to work around by re-running the test on subsamples until it rejects; it is a limit that argues for pairing the test with one whose null is the opposite, and for weighing the plot and the domain knowledge alongside both. ## How to report it A usable report says which deterministic terms were included, the lag order used for the augmentation, the statistic, the critical value it was compared against, and the conclusion phrased in the right direction: "we reject the unit root at the 5% level, so the evidence supports stationarity", or "we cannot reject a unit root; the data are consistent with non-stationarity, and given the sample length that is weak evidence". Never write "the test proves the series is non-stationary".

  • What does the augmented part add over the original Dickey-Fuller test?
    Lagged difference terms. The original test assumes the errors are white noise, which real series rarely satisfy; leftover serial correlation distorts the statistic's distribution and the resulting significance level. The augmentation adds lagged changes purely to absorb that correlation so the residuals behave. Their coefficients are nuisance parameters and are never interpreted, but the lag order chosen affects both the size and the power of the test.
  • Why can't you compare the statistic to a standard t-table?
    Because under the null the regressor is a non-stationary level, so the usual asymptotic theory breaks down. The estimator follows the non-standard Dickey-Fuller distribution, which is shifted left of the normal, so the correct critical values are considerably more negative than ordinary t values — and they differ again depending on whether a constant, a trend, or neither was included. Using a t-table rejects the unit root far too often.
  • When does the test have poor power?
    When the true autoregressive coefficient is close to but below one, and when the sample is short. A coefficient of 0.95 generates paths that look much like a genuine random walk over a few hundred observations, so the test fails to reject even though the series is technically stationary. Misspecified deterministic terms — omitting a trend the data clearly have — also push the test toward not rejecting.

saying these in an interview costs you the question

  • Says failing to reject proves the series has a unit root
  • States the null hypothesis is stationarity
  • Compares the statistic against ordinary t critical values
  • Reports a result without saying whether a constant or trend was included
  • Re-runs the test on subsamples until one rejects

context