skip to content

ADF fails to reject a unit root and KPSS rejects stationarity on the same series — what do you conclude?

level: seniorimportance: should knowfreq 45%

answer

  1. the two nulls point opposite ways
  2. each covers the other's weak side
  3. agreement here, not contradiction
  4. difference once, then test again
  5. both rejecting means neither model fits

basics

~10 s

The two tests carry opposite null hypotheses, so both verdicts point the same way: the evidence supports a unit root. Difference the series once, then re-run both tests on the differenced series before modelling.

solid answer

~50 s

This pair is agreement, not conflict. The augmented Dickey-Fuller test puts the unit root in the null, so failing to reject leaves non-stationarity standing. The KPSS test reverses the roles — its null is stationarity around a level or a deterministic trend — so rejecting it is affirmative evidence against stationarity. Two tests built on opposite defaults reaching the same conclusion is much stronger than either alone, because the weak side of one is the strong side of the other. The action is to difference once and re-run both on the result: you want ADF to reject and KPSS not to reject on the differenced series before you stop. If instead both tests fail to reject, the sample is simply uninformative; if both reject, neither pure model fits and you should stabilise the variance or reconsider the deterministic terms before trusting either verdict.

go deeper

for a junior

Be ready to state that the two tests have opposite nulls: one assumes a unit root, the other assumes stationarity. Getting that direction right is most of the value at this level.

for a middle

Expect to walk through all four outcome combinations and say what each licenses, including that both failing to reject means the sample could not decide.

for a senior

Show the working habit: same deterministic specification for both tests, difference once, re-test the difference, and stop on evidence rather than on whichever result you wanted.

for a principal

Own the standard the team applies when the tests are inconclusive — which default transformation is used, how it is documented, and how the resulting forecast-interval risk is communicated to the people acting on the numbers.

## Why two tests Every hypothesis test has a strong side and a weak side. A rejection is affirmative evidence; a non-rejection only says the data did not overturn the default. The augmented Dickey-Fuller and KPSS tests are used together precisely because their defaults are opposite, so each one's weak side is covered by the other's strong side. - **Augmented Dickey-Fuller.** Null: the series has a unit root. Rejecting supports stationarity. - **KPSS.** Null: the series is stationary around a level or around a deterministic trend. Rejecting supports a unit root. KPSS works by fitting the deterministic part — a constant, optionally plus a trend — and forming the cumulative sums of the residuals. Under stationarity those partial sums stay contained; under a unit root they wander widely, so the statistic is large. It is an upper-tail test: a large statistic rejects stationarity. ## Reading the pair With opposite nulls there are four possible outcomes, and each has a distinct meaning. | ADF | KPSS | Reading | | --- | --- | --- | | rejects | does not reject | Both point to stationarity. Model the level as is. | | does not reject | rejects | Both point to a unit root. Difference and re-test. | | does not reject | does not reject | The data are uninformative. Neither test had the power to decide. | | rejects | rejects | Neither simple model fits. Something else is going on. | The case in the question is the second row, and it is the *clean* one. The ADF result on its own would be weak — a non-rejection can arise simply from a short sample. But KPSS rejecting is a positive finding from the opposite direction. When both defaults, chosen to disagree, land on the same answer, the conclusion is about as firm as this family of tests gets. ## What to do next Difference the series once and re-run both tests on the differenced series. The target is the first row of the table: ADF rejecting and KPSS not rejecting on the difference. That confirms one difference was enough — that the series was integrated of order one — and it is the check most people skip. Stopping after the first pair of tests, differencing, and modelling without re-testing is how both under- and over-differenced series reach production. One discipline matters here: decide in advance how many differences you are willing to take, and on what evidence. Differencing until some test finally cooperates is a multiple-testing exercise dressed up as diagnostics, and it reliably over-differences. ## The other two outcomes **Neither rejects.** This is the honest "we do not know" cell. ADF has weak power against a coefficient near one, KPSS is not sensitive on short samples, and together they simply cannot separate the hypotheses from the data available. Do not read it as a green light for stationarity. Lean on the plot, on how long the series is, and on domain reasoning about whether shocks to this quantity should be permanent. Where the two treatments would produce materially different forecasts, prefer the one whose failure mode you can live with. **Both reject.** Now the tests contradict, which means neither of the two simple models — stationary around a deterministic path, or a pure unit root — describes the series. Common causes worth checking before believing either verdict: a spread that grows with the level, which no amount of differencing fixes and which a variance-stabilising transform applied first would; misspecified deterministic terms, such as running one test with a trend and the other without, which alone can generate the contradiction; or genuinely long-memory behaviour that sits between the stationary and unit-root cases. The first thing to verify is that both tests were run with the *same* deterministic specification, because mismatched specifications are the most common cause of an apparent contradiction. ## Practical reporting A usable stationarity note states, for each test: the deterministic terms included, the lag or bandwidth choice, the statistic, the critical value, and the direction of the conclusion. Then it states the pair reading and the resulting transformation. A single sentence like "ADF p = 0.61, so the series is non-stationary" is a red flag on both counts — it treats a non-rejection as proof and it hides the specification that produced it.

  • What do you conclude when both tests fail to reject their nulls?
    That the data are uninformative rather than that the series is stationary. The Dickey-Fuller test has weak power against a coefficient close to one, and KPSS is not sensitive on short samples, so on a few dozen observations both can come back quiet regardless of the truth. Fall back on the plot, on whether shocks to this quantity should plausibly be permanent, and on which error you can better afford.
  • What does it mean when both tests reject?
    That neither simple model fits. Before believing either verdict, check that both tests used the same deterministic specification — a trend in one and not the other alone produces this pattern. Then look for a spread that grows with the level, which differencing will never fix and which a variance-stabilising transform applied first would, or for long-memory persistence that sits between the stationary and unit-root cases.
  • After one difference, KPSS still rejects. Do you difference again?
    Not reflexively. First check whether the remaining structure is periodic, in which case a difference at the seasonal lag rather than another lag-one difference is the right tool. Then check whether the spread changes across the sample, which calls for a variance-stabilising transform rather than more differencing. A second lag-one difference is occasionally justified for a series whose growth rate itself wanders, but it is the last option, not the next one.

Two reviewers with opposite priors — one assuming a paper should be accepted, one assuming it should be rejected — reaching the same verdict is far more convincing than either verdict alone.

saying these in an interview costs you the question

  • Assumes both tests share the same null hypothesis
  • Reads two agreeing verdicts as a contradiction
  • Treats a KPSS non-rejection as proof of stationarity
  • Differences repeatedly until some test finally cooperates
  • Never re-runs the tests on the differenced series
  • Compares tests run with different deterministic terms

context