skip to content

Why can regressing one random walk on an unrelated random walk give a high R-squared?

level: middleimportance: must knowfreq 58%

answer

  1. neither series returns to a mean
  2. the residuals are the problem
  3. standard errors badly understated
  4. compare R-squared against Durbin-Watson
  5. difference first, or test cointegration

basics

~20 s

A random walk never returns to a mean, so over any window it looks trended and least squares lines up the two drifts. The residuals stay non-stationary, so standard errors are far too small and the fit statistics are meaningless.

solid answer

~50 s

A random walk has no mean to return to, so over any finite stretch it looks like it is trending. Regress two independent walks on each other and least squares happily lines up one wandering path against the other: R-squared is often large and the slope's t-statistic frequently far exceeds 2. The inference is invalid because ordinary standard errors assume roughly independent residuals, while here the residuals are themselves a wandering, non-stationary series, so the standard error is badly understated. Worse, the problem grows with the sample: the t-statistic diverges rather than settling down, so more data makes the false finding look stronger. The classic tells are a Durbin-Watson statistic close to 0 and an R-squared larger than the Durbin-Watson value. The fix is to regress the first differences, or — if the two series really do share a long-run equilibrium — to test for cointegration and fit an error-correction model instead.

code

python · 21 lines
python
import random

def walk(n):
    x, out = 0.0, []
    for _ in range(n):
        x += random.gauss(0, 1)
        out.append(x)
    return out

def slope_t(x, y):                       # OLS y = a + b*x, t-statistic for b
    n = len(x)
    mx, my = sum(x) / n, sum(y) / n
    sxx = sum((xi - mx) ** 2 for xi in x)
    b = sum((xi - mx) * (yi - my) for xi, yi in zip(x, y)) / sxx
    a = my - b * mx
    s2 = sum((yi - a - b * xi) ** 2 for xi, yi in zip(x, y)) / (n - 2)
    return b / (s2 / sxx) ** 0.5

random.seed(0)
hits = sum(abs(slope_t(walk(200), walk(200))) > 2 for _ in range(1000))
print(hits / 1000)                       # nowhere near 0.05

go deeper

for a junior

Be ready to say that a regression on two trending series can look excellent while meaning nothing, and that first differences are the standard first move before regressing time-ordered data.

for a middle

Explain the mechanism: no mean reversion means the residual is itself non-stationary, so the standard error is understated and the t-statistic is not comparable to 2. Know the Durbin-Watson tell.

for a senior

Show the diagnostic habit — residual plots, out-of-sample checks, an R-squared-versus-Durbin-Watson sanity check — and explain that lengthening the sample strengthens rather than dissolves the false finding.

for a principal

Own the framing: when a dashboard or a model built on levels reports a stunning relationship between two drifting metrics, decide whether to difference, to model an equilibrium, or to declare the finding unusable, and set that standard for the team.

## The setup A **random walk** is a series in which each value is the previous value plus an independent random shock: `x_t = x_{t-1} + e_t`. Because the shock is added to the *level* and never decays, the series has no mean to return to and its variance grows without bound as time passes. Such a series is called **non-stationary**, or **integrated of order one** (written I(1)), because taking first differences — `x_t - x_{t-1} = e_t` — leaves a stationary, memoryless series. Now generate two random walks completely independently. Nothing connects them: no shared shocks, no mechanism, no common driver. Regress one on the other with ordinary least squares and inspect the slope's t-statistic and the R-squared. ## What actually happens Granger and Newbold ran this experiment in 1974 and found that the null hypothesis of a zero slope was rejected far more often than the nominal 5% error rate — the majority of independent pairs produced a "significant" slope. R-squared values above 0.5 are routine, and t-statistics in double digits are not unusual. Yule had made the same point in 1926 with real data, reporting a correlation of about 0.95 between the standardised mortality rate in England and Wales and the proportion of marriages performed in the Church of England over 1866-1911 — two series that share nothing but a downward drift. He called these **nonsense correlations**. ## Why least squares is fooled Three things go wrong at once. 1. **A random walk looks trended over any finite window.** Slice a wandering path anywhere and it goes broadly up or broadly down over that slice. Two such slices can be lined up by a straight line, and the fitted line explains a lot of the variation in the vertical spread, which is exactly what R-squared measures. 2. **The residuals inherit the non-stationarity.** If neither series mean-reverts and no true linear combination of them mean-reverts either, the regression residual `y_t - a - b*x_t` is itself a wandering I(1) series. Ordinary standard errors are derived assuming residuals are close to independent draws; when consecutive residuals are almost perfectly correlated, the effective number of independent observations is a tiny fraction of the nominal sample size, and the reported standard error is far too small. 3. **More data makes it worse, not better.** With well-behaved stationary data, a false slope estimate shrinks toward zero and its t-statistic settles into a t distribution. Here the t-statistic *diverges* as the sample lengthens — roughly in proportion to the square root of the sample size — and R-squared converges not to zero but to a random quantity that differs from sample to sample. So the reassuring instinct "I have ten years of daily data, the significance must be real" is precisely backwards. ## How to spot it - **Durbin-Watson near 0.** The Durbin-Watson statistic is approximately `2 * (1 - rho)`, where `rho` is the first-order autocorrelation of the residuals. Values near 2 indicate no autocorrelation; values near 0 indicate strong positive autocorrelation, the signature of a non-stationary residual. Granger and Newbold's rule of thumb: if R-squared exceeds the Durbin-Watson statistic, suspect a spurious regression. - **Plot the residuals.** A spurious fit leaves residuals that drift in long, slow excursions above and below zero rather than scattering around it. - **Hold out the tail of the series.** A spurious relationship almost always fails out of sample, because the drift that produced the fit does not have to continue in the same direction. ## What to do instead - **Difference both series and regress the differences.** Both differenced series are stationary, standard inference is restored, and any surviving relationship is a genuine short-run co-movement. - **Test for cointegration first.** Sometimes two I(1) series really do share a long-run equilibrium — a linear combination of them *is* stationary. Differencing blindly would throw that away; the right model keeps the level gap as an error-correction term. - **Specify dynamics properly.** Including lags of both the dependent and the explanatory series often absorbs the autocorrelation that the naive levels regression dumped into the residual. - **Adding a deterministic time trend does not fix it.** A `t` regressor removes a straight-line trend; a random walk carries a *stochastic* trend whose direction changes at random. The residual stays I(1) and the inference stays invalid. ## Interview framing The question is really testing whether you check the stationarity of your inputs before you trust a regression on time-ordered data, and whether you know that the diagnostic lives in the *residuals*, not in the goodness-of-fit number.

  • How does the Durbin-Watson statistic help you spot a spurious regression?
    Durbin-Watson is roughly `2 * (1 - rho)` for the residuals' first-order autocorrelation `rho`, so a value near 0 says consecutive residuals are almost perfectly correlated — the residual is wandering rather than scattering. Granger and Newbold's rule of thumb is that an R-squared larger than the Durbin-Watson statistic marks the regression as suspect.
  • Does adding a deterministic time trend to the levels regression fix the problem?
    No. A `t` regressor absorbs a straight-line drift, but a random walk carries a stochastic trend whose direction changes at random and is not a function of `t`. The residual remains non-stationary, the standard errors remain understated, and the t-statistic remains invalid. Difference the series or test for cointegration instead.
  • If differencing destroys the long-run relationship you actually care about, what else can you do?
    Test whether the two series are cointegrated — whether some linear combination of their levels is stationary. If it is, fit an error-correction model: the differences carry the short-run dynamics while the lagged level gap carries the long-run equilibrium, so you get valid inference without discarding the level information.

Two people scribbling independent doodles will often trace roughly the same wandering shape over a short stretch of paper. A straight line drawn through the pair fits both nicely, though neither pen was ever watching the other.

saying these in an interview costs you the question

  • Says a large R-squared proves a real relationship exists
  • Trusts the t-statistic without ever looking at the residuals
  • Claims a longer sample would settle the question
  • Adds a time trend and declares the series stationary
  • Confuses a stochastic trend with a deterministic one

context