skip to content

Why does first differencing turn a random walk into a stationary series?

level: middleimportance: must knowfreq 68%

answer

  1. the level is a running total
  2. nothing is ever forgotten
  3. variance proportional to elapsed time
  4. coefficient exactly one on the lag
  5. the difference is the shock itself

basics

~20 s

A random walk accumulates every past shock, so its variance grows with time and its covariance depends on the date rather than the lag. Differencing once strips the accumulation away and leaves the shock itself, which is stationary noise.

solid answer

~50 s

A random walk is `X_t = X_{t-1} + e_t`, so starting from a fixed `X_0` the value is the running sum of all shocks so far. That gives `Var(X_t) = t * sigma^2`, a variance that grows without bound, and a covariance between two points that depends on how far along the path they are rather than on the gap between them — both conditions of weak stationarity fail. In autoregressive terms the coefficient on the lagged value equals one, a **unit root**: the autoregressive polynomial `1 - phi*B` has a root at `B = 1`, on the unit circle. The first difference is `X_t - X_{t-1} = e_t`, exactly the white noise, with constant mean and constant variance. This is why a daily stock closing-price series gets modelled through its log returns: the return is the first difference of the log price, and it is the stationary object.

code

python · 22 lines
python
import random, statistics

random.seed(7)

def walk(n):
    x, path = 0.0, []
    for _ in range(n):
        x += random.gauss(0, 1)
        path.append(x)
    return path

paths = [walk(400) for _ in range(2000)]

# variance of the LEVEL across replicates, at two time points
for t in (99, 399):
    print("t =", t + 1, "var of level:",
          round(statistics.pvariance([p[t] for p in paths]), 1))

# variance of the FIRST DIFFERENCE along one path
one = paths[0]
steps = [one[i] - one[i - 1] for i in range(1, len(one))]
print("var of first difference:", round(statistics.pvariance(steps), 2))

go deeper

for a junior

Be ready to write X_t = X_{t-1} + e_t, unroll it into a sum of shocks, and say that the first difference is the shock itself. Knowing that prices are modelled through returns is enough at this level.

for a middle

Expect to derive Var(X_t) = t * sigma^2, connect the coefficient of one to the term unit root, and explain that one difference removes one unit root — the meaning of an I(1) series.

for a senior

Demonstrate that you difference the minimum amount the evidence supports, re-test afterwards, and reach for a variance-stabilising transform rather than another difference when the spread, not the level, is what moves.

for a principal

Own the modelling stance: whether the team forecasts levels through differenced models or forecasts changes directly, and how that choice propagates into forecast interval width at long horizons.

## What a random walk is The random walk is the canonical non-stationary series: `X_t = X_{t-1} + e_t` where `e_t` is white noise with mean 0 and variance `sigma^2`. Unrolling the recursion from a fixed starting value `X_0 = 0` gives `X_t = e_1 + e_2 + ... + e_t` The value at time `t` is the running total of every shock that has ever hit the series. Nothing is ever forgotten. ## Why it is not stationary Because the shocks are uncorrelated, the variance of a sum is the sum of the variances: `Var(X_t) = t * sigma^2` The spread grows linearly with time and is unbounded, so the constant-variance condition fails immediately. The covariance is `Cov(X_t, X_s) = min(t, s) * sigma^2`, which depends on the positions `t` and `s` themselves and not only on the gap `|t - s|`, so the lag-only condition fails too. The mean does stay at zero when there is no drift, which is why the mean alone is a poor stationarity check — a driftless random walk has a constant mean and is still emphatically non-stationary. ## The unit-root view Write the general first-order autoregression `X_t = phi * X_{t-1} + e_t`. The process is stationary when `|phi| < 1`: shocks decay geometrically, because a shock at time `t` contributes `phi^k` to the value `k` steps later, and `phi^k` shrinks to zero. The random walk is the boundary case `phi = 1`. In the polynomial form `(1 - phi*B) X_t = e_t`, where `B` is the backshift operator with `B X_t = X_{t-1}`, the root of `1 - phi*B = 0` is `B = 1/phi`. At `phi = 1` that root equals 1 and sits exactly on the unit circle — a **unit root**. Beyond one, at `phi = 1.2`, the process is explosive rather than merely non-stationary. The practical consequence is about memory. Under `|phi| < 1`, a shock is transient and the series returns to its mean. Under a unit root, a shock is **permanent**: it shifts the whole future path by the size of the shock and is never worked off. This is the single most useful intuition for telling the two regimes apart from a plot. ## What differencing does Define the difference operator `dX_t = X_t - X_{t-1}`. Applying it to the random walk gives `dX_t = (X_{t-1} + e_t) - X_{t-1} = e_t` which is exactly the white noise: mean 0, variance `sigma^2` at every time point, and zero covariance at every non-zero lag. One difference removes one unit root, which is what the label **I(1)**, integrated of order one, means: a series that becomes stationary after exactly one difference. Differencing works here because it undoes the accumulation — the level accumulates shocks, and the difference of an accumulation is the increment. With drift the story is the same. For `X_t = mu + X_{t-1} + e_t` the mean grows linearly as `t * mu`, and the difference is `mu + e_t` — stationary with a non-zero mean. One difference still suffices; the drift becomes the mean of the differenced series. ## What differencing does not do Differencing addresses a moving level, not a moving spread. If the size of the fluctuations grows with the level of the series, the differenced series will still show changing variance, and the cure for that is a variance-stabilising transform such as a log or a Box-Cox transform applied *before* differencing, not another difference. Differencing is also not free. Each difference costs an observation, and applying more differences than the series needs inflates the variance of the result and pushes the model toward a non-invertible moving-average structure that is awkward to estimate. Two differences are occasionally justified for a series whose growth rate itself wanders; three essentially never are. ## The financial illustration Daily closing prices of a heavily traded stock are the textbook empirical random walk: the best forecast of tomorrow's price is today's price, past moves carry almost no information about the direction of the next move, and the price level has no fixed centre to revert to. That is why nobody models the price level directly. The log return `r_t = log(P_t) - log(P_{t-1})` is the first difference of the log price, and it hovers around a near-zero mean with a stable unconditional spread. The pattern generalises: when a series is I(1), model the change and reconstruct the level by cumulating the forecasts.

  • How do shocks behave differently in a unit-root series and in a stationary autoregression?
    In a stationary first-order autoregression with coefficient `phi` below one in absolute value, a shock contributes `phi^k` after `k` steps and decays geometrically, so the series reverts to its mean. Under a unit root the coefficient is one, the contribution stays at full size forever, and the level is permanently displaced. Permanent versus fading shocks is the operational difference and it drives whether you difference or detrend.
  • Does a random walk with drift still need only one difference?
    Yes. For `X_t = mu + X_{t-1} + e_t` the mean grows linearly with time, but the first difference is `mu + e_t`, which is stationary with mean `mu` and variance equal to the shock variance. The drift survives as the mean of the differenced series, which is exactly what you want — it becomes the average per-period change the model then forecasts.
  • If one difference helps, why not difference twice to be safe?
    Because over-differencing is a real cost, not a free precaution. Differencing a series more than it needs inflates the variance of the result and induces a moving-average unit root in the errors, which is non-invertible and makes parameter estimation ill-behaved. You also lose another observation. Difference the minimum number of times the evidence supports, and re-check after each one.

The level is a bank balance and the shocks are daily deposits and withdrawals. The balance drifts anywhere over years; the daily transaction amounts keep the same distribution the whole time.

saying these in an interview costs you the question

  • Says a random walk is stationary because its expected change is zero
  • Claims differencing also fixes a variance that grows with the level
  • Treats a unit root as any coefficient above zero rather than exactly one
  • Believes shocks to a unit-root series eventually wash out
  • Differences repeatedly on the assumption that more is safer

context