skip to content

What does it mean to say oil prices Granger-cause airline stock returns?

level: middleimportance: should knowfreq 54%

answer

  1. a claim about forecasting, not mechanism
  2. does adding lags of X help predict Y
  3. joint test on the lagged coefficients
  4. an unobserved common driver mimics it
  5. expectations can reverse the apparent order

basics

~20 s

Past oil prices improve the forecast of airline returns beyond what the returns' own history explains. It is predictive precedence, established by testing whether lagged oil terms jointly add explanatory power, not evidence of a causal mechanism.

solid answer

~50 s

Granger causality is a forecasting claim, not a mechanism claim. You regress airline returns on their own lags, then add lags of oil prices, and test whether the oil coefficients are jointly zero. If they are not, past oil carries information about future airline returns that the returns' own history does not, and we say oil Granger-causes airline returns. The name is unfortunate: the test only establishes that one series systematically moves first in a way that helps prediction. It can be produced by an unobserved common driver — a global demand shock that moves both — and it can be reversed by anticipation, since markets price expected fuel costs before the spot price moves. It also assumes the sampling frequency is fine enough to see the ordering, and both series must be stationary or handled with a method built for unit roots, or the test over-rejects. Bidirectional Granger causality is common and perfectly coherent.

go deeper

for a junior

Be ready to state the definition in forecasting terms — past values of one series help predict the other beyond its own past — and to say plainly that the word 'cause' in the name overclaims.

for a middle

Explain the mechanics: regress the target on its own lags, add lags of the candidate, and jointly test those coefficients. Know that the inputs must be stationary and that feedback in both directions is possible.

for a senior

Demonstrate judgment about when a positive result is untrustworthy — an unobserved common driver, anticipation inverting the order, a sampling interval coarser than the transmission — and how you would probe each.

for a principal

Own the framing for the organisation: decide when predictive precedence is a good enough basis to act on, when a genuine causal question needs a different design entirely, and how to keep the two apart in reporting.

## The definition Clive Granger's idea, from 1969, is deliberately operational: **X Granger-causes Y if past values of X help predict Y better than the past of Y alone.** Nothing more is claimed. Formally, take the model `y_t = a + b_1*y_{t-1} + ... + b_p*y_{t-p} + c_1*x_{t-1} + ... + c_p*x_{t-p} + e_t` and test the joint null `c_1 = c_2 = ... = c_p = 0` with an F-test (or the equivalent Wald test). Rejecting the null means the lags of X add predictive content, and X is said to Granger-cause Y. Applied to oil and airlines: if last week's oil prices sharpen the forecast of this week's airline returns beyond what airline returns' own history gives, oil Granger-causes airline returns. ## Why the name misleads Granger himself preferred to talk about *temporal precedence in predictability*, and the confusion he predicted is exactly what interviewers probe. Three distinct failure modes: **Omitted common driver.** Suppose global demand conditions move both oil and airline profitability, and oil markets happen to reprice a day faster than equities. Oil will Granger-cause airline returns even though the causal arrow runs from demand to both. The test sees only the two series it is given; it cannot condition on anything outside the model. **Anticipation.** Financial series price expectations, not just realised events. If traders expect fuel costs to fall next quarter, airline shares can move *before* oil does, and the test will report airline returns Granger-causing oil prices — the reverse of the physical mechanism. Any forward-looking market can invert the apparent direction. **Sampling frequency.** If the real transmission happens within a day and you have monthly data, both series move within the same observation and neither precedes the other; the test sees nothing, or attributes the ordering to noise. Conversely, aggregation can create apparent precedence where none exists at the underlying frequency. ## Practical requirements - **Stationarity.** The standard F-test assumes stationary inputs. Run it on two drifting price levels and the test statistic no longer has its usual distribution; you get inflated rejection rates for the same reason a levels regression between two wandering series looks significant. The usual practice is to test on returns or differences. - **A unit-root-safe alternative.** If you must work in levels, the Toda-Yamamoto procedure fits a model in levels with extra lags equal to the maximal order of integration and tests only the original lags; that Wald statistic has a standard asymptotic chi-squared distribution regardless of whether the series have unit roots or are cointegrated. - **Lag length matters.** The conclusion can flip with the number of lags included, so choose the lag order by an information criterion and report the sensitivity rather than picking the length that gives the answer you wanted. - **Direction is not exclusive.** Testing both directions can find that each series Granger-causes the other. That is feedback, and it is a perfectly legitimate finding — it simply means each carries predictive information about the other. ## What it is good for Despite the caveats, the test earns its place. It is a disciplined way to ask whether a candidate leading indicator actually carries usable predictive content beyond the target's own history — which is precisely the bar a feature must clear before it belongs in a forecasting model. Framed as "does adding this series improve my forecast?", it is honest and useful. Framed as "does this series cause the other?", it overclaims. ## Interview framing The strongest answer states the definition in forecasting terms in one sentence, then names at least one concrete way the test can point the wrong way — a common driver, or anticipation reversing the order — and finishes with the stationarity requirement. Candidates who simply say "it proves causality" or "it proves nothing at all" both miss: it is a real, testable, useful claim about predictive precedence.

  • What goes wrong if you run the Granger test on two non-stationary price levels?
    The F-statistic loses its usual reference distribution, so you over-reject and find predictive precedence that is not there — the same failure that makes a levels regression between two drifting series look significant. Test on differences or returns instead, or use the Toda-Yamamoto approach, which adds extra lags in levels so that the Wald test recovers a standard chi-squared distribution.
  • Can two series Granger-cause each other, and what would that mean?
    Yes, and it is common. Bidirectional Granger causality means each series' history improves the forecast of the other beyond its own past — a feedback relationship. It is not a contradiction, because the claim is about predictive content in each direction, not about a single causal arrow that must point one way.
  • How would anticipation reverse the direction the test reports?
    Forward-looking markets price expected conditions. If traders expect fuel costs to rise, airline shares can fall before the oil price actually moves, so the test reports airline returns Granger-causing oil. The physical mechanism still runs from fuel cost to airline profit; the observable ordering has simply been inverted by expectations.
  • How sensitive is the result to the number of lags you include?
    Quite sensitive. Too few lags and a genuine relationship with a longer delay is missed; too many and the joint test loses power while soaking up degrees of freedom. Choose the lag order with an information criterion on the unrestricted model, and report whether the conclusion survives nearby lag lengths rather than presenting a single specification.

saying these in an interview costs you the question

  • Reads Granger causality as proof of a causal mechanism
  • Runs the standard test on two drifting price levels
  • Ignores an unobserved driver that moves both series
  • Says bidirectional Granger causality is a contradiction
  • Reports one lag length without checking sensitivity

context