skip to content

What does a cross-correlation peak at lag -14 between daily ad spend and sign-ups tell you?

level: juniorimportance: should knowfreq 46%

answer

  1. correlation computed at every relative shift
  2. the sign convention decides who leads
  3. shared trend correlates at every lag
  4. each series' own memory smears the peak
  5. prewhiten before trusting the location

basics

~20 s

The two series align best at a two-week offset. Under the usual convention, where the lag shifts ad spend, spend from 14 days earlier tracks today's sign-ups, so spend leads. Confirm the sign convention before claiming a direction.

solid answer

~50 s

The cross-correlation function computes the correlation between the two series at every relative shift; a peak at lag -14 means they align best when one is moved two weeks against the other. Under the common convention where the lag indexes the first named series — `corr(spend_{t+k}, signups_t)` — a peak at `k = -14` means spend two weeks *before* a given day is what tracks that day's sign-ups, so ad spend leads sign-ups by about a fortnight. The first thing I would do is verify the sign convention in whatever computed it, because the two conventions are mirror images and getting it backwards inverts the business conclusion. The second is to be sceptical of the peak itself: if both series trend upward or share weekly seasonality, almost every lag will look correlated and the peak position is unreliable. Remove trend and seasonality — ideally prewhiten — before reading a lead-lag off the plot.

go deeper

for a junior

Be ready to explain that the cross-correlation function is one correlation per relative shift, that an off-zero peak means one series leads the other, and that you must confirm which series the lag indexes.

for a middle

Explain why raw cross-correlations mislead: shared trend correlates the pair at every lag, and each series' own autocorrelation smears the peak and narrows the significance bands. Describe prewhitening as the remedy.

for a senior

Show that you validate a lead before anyone builds on it — check it on held-out weeks, check that the sampling interval can resolve it, and separate 'spend moves first' from 'spend drives sign-ups'.

for a principal

Own the decision of whether a discovered lead is stable enough to plan budget or staffing around, and what monitoring would tell you when the lag has drifted or vanished.

## What the cross-correlation function is Given two time series measured on the same clock, the **cross-correlation function** (CCF) is the correlation between them computed at every relative shift, or **lag**. Instead of one number, you get a curve: one correlation per lag, typically plotted with lag on the horizontal axis and the correlation on the vertical, with bands marking what is distinguishable from zero. A peak somewhere off zero says the two series match best when one is slid along the time axis relative to the other — a **lead-lag** relationship. ## Reading the sign This is the part that trips people up, and interviewers know it. There are two conventions in circulation, and they are exact mirror images: - `CCF(k) = corr(x_{t+k}, y_t)` — a negative peak means the *earlier* values of `x` line up with `y`, so `x` leads. - `CCF(k) = corr(x_t, y_{t+k})` — a *positive* peak means `x` leads. So "a peak at lag -14 between ad spend and sign-ups" is only interpretable once you say which series the lag indexes. Under the first convention, with spend as `x`, the peak at -14 says spend measured 14 days before a given day is what correlates with that day's sign-ups: **ad spend leads sign-ups by two weeks**, which is a plausible story for a consideration-phase product. Under the mirror convention the same number would say sign-ups lead spend, which would instead suggest budget is being allocated in response to demand. The safe habit is to confirm the direction with a sanity check you already believe — shift the series by hand and look at a plot — rather than trusting your memory of the convention. ## Why the peak can be an artefact A raw CCF between two business series is usually not trustworthy on its own, for two reasons. **Shared trend.** If both series grow over the year, they are positively correlated at essentially every lag. The CCF then sags in a broad hump rather than showing a sharp peak, and the exact location of the maximum is close to arbitrary — it can move by weeks between samples. **Autocorrelation within each series.** Ad spend today looks like ad spend yesterday; sign-ups today look like sign-ups yesterday. That internal memory smears any genuine relationship across many neighbouring lags and inflates the correlations, so the CCF shows a wide band of large values instead of a clean spike, and the sampling bands drawn around zero are too narrow to trust. **Shared seasonality.** If both series have a weekly rhythm — spend paused at weekends, sign-ups lower at weekends — the CCF will show peaks at multiples of 7 that reflect the calendar rather than any causal chain. ## Prewhitening The standard remedy is **prewhitening**. Fit a model to the input series (spend) that reduces it to something close to uncorrelated noise, then apply that *same* filter to both series, and compute the CCF on the filtered pair. Because the input's own memory has been stripped out, a peak in the resulting CCF reflects the transfer between the series rather than the internal structure of either one, and the significance bands become meaningful. Removing trend and seasonality first — by differencing, by subtracting a fitted seasonal component, or by including calendar terms — is the minimum version of the same idea. ## Other cautions - **Sampling interval sets the resolution.** Daily data cannot resolve a lead of a few hours, and rolling the data up to weekly buckets can hide a 3-day lead entirely or shift where the peak appears. - **A lead is evidence about ordering, not about mechanism.** That spend moves first is consistent with spend driving sign-ups, but also with both responding to something that moves spend first — a product launch calendar, for example, that raises budgets and then attracts users. - **Alignment is not a forecast.** A lead of 14 days is only exploitable if it is stable; check whether the same peak appears in a held-out stretch of the series before building anything on it. ## Interview framing A good answer names the direction confidently *and* immediately qualifies it: which convention, and what was done about trend, seasonality and autocorrelation before the plot was read.

  • Why can a cross-correlation function look significant at almost every lag?
    Because both series carry trend and internal autocorrelation. A shared upward drift makes them positively correlated at every shift, and each series' own memory spreads any genuine association across neighbouring lags. The result is a broad hump rather than a spike, and the significance bands — which assume uncorrelated inputs — are far too narrow.
  • What is prewhitening, and why do it before reading the peak?
    Prewhitening means fitting a model to the input series so that what remains is close to uncorrelated noise, then applying that same filter to both series and computing the cross-correlation on the filtered pair. It strips out the input's internal memory, so the surviving peak reflects the relationship between the series rather than the structure inside one of them.
  • How does the sampling interval limit what lead-lag you can detect?
    The CCF can only resolve shifts that are multiples of the sampling interval, so daily data cannot see an intraday lead at all. Aggregating to a coarser bucket can also move or erase a peak: a genuine 3-day lead largely disappears once the data is rolled up to weeks, and misaligned bucket boundaries can invent apparent leads.

saying these in an interview costs you the question

  • Reports a lead direction without checking the lag convention
  • Reads a peak off two strongly trending series untouched
  • Treats the standard significance bands as valid for autocorrelated inputs
  • Calls the leading series the cause of the other
  • Assumes the peak lag stays stable out of sample

context