skip to content

How do you use an event-study plot to defend a difference-in-differences design?

level: seniorimportance: should knowfreq 52%

answer

  1. one coefficient per period, not one number
  2. leads before, lags after adoption
  3. one period is the omitted reference
  4. flat, tight leads are the evidence
  5. wide intervals make flatness meaningless

basics

~20 s

Estimate one coefficient per period relative to the adoption date, with one pre-period as the omitted reference, and plot them with confidence intervals. Flat, tight pre-adoption coefficients support parallel trends; the post-adoption path shows how the effect evolves.

solid answer

~50 s

Instead of collapsing everything into one before-versus-after number, I estimate a coefficient for each period relative to the adoption date — leads before it, lags after it — omitting one pre-period, usually the period immediately before adoption, as the reference. Then I plot all of them with intervals. Take a feature shipped in Canada with the US as the comparison: if the leads drift upward, Canadian usage was already climbing faster than US usage before the feature existed, and the single difference-in-differences number is reading that slope as an effect. Flat, tightly estimated leads are the credibility evidence. Two cautions. Wide intervals make flat leads meaningless — a non-significant pre-trend is not evidence of no pre-trend, and these checks are usually underpowered. And the lags carry information of their own: a jump that decays is a novelty effect, a slope that keeps climbing is gradual adoption, a jump that persists is a level shift.

go deeper

for a junior

Know that an event study replaces the single before-after number with one coefficient per period around the adoption date, and that the pre-adoption ones are what you inspect first.

for a middle

Be ready to describe the specification: relative-time indicators, one omitted reference period, and why every coefficient is read relative to that anchor rather than in absolute terms.

for a senior

Show judgment about power: say what magnitude of pre-trend the data could not have ruled out, and describe what you would actually do when the leads slope rather than just reporting them.

for a principal

Own how these plots are used organisationally: whether a sloping pre-trend blocks a launch decision, and how to stop teams from shopping for the comparison group whose leads look flattest.

## The specification An event study is difference-in-differences with the single post indicator replaced by a full set of period indicators measured relative to the adoption date. For a unit `i` in calendar period `t` that adopts at time `g`, define relative time `k = t - g`. You fit something of the form `Y = unit effects + period effects + sum over k of (beta_k * 1[relative time = k]) + error` with one value of `k` omitted. The `beta_k` for negative `k` are the **leads** — differences between treated and comparison units before anything happened. The `beta_k` for `k >= 0` are the **lags** — the effect path after adoption. **Why one period must be omitted.** The full set of relative-time indicators is collinear with the unit and period fixed effects, so the coefficients are only identified up to a normalisation. You drop one, conventionally `k = -1` (the period just before adoption), and every plotted coefficient is then read as *relative to that period*. This is not a modelling nicety you can skip; without it the model does not have a unique solution. It also means the plot is anchored, not absolute: shifting the reference period shifts the whole curve. **Binning the endpoints.** Far-out relative periods are estimated off very few units and become extremely noisy. The usual fix is to bin everything beyond some horizon into a single terminal lead and a single terminal lag, so the tails of the plot are estimated with enough data to mean something. ## Reading the leads The leads are the evidence for parallel trends. Three patterns matter. **Flat and tight.** Coefficients hovering near zero with narrow intervals across several pre-periods. This is what you want: before the treatment existed, the two groups moved together. **Sloping.** Coefficients drifting in one direction as adoption approaches. This is the Canada-versus-US case: usage in the treated market was already growing faster, and a two-period double difference would extrapolate that pre-existing divergence into the post-period and call it an effect. The direction matters — a rising pre-trend inflates a positive estimate. **A spike right before adoption.** Often anticipation: units changed behaviour once they knew the treatment was coming, or the treatment was targeted at units that had just moved. Either way, the period immediately before adoption is a bad reference, and you should re-anchor further back and see what the picture looks like. ## The honest caveats **Underpowered checks.** The leads are estimated on the same, often small, sample as the main effect, so their intervals are frequently wide enough to contain economically enormous pre-trends. 'No significant pre-trend' then carries almost no information. The right way to report it is to say how large a pre-trend the data could not have detected, not to say the test passed. **Selection on passing.** If you only publish results whose pre-trend check came out flat, you have selected on a noisy statistic. Conditioning on a passed pre-test can leave the surviving estimates *more* biased than the unconditional ones, because the cases that survive are those where noise happened to flatten a real trend. **Flat leads are not sufficient.** Parallel trends is about the post-period counterfactual. Nothing about pre-period agreement guarantees the groups would have kept agreeing. The leads make the assumption plausible; they do not deliver it. ## What to do when the leads slope In rough order of preference: 1. **Get a better comparison group.** A sloping pre-trend usually means the comparison unit is substantively wrong — a different growth stage, a different market maturity. Reselect on the mechanism rather than on the fit. 2. **Extend the pre-period.** A slope over three periods may be noise; over twelve it is a fact about the world. More pre-periods also make the check informative rather than decorative. 3. **Allow group-specific linear trends.** You can add a separate linear time trend per unit, which absorbs a steady pre-existing slope. The serious warning is that this also absorbs any treatment effect that grows linearly, so it can flatten a genuine effect to nothing. Report both with and without. 4. **Report a range rather than a point.** Ask how big a differential trend would have to be to explain the whole estimate away, and say whether that magnitude is plausible. 5. **Abandon difference-in-differences here.** Sometimes the honest answer is that the comparison group does not support a causal claim, and the readout should be labelled descriptive. ## Reading the lags The post-adoption coefficients are not just a robustness display; they are the substantive result. A single collapsed number hides whether the effect was a one-off jump that faded, a step change that persisted, or something that took three periods to appear. Product decisions differ sharply between those shapes, and the event-study plot is what lets you distinguish them.

  • Why must one pre-treatment period be omitted from an event-study specification?
    The full set of relative-time indicators is collinear with the unit and period fixed effects, so the coefficients are identified only up to a normalisation. Dropping one period — conventionally the one just before adoption — fixes that anchor, and every plotted coefficient is then read relative to it. Move the reference and the whole curve shifts.
  • Your pre-period leads slope upward. What are your options?
    Reselect the comparison group on substance rather than fit; extend the pre-period so you can tell a real slope from noise; allow group-specific linear trends, while warning that this also absorbs any linearly growing treatment effect; or state how large a differential trend would be needed to explain the estimate away. If none is defensible, label the readout descriptive.
  • Is a non-significant pre-trend test enough to accept parallel trends?
    No. These checks run on the same small sample as the main effect and are typically underpowered, so wide intervals can hide a large pre-trend. Report the magnitude the data could not have detected rather than a pass or fail. Publishing only estimates that survive the check also selects on noise and can worsen bias.
  • What can the post-adoption coefficients tell you that the single estimate cannot?
    The shape of the effect over time. A jump that decays suggests novelty; a curve that keeps climbing suggests gradual adoption; a jump that holds suggests a persistent level shift. A collapsed average of those three paths would look identical, yet each implies a different rollout decision.

saying these in an interview costs you the question

  • Treats flat leads as proof that parallel trends holds
  • Reads wide, noisy leads as evidence of no pre-trend
  • Forgets that one period must be the omitted reference
  • Plots only post-adoption coefficients and calls it an event study
  • Adds group-specific trends without noting they absorb real effects

context