Why can two-way fixed effects be biased when a policy rolls out at staggered dates?
answer
- a weighted average of many 2x2s
- some comparisons use already-treated units
- weights can be negative
- dynamic effects are what break it
- compare against never- or not-yet-treated
basics
~20 sWith staggered adoption, the two-way fixed-effects coefficient averages many two-group comparisons, some using already-treated units as the control. When effects change over time those comparisons can take negative weight and pull the estimate the wrong way.
solid answer
~50 sA single treatment indicator in a regression with unit and period fixed effects looks like difference-in-differences generalised to many units and dates, but with staggered adoption it is not one clean comparison. The Goodman-Bacon decomposition shows the coefficient is a weighted average of every possible two-group two-period comparison in the panel — including ones where a later-adopting unit is compared against a unit that was *already treated*. In those forbidden comparisons the supposed control's outcome already contains its own treatment effect. If effects grow over time, the already-treated group's rising path is subtracted from the newly treated group's, which enters the average with the wrong sign and can carry a negative weight. The estimate can then be attenuated or even flipped when every unit-level effect is positive. The fix is a heterogeneity-robust estimator: compute a separate effect per adoption cohort and period, using only never-treated or not-yet-treated units as comparisons, then aggregate deliberately.
go deeper
Know that a rollout arriving at different dates for different units is not the same as one common start date, and that the simple treatment-indicator regression is not automatically the effect.
Be ready to explain that the fixed-effects coefficient is a weighted average of many two-group comparisons, and that some of those use units which have already been treated.
Demonstrate the mechanism — dynamic effects turning already-treated comparisons negative — and name a concrete remedy: cohort-by-period effects against never- or not-yet-treated units, aggregated deliberately.
Own the call on when the extra machinery is worth it: whether the rollout shape and effect dynamics in your setting could plausibly flip a sign, and what you standardise across teams reporting rollout readouts.
## The setup A state-by-state policy rollout is the canonical case: some states adopt in 2019, some in 2021, some never. The reflexive specification is `Y = unit effects + period effects + beta * Treated_it + error` where `Treated_it` is 1 once unit `i` has adopted. This is called two-way fixed effects (TWFE), and for a decade `beta` was reported as *the* difference-in-differences estimate. With a single common adoption date it is exactly that. With staggered dates it is something else. ## The decomposition The Goodman-Bacon decomposition establishes that with staggered timing, `beta` equals a weighted average of all the two-group two-period comparisons available in the panel. Those comparisons come in three flavours: 1. **Treated cohort versus never-treated units.** Clean. This is what everyone thinks they are estimating. 2. **Earlier-adopting cohort versus later-adopting cohort, before the later one adopts.** Also clean — the later cohort is genuinely untreated over that window. 3. **Later-adopting cohort versus earlier-adopting cohort, after the earlier one has already adopted.** This one is the problem. The comparison group is already treated. The weights are functions of group sizes and of how much treatment variance each comparison contributes, and they sum to one. ## Why the third comparison misleads In a forbidden comparison, the double difference is `(late group's change) - (early group's change)` but the early group's change *includes the evolution of its own treatment effect*. If the effect is constant over time, that evolution is zero and nothing goes wrong — the already-treated group's effect is a fixed level shift that differences out. That is why TWFE is fine under homogeneous, time-constant effects. Once effects are **dynamic** — they build over periods, or decay — the early group's effect is still moving during the comparison window. Subtracting a rising treatment effect from the late group's change subtracts something that has nothing to do with the late group's treatment. The comparison contributes with an effectively negative sign. Aggregate that with positive weights on the clean comparisons and you can get an estimate that is attenuated toward zero, or, when the dynamic component is large enough and the forbidden comparisons carry enough weight, a `beta` with the **opposite sign from every unit-level effect in the data**. That is the striking result: a policy that helped every state can produce a negative TWFE coefficient. ## Heterogeneity across cohorts A second, related failure: even without dynamics, if different adoption cohorts have genuinely different effects — early adopters gained more because they were better suited to the policy — TWFE weights those cohorts by how much treatment variance they supply, not by anything you would choose. Units treated near the middle of the panel get the most weight; units treated at the very start or the very end get almost none. The resulting number is a weighted average with weights nobody would defend if they were written down. ## What to do instead **Group-time estimators.** Estimate a separate average effect for each adoption cohort `g` and each period `t`, always using a clean comparison set: never-treated units, or units not yet treated as of `t`. This gives a grid of `ATT(g, t)` estimates. Then aggregate them with weights *you* choose — by cohort, by time since adoption (which yields an event-study path), or by an overall average weighted by cohort size. The Callaway and Sant'Anna estimator is the standard reference for this approach. **Interaction-weighted event studies.** The Sun and Abraham estimator fixes the analogous contamination in a naive event-study regression, where a relative-time coefficient can absorb effects from other cohorts at other relative times, by interacting relative-time indicators with cohort and reweighting. **Stacked designs.** For each adoption cohort, build a small clean dataset of that cohort plus its not-yet-treated comparisons over a fixed window around its adoption, then stack all those datasets and estimate once. Easy to explain and audit, at the cost of using some units repeatedly. **Drop the already-treated.** The blunt version: restrict comparisons to never-treated units only. Simple and clean, but it throws away data and requires a never-treated group to exist, which it often does not in a full rollout. ## What an interviewer is listening for The mechanism, not the citation. Say: with staggered timing the coefficient is an average of many 2x2 comparisons; some use already-treated units as controls; when effects vary over time those comparisons enter with the wrong sign; therefore estimate cohort-by-period effects against clean comparisons and aggregate explicitly. Then add the honest boundary: with a single common adoption date, or genuinely constant homogeneous effects, plain TWFE is fine and this whole discussion is moot.
- Which two-by-two comparison in a staggered design is the problematic one?A later-adopting cohort compared against a cohort that has already adopted. The supposed control's outcome already contains its own treatment effect, so if that effect is still evolving, its movement is subtracted from the newly treated cohort's change and enters the aggregate with the wrong sign.
- When is plain two-way fixed effects still fine despite staggered timing?When treatment effects are constant over time and the same across cohorts. Then an already-treated comparison unit contributes only a fixed level shift, which differences out, and every two-by-two comparison estimates the same quantity. A single common adoption date is also safe, because no forbidden comparisons exist.
- How does a group-time estimator fix the problem?It estimates a separate average effect for each adoption cohort in each period, always against never-treated or not-yet-treated units, so no already-treated unit is ever used as a control. You then aggregate those cell estimates with weights you choose explicitly — by cohort, by time since adoption, or overall — rather than accepting whatever weights the regression implies.
- Does adding more fixed effects or covariates solve the negative-weight problem?No. The problem is which comparisons the estimator is implicitly forming, not omitted variables. Extra fixed effects change the weighting but leave already-treated units serving as controls. The fix has to change the comparison set, which means a different estimator, not a richer specification.
saying these in an interview costs you the question
- Assumes the fixed-effects coefficient is a simple average of unit effects
- Ignores that already-treated units are being used as comparisons
- Says staggered timing costs precision but never bias
- Thinks adding covariates or more fixed effects removes the problem
- Cannot say why constant effects make the problem disappear