skip to content

Quasi-Experimental Designs

Designs that borrow identification from the world: a policy start date, an eligibility cutoff, a lottery-like nudge. Interviewers ask which assumption each design buys and how you would test it.

on this pageshow

explore

questions

17

What does a difference-in-differences estimate compute from a two-group, two-period panel?

level: juniorimportance: must knowfreq 70%

answer

  1. two subtractions, not one
  2. cancels a fixed gap between groups
  3. cancels a shock hitting both groups
  4. the interaction coefficient in the regression

basics

~20 s

It subtracts the comparison group's before-to-after change from the treated group's before-to-after change. That double difference cancels the fixed level gap between the groups and any shock that moved both of them over the same window.

solid answer

~50 s

Difference-in-differences uses four means: the treated group before and after, and the comparison group before and after. The estimate is `(treated_after - treated_before) - (control_after - control_before)`. The first difference removes anything about the treated group that stayed constant across the window — its baseline level, its composition, its market. The second difference removes whatever moved both groups between the two periods — seasonality, a macro shift, a platform-wide change. What is left is attributed to the treatment. The same number falls out of a regression of the outcome on a treated-group indicator, a post-period indicator, and their interaction: the interaction coefficient is the estimate, and the regression form gives you standard errors and room for covariates. The two groups do not need equal starting levels; the design assumes only that their changes would have matched.

go deeper

for a junior

Be ready to compute the double difference from four means on the spot, and to say in one sentence what each of the two subtractions removes.

for a middle

Expect to write the regression form — group indicator, post indicator, interaction — and explain why the interaction coefficient is the estimate and what the other two coefficients absorb.

for a senior

Show you know what the number is an estimate of: the effect on treated units over the observed window, with standard errors clustered at the level treatment was assigned.

for a principal

Own the framing decision: whether a two-period double difference is the right readout at all, given the rollout shape, the length of panel you have, and how reversible the decision is.

## The four numbers Difference-in-differences (DiD) is the simplest credible way to get a causal number when you could not randomise. You need two groups — one that gets the treatment and one that does not — observed in two periods, one before the treatment starts and one after. That gives you a 2x2 table of average outcomes: | | Before | After | | --- | --- | --- | | Treated group | A | B | | Comparison group | C | D | The estimate is `(B - A) - (D - C)`. Equivalently, it is `(B - D) - (A - C)`: the post-period gap between the groups minus the pre-period gap. Both orderings give the same number, and each one makes a different intuition visible. The first says *compare the changes*. The second says *compare the gaps*. ## What each subtraction buys The first difference, `B - A`, is a simple before-and-after on the treated group. It removes everything about that group that did not change over the window: how big its market was, how its users were composed, whatever fixed advantage or disadvantage it started with. What it does not remove is time. If the whole market grew 8% that quarter, `B - A` contains that 8%. The second difference, `D - C`, is the comparison group's before-and-after. Under the design's assumption, that change is an estimate of what would have happened to the treated group anyway. Subtracting it strips out the common movement. This is the step that makes DiD more than a naive before-after. So DiD needs neither group to look like the other in *level*. A treated market that is three times the size of the comparison market is fine — the size difference sits in both `A` and `B` and cancels in the first difference, and it sits in the pre-period gap and cancels in the second ordering too. What must match is the *change*, and that requirement is the parallel-trends assumption. ## The regression form In practice you rarely compute four means by hand. You fit `Y = b0 + b1*Treated + b2*Post + b3*(Treated x Post) + error` on the stacked observations, where `Treated` is 1 for units in the treated group and `Post` is 1 for observations in the after period. Then: - `b0` is the treated-group-excluded baseline: the comparison group before treatment. - `b1` is the fixed level gap between the groups. - `b2` is the common time movement affecting both groups. - `b3` is the difference-in-differences estimate — the extra movement in the treated group after treatment began. The regression form is the one to reach for because it generalises: you can add covariates, add unit and period fixed effects to move beyond two groups and two periods, and it gives you standard errors. Those standard errors should be clustered at the level treatment was assigned — usually the unit (state, market, restaurant), not the observation — because outcomes within a unit are correlated over time and ignoring that badly understates the uncertainty. ## The canonical example The study that made the design famous compared fast-food employment in New Jersey, which raised its state minimum wage in April 1992, with fast-food employment in eastern Pennsylvania, which did not. Employment levels in the two areas were not identical to begin with — that was never the point. The claim was that, absent the wage change, employment in the two neighbouring areas would have moved together. Under that claim, the double difference is the effect of the wage increase, and the finding was that New Jersey employment did not fall relative to Pennsylvania's. ## What the number means The estimate is an **average treatment effect on the treated (ATT)**: the average effect for the units that actually received the treatment, over the post-period you observed. It is not the effect the comparison group would have experienced, and it is not the effect over a longer horizon. Two honest qualifications belong in any report of it: the window it covers, and the assumption it rests on. Finally, DiD is not magic. It is a subtraction that is only as good as the claim that the comparison group's change stands in for the treated group's counterfactual change. Everything hard about the design lives in defending that claim, not in the arithmetic.

  • Which coefficient in a regression with group, period and interaction terms is the difference-in-differences estimate?
    The coefficient on the interaction of the treated-group indicator with the post-period indicator. The group indicator absorbs the fixed level gap between the groups, the post indicator absorbs the movement common to both periods, and the interaction picks up the extra movement in the treated group after treatment started — that is the estimate.
  • Does difference-in-differences estimate the effect on everyone, or only on the treated?
    Only on the treated — the average treatment effect on the treated, over the post-period you observed. It says nothing about how the comparison group would have responded had they been treated, and nothing about periods outside your window. Report the horizon alongside the number.
  • Why is a simple before-and-after on the treated group usually not enough?
    Because everything else that changed between the two periods is baked into that difference: seasonality, a pricing change, a macro shift, a platform release. The comparison group's before-and-after change is your estimate of that common movement, and subtracting it out is the whole point of the design.

Two hikers set off from different altitudes. You do not compare where they stand — you compare how much each one climbed in the same hour, and the difference is what the extra gear was worth.

saying these in an interview costs you the question

  • Reports the treated group's before-after change as the effect
  • Claims the two groups must start at equal outcome levels
  • Treats the comparison group's post-period mean as the counterfactual
  • Calls the estimate causal without naming any assumption

context

open as a page

How does two-stage least squares turn an instrument into a causal estimate?

level: middleimportance: must knowfreq 62%

basics

~20 s

Two-stage least squares first predicts the treatment from the instrument, then relates the outcome to that prediction. Only the instrument-driven part of the treatment's variation is used, and that part is assumed free of the confounding.

open as a page

What assumptions must an instrumental variable satisfy to identify a causal effect?

level: middleimportance: must knowfreq 72%

basics

~20 s

A valid instrument needs relevance, meaning it actually shifts the treatment, and the exclusion restriction, meaning it affects the outcome only through that treatment. It must also be as good as randomly assigned with respect to unmeasured confounders.

open as a page

In a sharp regression discontinuity, what assumption identifies the effect and what does it estimate?

level: middleimportance: must knowfreq 70%

basics

~20 s

Identification rests on continuity: average outcomes with and without treatment must vary smoothly through the cutoff, so nothing but treatment jumps there. The estimate is the average treatment effect for units sitting at the cutoff, not for everyone.

open as a page

How do you check whether users manipulated the running variable around a regression discontinuity cutoff?

level: seniorimportance: must knowfreq 55%

basics

~20 s

The McCrary test checks whether the density of the running variable jumps at the cutoff. Bunching just above the threshold means units steered themselves across, so the two sides are no longer comparable and the design fails.

open as a page

What is the running variable in a regression discontinuity, and what makes age-65 Medicare eligibility a good one?

level: juniorimportance: should knowfreq 45%

basics

~20 s

The running variable is the score that decides treatment: everyone past a fixed cutoff gets it, everyone below does not. Age works well for Medicare eligibility at 65 because nobody can nudge their birth date across the threshold.

open as a page

How does a fuzzy regression discontinuity differ from a sharp one, and what does it estimate?

level: middleimportance: should knowfreq 50%

basics

~10 s

A sharp cutoff switches treatment on for certain; a fuzzy one only raises its probability. The effect is then the jump in the outcome divided by the jump in treatment probability at the cutoff.

open as a page

Why can two-way fixed effects be biased when a policy rolls out at staggered dates?

level: seniorimportance: should knowfreq 44%

basics

~20 s

With staggered adoption, the two-way fixed-effects coefficient averages many two-group comparisons, some using already-treated units as the control. When effects change over time those comparisons can take negative weight and pull the estimate the wrong way.

open as a page

Whose treatment effect does an instrumental variable actually estimate?

level: seniorimportance: should knowfreq 52%

basics

~20 s

Only the compliers, the units whose treatment status the instrument actually changed. That is the local average treatment effect. Units who would take the treatment regardless, or refuse it regardless, supply no identifying variation and contribute nothing to the estimate.

open as a page

What does a first-stage F statistic of 4 imply about a 2SLS estimate?

level: seniorimportance: should knowfreq 48%

basics

~20 s

A first-stage F of 4 signals a weak instrument. The two-stage least squares estimate is biased back toward the confounded naive comparison, standard errors inflate, and conventional confidence intervals undercover. Around 10 is the classic minimum.

open as a page

How would you use an interrupted time series to measure a site-wide redesign shipped on 1 March?

level: seniorimportance: should knowfreq 42%

basics

~20 s

Treat the ship date as a cutoff on calendar time: fit the pre-period trend, then allow both an immediate level shift and a later slope change. The main threat is anything else that shipped that day.

open as a page

A launch can only be evaluated with difference-in-differences — how do you decide whether to act on it?

level: principalimportance: should knowfreq 34%

basics

~10 s

Grade the design before the number: why randomisation was impossible, how the comparison group was chosen, whether the pre-adoption path is flat. Then match the evidence bar to how reversible the decision is.

open as a page

When does synthetic control beat picking a single comparison group for a policy evaluation?

level: seniorimportance: nice to knowfreq 30%

basics

~20 s

Synthetic control helps when one treated unit has no credible twin. It builds the comparison as a weighted average of many untreated units, with weights chosen so the composite tracks the treated unit's pre-treatment path.

open as a page

How do you choose the bandwidth and functional form for a regression discontinuity estimate?

level: seniorimportance: nice to knowfreq 38%

basics

~20 s

Fit a straight line on each side of the cutoff inside a narrow window and read the gap. A wider window buys precision but imports bias from curvature; a global high-order polynomial can manufacture a jump.

open as a page

How would you design an encouragement experiment when you cannot randomize the treatment itself?

level: principalimportance: nice to knowfreq 30%

basics

~20 s

Randomize the invitation rather than the treatment, leave uptake voluntary, then use the nudge as an instrument for actual use. The nudge must move uptake substantially and must not affect the outcome by any other route.

open as a page