skip to content

Regression Discontinuity

When treatment switches on at a cutoff score, units just above and just below it are nearly comparable. Interviewers ask about bandwidth choice and whether people can game the running variable.

on this pageshow

questions

6

In a sharp regression discontinuity, what assumption identifies the effect and what does it estimate?

level: middleimportance: must knowfreq 70%

answer

  1. treatment is a step function of score
  2. nothing else may jump at the cutoff
  3. smooth counterfactual means through the threshold
  4. the estimate is local to the marginal unit
  5. untestable directly, but has testable implications

basics

~20 s

Identification rests on continuity: average outcomes with and without treatment must vary smoothly through the cutoff, so nothing but treatment jumps there. The estimate is the average treatment effect for units sitting at the cutoff, not for everyone.

solid answer

~50 s

In a sharp design, treatment is a deterministic step function of the running variable: everyone at or above the cutoff is treated, everyone below is not. Take a merit scholarship awarded to every applicant scoring at least 60 on an entrance exam. The identifying assumption is **continuity of average potential outcomes at the cutoff**: the average outcome an applicant would have with the scholarship, and the average outcome they would have without it, are both smooth functions of the exam score through 60. Under that assumption, the vertical gap between the fitted curve approaching 60 from above and the one approaching from below equals the average treatment effect for applicants whose score is 60. That is a *local* estimand — it is the effect at the threshold, not the effect for the whole applicant pool. Continuity is untestable directly, but it has testable implications you should check.

go deeper

for a junior

Be ready to state that treatment is switched on by a fixed threshold and that the effect is read as the jump in the outcome at that threshold, using a concrete rule such as a scholarship at a passing score.

for a middle

Explain the mechanics: separate trends on each side, the score centred at the cutoff, and the treatment coefficient reading the jump. State the continuity assumption in your own words rather than reciting it.

for a senior

Demonstrate that you check what you can — covariate smoothness, a placebo cutoff, the plot — and that you write the estimand into the report so nobody downstream reads the number as a population-wide effect.

for a principal

Own the decision framing: a locally identified number is worth commissioning when the choice on the table is where to draw the eligibility line, and worth much less when leadership wants a global rollout estimate.

## The sharp design, precisely Write the running variable as `X` (the entrance exam score), the cutoff as `c = 60`, and the treatment indicator as `D`. In a **sharp** design, `D = 1 if X >= 60, and D = 0 if X < 60` with no exceptions. Every applicant scoring 60 or more receives the merit scholarship; nobody scoring 59 receives it. Treatment is a deterministic function of the score. Each applicant has two potential outcomes: `Y(1)`, the graduation outcome they would have with the scholarship, and `Y(0)`, the outcome they would have without it. Only one is ever observed, which is the fundamental problem the design has to work around. ## The identifying assumption The assumption is **continuity of the conditional mean potential outcomes at the cutoff**: `E[Y(1) | X = x]` and `E[Y(0) | X = x]` are both continuous in `x` at `x = 60`. Read that in plain words. Among applicants who *would not* get the scholarship, average graduation is a smooth function of exam score — a 59.9 scorer and a 60.1 scorer would be nearly identical in the untreated world. The same holds in the treated world. There is nothing magical about the number 60 itself, other than the scholarship rule attached to it. Notice what this does *not* say. It does not say the score is unrelated to graduation; higher scorers almost certainly graduate more, and the fitted trend on each side absorbs exactly that. It does not require randomisation. It does not require the two curves to be linear, or parallel, or anything else about their shape away from the cutoff. ## What follows: the estimand Under continuity, `lim(x -> 60 from above) E[Y | X = x] - lim(x -> 60 from below) E[Y | X = x] = E[Y(1) - Y(0) | X = 60]` The observed jump equals the average treatment effect **for units at the cutoff**. This is the whole payoff, and it is the sentence to be able to say precisely. The estimand is local in a specific sense: it describes applicants whose score is 60, the marginal ones. It does not describe the applicant who scored 90 and would have got the scholarship under any plausible rule, nor the applicant who scored 30. This is not a flaw to apologise for — it is often the exact policy quantity, because the live decision is usually whether to move the line to 58 or 62, and the marginal applicant is the one that decision is about. ## Estimation in practice Estimation means approximating those two limits. Fit a regression of the outcome on the score using only observations within a window around 60, separately on each side, and read the gap at the cutoff. A common specification pools the two sides: `Y = a + b*(X - 60) + t*D + g*D*(X - 60) + error`, fitted for `|X - 60| <= h` Here `t` is the estimated jump at the cutoff, the coefficient of interest, and centring the score at 60 is what makes `t` read off at the threshold rather than at score zero. Allowing separate slopes on each side (`g`) matters: forcing one common slope pushes curvature into the jump. ## Why continuity cannot be tested, and what you check instead Continuity is a statement about counterfactual averages, one of which is never observed on either side. You cannot test it. You can test its implications: - **Covariate smoothness.** Characteristics fixed before the exam — prior schooling, age, region — should not jump at 60. If they do, applicants on the two sides differ for reasons other than the scholarship, and the continuity story is in trouble. - **Density of the running variable.** If applicants can nudge their own score, the number of applicants just above 60 will be inflated relative to just below. - **Placebo cutoffs.** Re-run the estimate at fake thresholds like 50 or 70, where no rule exists. Repeated "significant" jumps there mean your specification manufactures jumps. - **A picture.** Plot binned outcome means against score with the fitted curves. If the jump is not visible to the eye at a sensible bin width, be sceptical of a jump that only a regression finds. ## Common mistakes Saying "RD gives you the average treatment effect" without the qualifier is the single most common error, and it is the one interviewers listen for. A second is describing continuity as "units near the cutoff are randomly assigned" — a useful intuition, but strictly it is a stronger claim than continuity, and stating it as the assumption suggests the definition was memorised rather than understood. A third is treating the running variable as a confounder to control away with a covariate adjustment, which destroys the design: the score is the axis of the design, not a nuisance. ## Reporting Report the estimated jump with its standard error, the window used, the specification, the picture, and the checks above. Say explicitly, in words, which units the number describes: applicants scoring around 60.

  • Is continuity of potential outcomes the same as saying treatment is randomised near the cutoff?
    No. Local randomisation is a stronger claim — that treatment behaves as if randomly assigned within a small window — and it implies continuity, not the other way round. Continuity only requires the counterfactual means to have no break at the threshold, which is weaker and is the standard assumption behind the local-linear estimator.
  • Why centre the running variable at the cutoff in the regression?
    So that the treatment-indicator coefficient reads the jump at the cutoff rather than at score zero. With an uncentred score and separate slopes, the treatment coefficient is an extrapolated intercept difference far from the threshold, which is not the quantity the design identifies.
  • If a pre-treatment covariate jumps at the cutoff, what do you do?
    Treat it as evidence against continuity rather than something to control away. Investigate why: manipulation of the score, a second rule sharing the cutoff, or a data artefact such as re-scoring near the boundary. Adding the covariate as a control does not repair a design whose core assumption has failed.
  • How would you explain the estimand to a non-technical stakeholder?
    Say it measures what the scholarship does for applicants right at the borderline — the ones who barely qualified or barely missed. That is the group affected by moving the passing mark, so it answers the question about where to set the line, but it does not tell you the effect on top scorers.

saying these in an interview costs you the question

  • Calls the estimate the average treatment effect for everyone
  • Says the design assumes no relationship between score and outcome
  • Adds the running variable as a control to remove its effect
  • Forces one common slope across both sides of the cutoff
  • Claims the continuity assumption can be tested directly

context

open as a page

How do you check whether users manipulated the running variable around a regression discontinuity cutoff?

level: seniorimportance: must knowfreq 55%

basics

~20 s

The McCrary test checks whether the density of the running variable jumps at the cutoff. Bunching just above the threshold means units steered themselves across, so the two sides are no longer comparable and the design fails.

open as a page

What is the running variable in a regression discontinuity, and what makes age-65 Medicare eligibility a good one?

level: juniorimportance: should knowfreq 45%

basics

~20 s

The running variable is the score that decides treatment: everyone past a fixed cutoff gets it, everyone below does not. Age works well for Medicare eligibility at 65 because nobody can nudge their birth date across the threshold.

open as a page

How does a fuzzy regression discontinuity differ from a sharp one, and what does it estimate?

level: middleimportance: should knowfreq 50%

basics

~10 s

A sharp cutoff switches treatment on for certain; a fuzzy one only raises its probability. The effect is then the jump in the outcome divided by the jump in treatment probability at the cutoff.

open as a page

How would you use an interrupted time series to measure a site-wide redesign shipped on 1 March?

level: seniorimportance: should knowfreq 42%

basics

~20 s

Treat the ship date as a cutoff on calendar time: fit the pre-period trend, then allow both an immediate level shift and a later slope change. The main threat is anything else that shipped that day.

open as a page

How do you choose the bandwidth and functional form for a regression discontinuity estimate?

level: seniorimportance: nice to knowfreq 38%

basics

~20 s

Fit a straight line on each side of the cutoff inside a narrow window and read the gap. A wider window buys precision but imports bias from curvature; a global high-order polynomial can manufacture a jump.

open as a page