How do you choose the bandwidth and functional form for a regression discontinuity estimate?
answer
- a window, not the whole range
- one straight line on each side
- wider window, more data, more bias
- high-order global fits invent jumps
- show the estimate across many bandwidths
basics
~20 sFit a straight line on each side of the cutoff inside a narrow window and read the gap. A wider window buys precision but imports bias from curvature; a global high-order polynomial can manufacture a jump.
solid answer
~50 sThe default is a **local linear** fit: keep observations within a bandwidth `h` of the cutoff, fit a separate line on each side, often weighting nearer points more heavily, and read the vertical gap at the threshold. Bandwidth is a bias-variance dial — widen it and you gain observations and precision but import curvature far from the cutoff as bias; narrow it and bias falls while the standard error grows. Data-driven selectors choose `h` to minimise mean squared error, and because the resulting estimate is deliberately slightly biased, bias-corrected intervals should be reported rather than naive ones. The alternative some candidates reach for — a **global fourth-order polynomial** fitted across the whole range — is discouraged: it gives points far from the cutoff heavy influence at the boundary, and it can produce a large apparent jump purely from wiggle. The credibility check is a sensitivity plot of the estimate against bandwidth.
go deeper
Know that the fit uses only observations near the cutoff and that a straight line on each side is the usual choice, rather than one curve fitted across all the data.
Explain the bias-variance direction precisely — wider window means more data and more bias, narrower means less bias and a bigger standard error — and why each side gets its own slope.
Show the workflow you would defend in review: a pre-committed data-driven bandwidth, bias-corrected intervals, a sensitivity plot across bandwidths, and a refusal to headline a global high-order fit.
Set the norm that specification choices are fixed before estimates are seen, so that bandwidth selection cannot become an unlogged degree of freedom in decisions the organisation acts on.
## The estimation problem The quantity you want is the gap between two limits at the cutoff: the average outcome approached from above minus the average approached from below. Neither limit is observed directly, so both are approximated by regression, and every approximation choice — how far from the cutoff to look, what shape to fit, how to weight points — moves the answer. ## Local linear as the default The standard estimator restricts attention to observations with `|X - c| <= h` and fits a separate straight line on each side, then reads the two fitted values at `c`. Equivalently, fit `Y = a + b*(X - c) + t*D + g*D*(X - c) + error` on `|X - c| <= h` where `D` marks the treated side, and `t` is the estimated jump. Separate slopes matter: forcing a common slope makes any difference in curvature between the sides leak into `t`. A kernel weight is usually applied so that points nearer the cutoff count more — a triangular weight declining linearly to zero at the bandwidth edge is the common choice. The intuition is that the observations closest to the threshold carry the identifying information, and the far ones are there mainly to pin down the slope. ## The bandwidth tradeoff Bandwidth trades bias against variance, and the direction is worth stating exactly: - **Wider `h`**: more observations, so lower variance and a tighter interval. But the straight line must now approximate the true curve over a longer stretch, and if the outcome-score relationship bends, that misfit is absorbed into the estimated jump. Bias rises. - **Narrower `h`**: the straight-line approximation is better over a shorter stretch, so bias falls. But fewer observations mean a larger standard error, and past some point the estimate is too noisy to be useful. At the mean-squared-error-optimal bandwidth, the estimator is deliberately left with some bias — that is what minimising MSE means. Consequently a conventional confidence interval centred on the estimate under-covers, and the modern practice is to report a bias-corrected estimate with an interval widened to account for the correction. Say this if asked why the interval from an automatic selector looks wider than the naive one. ## Why global high-order polynomials are discouraged An appealing but poor alternative is to fit one high-order polynomial — say fourth order — in the running variable across the entire data range, with a treatment indicator for the jump. Three problems: 1. **Distant points get weight at the boundary.** The estimate at the cutoff depends on observations far away, which the design gives no reason to trust for the local comparison. 2. **Boundary behaviour is unstable.** Polynomial fits are noisiest at the edges of their range, and the cutoff is exactly an edge for each side's fit. 3. **The order is arbitrary and the answer moves with it.** A jump that appears at fourth order and vanishes at third order is not evidence about treatment; it is evidence about the polynomial. The practical consequence is that a global fourth-order fit can manufacture a visible jump at the cutoff from ordinary curvature. If you use one at all, use it as a secondary check, not as the headline specification, and show the local linear result beside it. ## The sensitivity plot The single most convincing artefact is a plot of the estimated jump, with its confidence band, against a range of bandwidths — say from a quarter of the selected value to twice it. What you want to see is a flat region: the estimate roughly constant, with the interval widening as `h` shrinks. What condemns an analysis is an estimate that changes sign or size systematically as the window moves, or one that is significant at a single bandwidth and nowhere else. Report the plot, not just the number. Complementary checks: repeat with and without kernel weighting, with a local quadratic instead of a local linear fit, and with a placebo cutoff where no rule exists. ## Practical constraints - **Discrete running variables.** If the score takes few distinct values, there may be no meaningful choice of a small bandwidth, and the standard errors need to reflect the coarseness of the score rather than the number of units. - **Too few observations near the cutoff.** Sometimes the honest conclusion is that the design cannot be run at a defensible bandwidth. Widening `h` until something is significant is specification searching. - **Pre-registration.** Choosing the bandwidth after seeing which one gives a significant answer invalidates the inference. Fix the rule — a named data-driven selector, plus the sensitivity range you will show — before looking at the estimates. ## How this is asked Usually as a follow-up after the basic design: "and how would you actually fit it?" A strong answer names local linear as the default, describes the bias-variance direction correctly, mentions data-driven bandwidth selection with bias-corrected inference, rejects the global high-order polynomial with a reason, and finishes on the sensitivity plot as the thing that makes the estimate believable.
- Why is a confidence interval from an MSE-optimal bandwidth widened rather than used as is?Because the MSE-optimal choice deliberately leaves some bias in the estimate, so an interval centred on it under-covers the true effect. Bias-corrected inference subtracts an estimate of that bias and inflates the interval to account for the uncertainty in the correction, which is why the reported band is wider than the naive one.
- What does it mean if the estimate flips sign as the bandwidth widens?That the result is driven by functional form rather than by a discontinuity. Either the curve bends sharply near the threshold, so the linear fit is misspecified at larger windows, or the local sample is too thin to pin anything down. Either way the honest report is that the design does not support a confident estimate.
- Is a local quadratic ever preferable to a local linear fit?It can reduce bias when the relationship is visibly curved through the window, at the cost of variance and of more unstable boundary behaviour. Use it as a robustness check alongside the local linear headline rather than as the default, and expect the two to agree if the design is sound.
saying these in an interview costs you the question
- Fits a global high-order polynomial as the headline specification
- Picks the bandwidth that makes the result significant
- Says a wider bandwidth reduces bias
- Forces a single slope across both sides of the cutoff
- Reports one estimate with no bandwidth sensitivity