skip to content

How do you check whether users manipulated the running variable around a regression discontinuity cutoff?

level: seniorimportance: must knowfreq 55%

answer

  1. count units on each side first
  2. a pile-up just above the bar
  3. test the score's distribution, not the outcome
  4. pre-treatment covariates should not jump
  5. donut, bounds, or abandon the design

basics

~20 s

The McCrary test checks whether the density of the running variable jumps at the cutoff. Bunching just above the threshold means units steered themselves across, so the two sides are no longer comparable and the design fails.

solid answer

~50 s

Start with a histogram of the running variable. Take a loyalty tier that unlocks at 1,000 points: if there is a visible pile-up just above 1,000 and a hollow just below, users are topping up to clear the bar, so the group above is self-selected — more motivated, higher-intent — and the design's continuity assumption is gone. The formal version is the **McCrary density test**, which estimates the density of the running variable separately on each side of the cutoff and tests whether it is continuous at the threshold; a significant jump is evidence of sorting. Back it with two more checks: pre-determined covariates should not jump at the cutoff, and placebo estimates at fake thresholds should find nothing. If manipulation is confined to a narrow band, a donut specification that drops observations immediately around the cutoff can salvage the design, at the cost of extrapolating to the threshold.

go deeper

for a junior

Know that the first plot to draw is a histogram of the score itself, and that a pile-up just above the threshold is bad news for the design rather than an interesting finding.

for a middle

Explain what the density test actually tests — a discontinuity in the distribution of the running variable at the cutoff — and why sorting makes the two sides non-comparable even when the sample is large.

for a senior

Bring the full diagnostic battery and the institutional reasoning: who sees the score, who can move it, what covariates say, and what you would do — donut, bounds, or walk away — when the checks fail.

for a principal

Own the standard of evidence: decide in advance which checks a cutoff analysis must pass before it can inform a decision, so that a failed density test stops a launch claim rather than becoming a footnote.

## Why manipulation breaks the design A regression discontinuity works because units just below and just above the threshold are otherwise alike. That claim dies the moment units can choose their side. If clearing 1,000 points unlocks a loyalty tier and users can see their balance, the ones who buy a cheap extra item to cross the line are not a random slice of the near-threshold population — they are the ones who wanted the reward most. Compare them with users sitting at 990 and the jump you measure mixes the tier's effect with the difference between people who pushed and people who did not. The technical statement is that continuity of average potential outcomes at the cutoff fails: something other than treatment changes discretely at 1,000, namely the composition of the population. ## The density check Manipulation leaves a fingerprint in the distribution of the running variable itself. If units move from just below to just above, the count on the low side is depleted and the count on the high side is inflated. Under no manipulation, the density of the running variable should pass smoothly through the cutoff — there is no reason for the number of users with 999 points to differ sharply from the number with 1,001. **The McCrary density test** formalises this. It estimates the density of the running variable on each side of the threshold using a local fit, and tests the null hypothesis that the density is continuous at the cutoff against the alternative that it jumps. Rejecting the null is evidence of sorting. Two honest caveats belong in any answer: - **Passing the test does not prove there was no manipulation.** If as many units are pushed up as are pulled down, or if manipulation is spread across a wide band, the density can look smooth while composition still changes. - **A density jump is not always manipulation.** Rounding, a data pipeline that snaps values to the threshold, or a second rule that also fires at 1,000 can produce a jump. Investigate the mechanism before concluding. Always look at the raw histogram too. Heaping at round numbers is common, and a spike exactly *at* the cutoff has a different story from a smooth excess just above it. ## Supporting checks **Covariate smoothness.** Variables determined before the running variable was realised — tenure at signup, acquisition channel, region — should be continuous at the threshold. If users just above 1,000 points were disproportionately acquired through a promotional channel, they differ in ways the design cannot absorb. Plot each covariate the same way you plot the outcome and run the same local estimate on it; a jump is a red flag, and one jump among many covariates tested is a multiple-comparisons matter to be honest about rather than to explain away. **Placebo cutoffs.** Re-run the whole estimate at thresholds where no rule exists — 800 points, 1,200 points. Finding "significant" jumps at several placebo values means the specification is manufacturing discontinuities out of curvature or noise, and the headline estimate is not evidence of anything. **Institutional reasoning.** The strongest evidence is often not statistical. Ask who knows the cutoff, when they learn their score, and whether they have the means and the incentive to change it. Age cannot be manipulated. A points balance shown in the app with a progress bar toward the next tier can be manipulated, and probably is. A risk score computed nightly from data the user never sees is somewhere in between. ## What to do when manipulation is present - **Donut RD.** Drop observations within a small band on either side of the cutoff and estimate the jump from the remaining data, extrapolating the fits to the threshold. This removes the units most likely to have sorted, at the price of leaning on functional form over the gap. Report the donut radius and show that the estimate is stable across a few radii. - **Bound rather than point-estimate.** If the direction of sorting is known — only upward crossing is plausible — the effect can be bounded rather than estimated, which is more honest than a point estimate the design cannot support. - **Abandon the design.** Sometimes the right answer is that this cutoff cannot identify anything, and the question should be answered another way. Saying that in an interview is a strength, not a failure. ## How this is asked Expect a scenario: "we give a perk at a visible threshold, can we use the cutoff to measure the perk's effect?" The complete answer names the manipulation risk first, proposes the density check and the covariate check, distinguishes a rule users can see from one they cannot, and only then talks about estimation. Candidates who jump straight to the regression have missed the part that decides whether the number means anything.

  • The density test passes. Can you conclude there was no sorting?
    No. A smooth density is reassuring but not conclusive: manipulation that moves similar numbers in both directions, or that is spread over a wide band, leaves the density looking continuous while composition still shifts. Pair it with covariate smoothness and with a plain argument about who could manipulate the score and why.
  • What is a donut RD and what does it cost you?
    You drop observations in a narrow band on both sides of the cutoff and estimate the jump from what remains, extrapolating each fit to the threshold. It removes the units most likely to have sorted, but the estimate now rests on functional form across the gap, so report several donut radii and show the result does not depend on the choice.
  • Which rules are least vulnerable to manipulation?
    Ones where the unit cannot observe or move the score before it binds: a birth date, a score computed from historical data the unit never sees, or a threshold announced only after assignment. Visible, real-time, incrementable scores with a progress indicator are the most vulnerable.
  • You find a covariate that jumps at the cutoff. Does adding it as a control fix the design?
    No. A covariate jump is a symptom that the populations on the two sides differ, which means continuity has failed for reasons you cannot enumerate. Controlling for the one covariate you happened to measure leaves the unmeasured differences in place. Diagnose the mechanism instead.

It is the queue at a height bar where children stand on tiptoe: the ones who just clear it are not the same children as the ones who just missed, no matter how close the measurements look.

saying these in an interview costs you the question

  • Only checks the outcome plot, never the score's distribution
  • Treats a passing density test as proof of no manipulation
  • Controls for a jumping covariate instead of questioning the design
  • Assumes any threshold rule is safe from sorting
  • Never asks whether units can see their own score

context