skip to content

Crashes fell 35% at junctions where speed cameras were installed after a record-bad year — how much credit do the cameras deserve?

level: seniorimportance: should knowfreq 52%

answer

  1. the selection rule is the problem
  2. a record-bad year is partly a bad-luck year
  3. compare with sites picked the same extreme way
  4. baseline from several pre-years, not the peak

basics

~20 s

Some of the drop is regression to the mean: the junctions were chosen for having an extreme year. Estimate the camera effect against untreated junctions selected by the same rule, not against the treated sites' own worst year.

solid answer

~50 s

The naive before-and-after number is biased upward because the sites were selected on an extreme. Annual crash counts at a single junction are small and noisy, so a 'record-bad year' list is systematically enriched with junctions that had an unlucky year on top of their real risk; the unlucky part does not repeat, and counts fall the following year with or without a camera. To separate the two, compare against junctions that had a similarly extreme year but received no camera, so both groups regress by the same amount. Failing that, build the baseline from several pre-treatment years rather than the selection year alone, or shrink the site's observed count toward the expected count for similar junctions — the empirical-Bayes approach standard in road-safety evaluation. Then check that the drop persists over several years, since regression produces a one-step return, not a trend.

go deeper

for a junior

Recognise that picking sites for their worst-ever year means some of the later drop was coming anyway, and say so. Naming the phenomenon and refusing to accept the raw 35% is already the right instinct.

for a middle

Explain why small annual counts make the selection so damaging, and describe the two practical baselines: a multi-year pre-period average, and a comparison group that met the same extreme criterion.

for a senior

Show you can build the estimate: choose the comparison group deliberately, explain why a random one makes things worse, and name persistence and mechanism specificity as the checks that separate artifact from effect.

for a principal

Take the decision angle. Decide what standard of evidence a safety programme must meet before it is scaled, and whether to stagger rollout so a defensible counterfactual exists without abandoning the targeting.

## Why the headline number cannot be trusted The estimate '35% fewer crashes after installation' compares each treated junction to itself in the year that got it selected. That year was chosen precisely because it was the worst on record. Crash counts at a single junction are small integers driven partly by underlying risk and partly by chance — which drivers passed, in what weather, at what moment. A ranking of junctions by last year's count therefore mixes 'genuinely dangerous' with 'ordinary junction that had a terrible year'. The second group's counts fall back toward their own long-run level next year with no intervention at all. This is regression to the mean, and in blackspot evaluations it has historically accounted for a large share of apparent treatment effects. The size of the problem scales with how noisy the counts are. A junction averaging 3 crashes a year has enormous relative year-to-year variation; a corridor averaging 300 has very little. Selection on an extreme bites hardest exactly where the metric is thinnest, which is also where blackspot lists tend to come from. ## What an honest estimate needs **A comparison selected the same way.** The single best fix is a comparison group of junctions that also had a record-bad year and did not get a camera — because a road authority could only fund so many, or because the eligibility threshold happened to fall between them. Both groups were selected on the same extreme, so both regress by the same amount; the difference between them is attributable to the camera. Note the subtlety: a comparison group of *randomly chosen* junctions is worse than useless here, because it did not undergo the extreme selection and so does not regress at all, leaving the entire bias in place. **A better baseline.** If no comparison group exists, stop using the selection year as the baseline. Average several pre-treatment years instead. The selection latched onto one year's fluctuation; a multi-year average estimates the junction's underlying rate with much less of that particular fluctuation in it. The bias does not vanish — the selection year is usually part of the average and the selection still favoured high-noise sites — but it shrinks substantially. **An explicitly modelled expectation.** The rigorous version predicts, from the whole network's history, the expected crash count for a junction with the treated one's traffic volume, layout and speed limit, then combines that expectation with the junction's own observed count, weighting the observed count by how reliable it is. The result is a shrunken estimate of the site's true rate — the empirical-Bayes baseline used in road-safety practice. It prices in directly how much of the record-bad year was noise. In standard-deviation terms it is the same rule as multiplying an observed deviation by the period-to-period correlation. **Randomisation where it is available.** If more junctions qualify than can be treated at once, allocating cameras at random among the qualifying set, or staggering rollout so a comparably extreme slice stays untreated for a period, gives a clean counterfactual while still targeting the sites that need it. ## Reading the evidence afterwards Several patterns help distinguish artifact from effect. - **Persistence.** Regression to the mean is a one-step return to the site's own long-run level. If crashes stay down for several years rather than rebounding partway or drifting, that is harder to explain as selection noise. - **Specificity.** A camera plausibly reduces speed-related crashes at and just before the camera. A drop spread evenly across every crash type, including those with no speed mechanism, looks more like a general regression. - **A dose relationship with the mechanism.** Measured speed reductions at the treated sites, correlating with the crash reduction, support a causal reading. - **The wrong dose relationship.** If the drop is largest at the junctions whose selection year was *most* extreme, that is the signature of the artifact rather than of the treatment. ## How to answer this in an interview Say first that you cannot answer '35% minus how much' without knowing the selection rule and the metric's year-to-year stability, and that the honest first move is to quantify the expected rebound rather than to argue about it. Then give the hierarchy: comparably-selected untreated sites if you can get them, a multi-year or model-based baseline if you cannot, and persistence plus mechanism checks either way. What interviewers are listening for is whether you spot the selection rule unprompted, whether you know that a random comparison group makes the bias worse rather than better, and whether you can say what the number would be under a pure-artifact null. Candidates who accept the 35% or who dismiss the whole evaluation as worthless both miss: the effect is estimable, just not by subtraction from the worst year on record.

  • Why does using several pre-treatment years as the baseline reduce the bias?
    The selection latched onto one year's fluctuation. Averaging several years estimates the junction's underlying rate with much less of that specific fluctuation in it, so the comparison baseline is closer to what the site would have done anyway. It does not remove the bias completely, since the selection year is usually inside the average and selection still favoured high-noise sites, but it shrinks it a lot.
  • Why is a randomly chosen comparison group worse than no comparison group here?
    Because it did not go through the extreme selection, so it does not regress. Treated sites fall for two reasons — the camera and the rebound — while random comparison sites fall for neither. The difference therefore keeps the full regression bias and dresses it up as a controlled estimate, which is more persuasive and equally wrong.
  • What evidence would convince you the drop was mostly a real effect?
    A reduction sustained across several later years rather than only the first, concentrated in the crash types a camera plausibly affects, matched by measured speed reductions at the treated sites, and absent at similarly extreme untreated sites over the same period. Regression predicts a single step back toward the site's long-run level, so persistence and mechanism are what separate the two.

saying these in an interview costs you the question

  • Compares only before and after at the treated sites
  • Draws comparison sites at random from the whole network
  • Treats the record-bad year as a normal baseline
  • Argues the drop must be real because it was large
  • Dismisses the evaluation as impossible rather than estimating the rebound

context