skip to content

When does synthetic control beat picking a single comparison group for a policy evaluation?

level: seniorimportance: nice to knowfreq 30%

answer

  1. one treated unit, many candidates
  2. weights fitted on the pre-period
  3. non-negative, summing to one
  4. no extrapolation beyond the donors
  5. permute the treatment across donors

basics

~20 s

Synthetic control helps when one treated unit has no credible twin. It builds the comparison as a weighted average of many untreated units, with weights chosen so the composite tracks the treated unit's pre-treatment path.

solid answer

~50 s

The classic setting is one treated unit and a long pre-period — one state, one country, one large market. No single untreated unit looks like it, and whichever one you pick, someone can argue you picked it to get the answer you wanted. Synthetic control replaces that judgment call with an optimisation: from a donor pool of untreated units, choose non-negative weights summing to one so the weighted composite matches the treated unit's pre-treatment outcome path and relevant predictors as closely as possible. The counterfactual after the intervention is that same weighted combination carried forward, and the estimated effect is the gap between the treated unit and its synthetic version. The canonical application is California's 1988 tobacco-tax initiative, where synthetic California is a weighted blend of a few other states. The constraints matter: non-negative weights summing to one keep the counterfactual inside the range the donors actually span, so you never extrapolate.

go deeper

for a junior

Know the shape of the idea: when one unit gets treated and no single comparison unit fits, the comparison is built as a weighted blend of several untreated units.

for a middle

Be ready to explain how the weights are chosen — fitted on pre-intervention outcomes and predictors — and why they are constrained to be non-negative and to sum to one.

for a senior

Show you treat pre-period fit as a go/no-go diagnostic, keep the donor pool free of treated and spillover-affected units, and use permutation across donors rather than a t-test for inference.

for a principal

Own when this design is worth its cost: it buys credibility for a single-market decision, but with many treated units a heterogeneity-robust difference-in-differences is simpler and easier to defend.

## The problem it solves Difference-in-differences with one treated unit and one hand-picked comparison unit has a credibility problem that no amount of statistics fixes: the choice of comparison drives the answer, and the choice was made by a person who could see the data. If California is the treated state, is the right comparison Nevada? Oregon? The national average? Each gives a different number, and each has an advocate. Synthetic control turns the choice into an estimated object. Rather than pick one donor, construct a weighted average of many and let the pre-treatment fit choose the weights. ## The construction You need: - **One treated unit**, observed for many periods before and after the intervention. - **A donor pool** of units that were never treated, are not affected by the treatment through spillover, and did not experience their own large idiosyncratic shocks during the window. - **A long, clean pre-period** — the longer it is, the more constraining the fit and the less likely a good match is coincidence. You then choose weights `w_j >= 0` for each donor `j`, with `sum(w_j) = 1`, to minimise the discrepancy between the treated unit and the weighted donors over the pre-treatment period, on the outcome and on a set of predictor variables. The synthetic unit's post-intervention path is the same weighted combination of donor outcomes, and the estimated effect in each post period is `effect_t = treated_t - sum over j of (w_j * donor_j,t)` ## Why the constraints **Non-negative weights** rule out combinations that subtract one donor from another. **Weights summing to one** keep the synthetic unit inside the convex hull of the donors — a genuine weighted average of real units, not an extrapolation beyond anything observed. Together these prevent the estimator from manufacturing a counterfactual that no combination of real places could produce, which is exactly the failure mode of an unconstrained regression fit to a short pre-period. They also tend to yield **sparse** weights: a handful of donors get non-trivial weight and the rest get zero, which makes the counterfactual describable in a sentence. ## The canonical case California passed a ballot initiative in 1988 raising its cigarette excise tax, effective the following year. No other state was a plausible stand-in for California on its own. The synthetic control approach built a synthetic California as a weighted average of a small number of other states chosen to reproduce California's pre-1989 per-capita cigarette sales and its predictors. The two series track closely through the pre-period, then diverge after the tax takes effect, and the widening gap is the estimated effect. That pre-period tracking is the whole argument. A synthetic unit that follows the treated unit for a decade and then separates precisely at the intervention date is far harder to dismiss than a single comparison state chosen after the fact. ## Diagnostics and inference **Pre-treatment fit is the first check.** If the synthetic unit cannot reproduce the treated unit before the intervention, the donor pool does not contain the ingredients, and the post-period gap is uninterpretable — you cannot tell an effect from a fit failure. Poor fit is a reason to stop, not to widen the intervals. **Inference is not a t-test.** With one treated unit you have no sampling distribution in the usual sense. The standard approach is permutation: reassign the intervention in turn to each donor, run the whole procedure, and collect the gaps. If the treated unit's post-intervention gap is extreme relative to the distribution of donor gaps — usually normalised by each unit's pre-period fit quality, so a unit that fitted badly does not look impressive by accident — the result is unlikely to be an artefact of the method. Reporting a conventional standard error on a single treated unit is a red flag. **Donor-pool hygiene.** Exclude units that were treated themselves, units plausibly affected by spillover from the treatment, and units that went through their own shock during the study window. A donor contaminated by the treatment biases the counterfactual toward the treated path and shrinks the estimated effect. ## When not to reach for it With many treated units, difference-in-differences and its heterogeneity-robust variants are simpler and give you standard machinery for inference. With a short pre-period, the weights are fitted on too little information and a good match means little. And when the treated unit is extreme on the outcome — the largest market by far — no convex combination of donors can reach it, and the method will tell you so through a poor fit rather than silently extrapolating. That honesty is a feature, but it does mean the design is unavailable in exactly the cases where the treated unit is most unusual.

  • Why are synthetic control weights constrained to be non-negative and to sum to one?
    So the counterfactual stays a genuine weighted average of real units rather than an extrapolation beyond anything observed. Unconstrained fits can produce a synthetic unit no combination of real places could ever look like. The constraints also tend to make the weights sparse, so the counterfactual can be described in one sentence.
  • What disqualifies a unit from the donor pool?
    Being treated itself, being plausibly affected by the treatment through spillover, or going through a large idiosyncratic shock during the study window. A contaminated donor drags the synthetic counterfactual toward the treated unit's own path, which shrinks the estimated effect toward zero.
  • How do you judge whether a single treated unit's gap is meaningful?
    By permutation: reassign the intervention to each donor in turn, rerun the whole procedure, and compare the treated unit's post-intervention gap with the distribution of donor gaps, normalised by each unit's pre-period fit quality. A conventional standard error on one treated unit has no basis and should not be reported.
  • What does a poor pre-treatment fit tell you?
    That the donor pool cannot reproduce the treated unit, so any post-intervention gap could just as easily be fit failure as effect. The response is to widen or improve the donor pool, or to conclude the design is unavailable — not to report the gap with wider intervals as though it were still an estimate.

saying these in an interview costs you the question

  • Fits the weights using post-intervention outcomes as well as pre
  • Leaves treated or spillover-affected units in the donor pool
  • Reports a conventional t-test on a single treated unit
  • Accepts a poor pre-period fit and interprets the gap anyway
  • Uses it with a short pre-period and calls a close match convincing

context