skip to content

What does a caliper do in propensity score matching, and how wide should it be?

level: middleimportance: should knowfreq 47%

answer

  1. nearest neighbour always finds someone
  2. a ceiling on match distance
  3. measured on the logit scale
  4. roughly 0.2 standard deviations
  5. unmatched treated units get dropped

basics

~20 s

A caliper caps how far apart a treated unit and its match may sit on the propensity score, commonly at 0.2 standard deviations of the score's logit. Treated units with no control inside that distance are left unmatched.

solid answer

~50 s

Plain nearest-neighbour matching always returns somebody, however unsuitable: with a thin control pool the nearest control can have a wildly different score. A caliper sets a maximum allowed distance, and a treated unit whose closest control lies outside it is simply left unmatched. The width is normally stated on the logit scale, with `0.2` standard deviations of the logit of the propensity score the common recommendation and roughly `0.1` to `0.25` the defensible range. The logit is used because estimated scores bunch up near 0 and 1, where a gap of 0.02 separates very different units, while the same gap mid-range is trivial. Tightening the caliper cuts bias from bad pairings but drops more treated units, which raises variance and quietly narrows the population the estimate describes, so always report how many treated units went unmatched and how they differ.

go deeper

for a junior

Know what the term means: a caliper is a limit on how far apart two units' propensity scores may be before the pair is refused, and it can leave some treated units with no match at all.

for a middle

Explain the mechanics, including that the width is expressed in standard deviations of the logit of the score, around 0.2 by convention, and that tightening it trades fewer matched units for closer ones.

for a senior

Show that you check the consequences: count and profile the unmatched treated units, and re-run at a couple of caliper widths to see whether the conclusion depends on a tuning choice.

for a principal

Own the framing that the caliper decides which population the study can speak about, and insist that write-ups name that population rather than presenting a trimmed result as the effect on all treated units.

## The failure mode a caliper prevents Nearest-neighbour matching is defined as: for each treated unit, take the control whose propensity score is closest. Note what that definition does not say. It does not say the control has to be *close*. If a treated unit has a score of 0.92 and the nearest control in the entire pool sits at 0.41, nearest-neighbour matching cheerfully forms that pair, and downstream the pair is treated as if it were a valid like-for-like comparison. Nothing in the output flags it. The imbalance shows up, diluted, in the balance table, and only if you happen to look at the distribution of matched distances do you see what happened. A caliper closes that hole. It is a maximum distance: if the best available control is further away than the caliper, no match is made and the treated unit drops out of the analysis. ## Why the width is stated on the logit scale Estimated propensity scores are probabilities, and they pile up at the ends of the range. Near 0.5, moving from 0.50 to 0.54 changes very little about what kind of unit you are looking at. Near the top, moving from 0.95 to 0.99 is an enormous change: the odds of treatment go from about 19 to about 99. A single distance measured on the raw probability scale therefore means different things in different parts of the range, and a caliper set that way is effectively far too permissive in the tails, which is exactly where matching is most fragile. The logit transform, `log(e / (1 - e))`, stretches the tails out and compresses the middle, giving a scale on which a fixed distance carries comparable meaning everywhere. The standard recipe is: 1. Transform every unit's estimated score to its logit. 2. Compute the standard deviation of those logits. 3. Set the caliper to 0.2 of that standard deviation. The 0.2 figure comes from simulation work comparing caliper widths on bias and mean squared error, and it is a recommendation rather than a law; values from roughly 0.1 to 0.25 standard deviations are all defensible, with tighter values preferred when the control pool is rich enough to afford them. ## The tradeoff the width controls The caliper is a bias–variance dial, and being able to say which way it turns is the point of the question. **Tighter caliper.** Only very close pairs survive, so bias from mismatched pairs falls. But more treated units fail to find any partner, the matched sample shrinks, and the estimate becomes noisier. Taken to the extreme, a caliper of nearly zero leaves you with a handful of near-exact pairs and an interval too wide to act on. **Wider caliper.** More treated units keep a partner, so the matched sample is bigger and the estimate more precise. But the pairs are worse, residual imbalance survives into the analysis, and the bias it causes does not shrink with sample size. A precise, confidently wrong number is the worst outcome available. ## The consequence people forget: the population changes Every treated unit the caliper discards is a unit the study no longer describes. If 30 of 200 treated units go unmatched, the estimate is an effect among the 170 treated units that had a comparable control, and those 170 are systematically different from the 30 — typically less extreme, closer to the middle of the score distribution. The number is still meaningful, but it is a number about a narrower population, and the write-up has to say which one. So the reporting obligation that comes with a caliper is: how many treated units were dropped, what fraction of the treated group that is, and how the dropped units differ from the retained ones on the covariates that matter. A single sentence of the form "23 of 200 treated accounts had no control within the caliper; they were larger and had longer tenure than the matched accounts" is worth more than any amount of methodological hedging. ## Sensitivity to the choice Because 0.2 is a convention, an interviewer will be pleased if you volunteer that you re-run the analysis at two or three caliper widths and look at whether the conclusion moves. If the effect is stable across 0.1, 0.2 and 0.25 while the matched sample size changes materially, that is real evidence of robustness. If the sign flips between widths, you have learned that the result is an artefact of a tuning choice, which is far better to learn before publication than after. ## What not to do The tempting response to a large number of discarded treated units is to widen the caliper until everyone keeps a partner. That does not recover the missing information; it manufactures pairs that should never have been formed and hides an overlap problem inside them. Discarding is at least visible and countable. Forced matching converts a limitation you could have reported into a bias nobody can see.

  • Why is the caliper usually set on the logit of the score rather than the score itself?
    Estimated scores pile up near 0 and 1, where a gap of 0.02 can separate very different units, while mid-range the same gap is trivial. The logit stretches the tails and compresses the middle, so one fixed distance carries comparable meaning across the whole range and the caliper is not effectively looser exactly where matching is most fragile.
  • What is the cost of tightening the caliper from 0.25 to 0.05 standard deviations?
    Fewer, closer pairs. Bias from mismatched pairs falls, but more treated units fail to find a partner, so the matched sample shrinks and the estimate gets noisier. It also shifts the population under study: the effect now describes only treated units dense enough in score to keep a near neighbour, which may not be the group the decision concerns.
  • Should you drop the caliper entirely so that every treated unit keeps a match?
    Rarely. Forcing a match for a treated unit with no comparable control creates no information; it buries an overlap problem inside a pair that should never have existed, and the resulting bias is invisible in the output. Leaving the unit unmatched is at least countable, so you can report who was lost and how they differed.

saying these in an interview costs you the question

  • Assumes nearest-neighbour matching guarantees a comparable control
  • Sets the caliper on the raw probability scale without thinking
  • Reports the effect without saying how many treated units were dropped
  • Treats 0.2 standard deviations as a law rather than a convention
  • Widens the caliper until the matched sample size looks respectable

context