skip to content

What is a propensity score, and why match on it rather than on the raw covariates?

level: juniorimportance: must knowfreq 76%

answer

  1. one scalar standing in for many covariates
  2. exact matching dies in high dimensions
  3. Rosenbaum-Rubin balancing property
  4. modelled from treatment, never from outcome
  5. probability of being treated given covariates

basics

~20 s

A propensity score is a unit's probability of receiving the treatment given its observed covariates. Conditioning on that single number balances the covariates that went into it, so matching happens in one dimension instead of many.

solid answer

~50 s

The propensity score is `e(x) = P(T = 1 | X = x)`: the probability that a unit with covariate vector `x` ends up treated. Rosenbaum and Rubin showed it is a balancing score, meaning that among units sharing the same score, the covariates that entered the score have the same distribution in the treated and untreated groups. That is what makes it useful. Exact matching on twenty covariates is hopeless, because almost every treated unit sits in a covariate cell with no untreated twin, but matching on one scalar is easy. And if treatment is unconfounded given the full covariate vector, it is also unconfounded given the score. Two caveats: the theorem is about the true score and you only ever have an estimated one, so balance has to be checked empirically; and the score can only balance covariates you measured and included.

go deeper

for a junior

Be ready to state the definition in one line, probability of treatment given the measured covariates, and to say why one number is easier to match on than twenty separate covariates.

for a middle

An interviewer expects the balancing-score argument: conditional on the score, the covariates that entered it have the same distribution in both groups, which is why one dimension can stand in for many.

for a senior

Show that you treat the score as an estimated nuisance quantity. You verify balance on the matched sample rather than trusting the theorem, and you say plainly that the score cannot fix anything you failed to measure.

for a principal

Own the framing that matching is a design decision made before any outcome is looked at, and be ready to say when observational adjustment simply cannot answer the question and an experiment is the honest recommendation.

## The problem the score solves In an observational study nobody randomised anything. Units that received the treatment differ systematically from units that did not: they are older, larger, more engaged, further along, or simply more likely to have been offered it. A raw comparison of outcomes therefore mixes the effect of the treatment with the effect of those differences. The obvious repair is to compare like with like — pair each treated unit with an untreated unit that has the same covariate values. This is exact matching, and it works beautifully with two or three covariates. It collapses as soon as you have twenty. The number of distinct covariate cells grows multiplicatively with the number of covariates and their levels, so almost every treated unit ends up alone in its cell with no untreated partner. This is the curse of dimensionality applied to matching. ## The definition The propensity score of a unit is e(x) = P(T = 1 | X = x) the probability that a unit whose observed pre-treatment covariates equal `x` receives the treatment. Three things follow directly from that definition and are worth saying out loud in an interview: - It is a function of the covariates only. The outcome never appears in it, and neither does anything measured after treatment. - It is a probability between 0 and 1. In a fair randomised experiment `e(x) = 0.5` for everyone regardless of `x`; the score varies only because assignment was not random. - It is a nuisance quantity. Nobody wants to know the score. It exists to make the two groups comparable. ## The balancing property Rosenbaum and Rubin (1983) proved the result that makes the whole method work: the propensity score is a *balancing score*. Formally, the observed covariates are independent of treatment given the score, X ⊥ T | e(X) In words: pick any group of units that share the same propensity score, and within that group the distribution of the covariates is the same for the treated and the untreated. Whatever combination of age, tenure, size and history pushed a unit's score to 0.3, treated and untreated units sitting at 0.3 look alike on those covariates on average. That is a strong and slightly surprising claim, and it is the reason a one-dimensional match can reproduce what a twenty-dimensional match would have given. You are not throwing information away by collapsing the covariates into a scalar; for the purpose of balancing them, the scalar is sufficient. The second Rosenbaum–Rubin result carries this to causal identification. If treatment is unconfounded given the full covariate vector — potential outcomes are independent of treatment once you condition on `X` — then treatment is also unconfounded given `e(X)` alone. So under unconfoundedness plus positivity (every unit has a score strictly between 0 and 1, so both treatment states are possible for everyone), comparing outcomes between matched units identifies a causal effect. ## Where the theory stops and the work starts Three limits matter, and candidates who name them unprompted stand out. **The score is estimated.** The balancing theorem is about the true score, which nobody knows. You fit a model, get fitted values, and match on those. Nothing guarantees the estimated score balances anything, which is exactly why the matched sample must be checked covariate by covariate afterwards. Balance is an empirical claim about your data, not a theorem you get to cite. **Only measured, included covariates are balanced.** A covariate that was never recorded, or was recorded and left out of the score, is not balanced by any of this. Matching on observables handles confounding by observables and nothing more. The long-running LaLonde job-training literature is the standing illustration: non-experimental comparison pools drawn from national survey data produced estimates far from the experimental benchmark, and the Dehejia–Wahba reanalysis only approached it once rich pre-treatment earnings histories were included — with later work showing the result was sensitive to which sample and specification were chosen. **A highly predictive score is bad news, not good news.** People trained on prediction problems assume that a score which separates treated from untreated cleanly is a well-built score. It is the opposite. If most units have scores near 0 or near 1, the two populations barely overlap, few treated units have any plausible partner, and matching will either drop them or pair them with someone very different. The score's job is comparability, not accuracy. ## Common confusions The score is not the probability of a good outcome; it models assignment, not response. It is not a similarity metric between a unit and the treated group as a whole — two units can share a score for entirely different covariate reasons, which is why balance is checked in aggregate rather than pair by pair. And it must be built from pre-treatment variables only: anything measured after treatment can itself be affected by the treatment, so conditioning on it removes part of the very effect you are trying to estimate.

  • Does matching on the propensity score deal with unmeasured confounding?
    No. The balancing property covers only covariates that entered the score, so anything you never measured is free to differ between matched groups. That is the standing critique of the LaLonde job-training evaluations: non-experimental comparison pools gave estimates far from the experimental benchmark, and the Dehejia-Wahba reanalysis only came close once rich pre-treatment earnings histories were included, and even then the result proved sensitive to the sample and specification chosen.
  • Is a propensity score that predicts treatment almost perfectly a good sign?
    No, it is close to fatal. Scores near 0 or 1 for most units mean the treated and untreated populations barely overlap, so few treated units have a plausible partner and matching either discards them or pairs them with someone very different. The score is a nuisance quantity whose job is comparable groups, not accurate prediction.
  • Should anything measured after treatment go into the propensity score?
    No. The score is a function of pre-treatment covariates only. A variable measured after treatment can itself have been changed by the treatment, so conditioning on it strips out part of the effect you are trying to estimate. Including the outcome is worse still, because it turns the design stage into outcome-driven fishing.

It is like sorting job applicants by the single probability that they would have been shortlisted, then comparing hired and non-hired people who had the same shortlisting odds, instead of hunting for someone with an identical CV.

saying these in an interview costs you the question

  • Calls the propensity score the probability of a good outcome
  • Treats high predictive accuracy of the score as the goal
  • Claims matching on the score removes unmeasured confounding
  • Puts post-treatment variables into the score
  • Assumes the balancing theorem guarantees balance in the actual sample

context