skip to content

Adjustment Estimators

Making treated and untreated groups comparable after the fact by pairing units, reweighting them or modelling the outcome, then checking balance. Interviewers ask what balance you got.

on this pageshow

explore

questions

15

What is a propensity score, and why match on it rather than on the raw covariates?

level: juniorimportance: must knowfreq 76%

answer

  1. one scalar standing in for many covariates
  2. exact matching dies in high dimensions
  3. Rosenbaum-Rubin balancing property
  4. modelled from treatment, never from outcome
  5. probability of being treated given covariates

basics

~20 s

A propensity score is a unit's probability of receiving the treatment given its observed covariates. Conditioning on that single number balances the covariates that went into it, so matching happens in one dimension instead of many.

solid answer

~50 s

The propensity score is `e(x) = P(T = 1 | X = x)`: the probability that a unit with covariate vector `x` ends up treated. Rosenbaum and Rubin showed it is a balancing score, meaning that among units sharing the same score, the covariates that entered the score have the same distribution in the treated and untreated groups. That is what makes it useful. Exact matching on twenty covariates is hopeless, because almost every treated unit sits in a covariate cell with no untreated twin, but matching on one scalar is easy. And if treatment is unconfounded given the full covariate vector, it is also unconfounded given the score. Two caveats: the theorem is about the true score and you only ever have an estimated one, so balance has to be checked empirically; and the score can only balance covariates you measured and included.

go deeper

for a junior

Be ready to state the definition in one line, probability of treatment given the measured covariates, and to say why one number is easier to match on than twenty separate covariates.

for a middle

An interviewer expects the balancing-score argument: conditional on the score, the covariates that entered it have the same distribution in both groups, which is why one dimension can stand in for many.

for a senior

Show that you treat the score as an estimated nuisance quantity. You verify balance on the matched sample rather than trusting the theorem, and you say plainly that the score cannot fix anything you failed to measure.

for a principal

Own the framing that matching is a design decision made before any outcome is looked at, and be ready to say when observational adjustment simply cannot answer the question and an experiment is the honest recommendation.

## The problem the score solves In an observational study nobody randomised anything. Units that received the treatment differ systematically from units that did not: they are older, larger, more engaged, further along, or simply more likely to have been offered it. A raw comparison of outcomes therefore mixes the effect of the treatment with the effect of those differences. The obvious repair is to compare like with like — pair each treated unit with an untreated unit that has the same covariate values. This is exact matching, and it works beautifully with two or three covariates. It collapses as soon as you have twenty. The number of distinct covariate cells grows multiplicatively with the number of covariates and their levels, so almost every treated unit ends up alone in its cell with no untreated partner. This is the curse of dimensionality applied to matching. ## The definition The propensity score of a unit is e(x) = P(T = 1 | X = x) the probability that a unit whose observed pre-treatment covariates equal `x` receives the treatment. Three things follow directly from that definition and are worth saying out loud in an interview: - It is a function of the covariates only. The outcome never appears in it, and neither does anything measured after treatment. - It is a probability between 0 and 1. In a fair randomised experiment `e(x) = 0.5` for everyone regardless of `x`; the score varies only because assignment was not random. - It is a nuisance quantity. Nobody wants to know the score. It exists to make the two groups comparable. ## The balancing property Rosenbaum and Rubin (1983) proved the result that makes the whole method work: the propensity score is a *balancing score*. Formally, the observed covariates are independent of treatment given the score, X ⊥ T | e(X) In words: pick any group of units that share the same propensity score, and within that group the distribution of the covariates is the same for the treated and the untreated. Whatever combination of age, tenure, size and history pushed a unit's score to 0.3, treated and untreated units sitting at 0.3 look alike on those covariates on average. That is a strong and slightly surprising claim, and it is the reason a one-dimensional match can reproduce what a twenty-dimensional match would have given. You are not throwing information away by collapsing the covariates into a scalar; for the purpose of balancing them, the scalar is sufficient. The second Rosenbaum–Rubin result carries this to causal identification. If treatment is unconfounded given the full covariate vector — potential outcomes are independent of treatment once you condition on `X` — then treatment is also unconfounded given `e(X)` alone. So under unconfoundedness plus positivity (every unit has a score strictly between 0 and 1, so both treatment states are possible for everyone), comparing outcomes between matched units identifies a causal effect. ## Where the theory stops and the work starts Three limits matter, and candidates who name them unprompted stand out. **The score is estimated.** The balancing theorem is about the true score, which nobody knows. You fit a model, get fitted values, and match on those. Nothing guarantees the estimated score balances anything, which is exactly why the matched sample must be checked covariate by covariate afterwards. Balance is an empirical claim about your data, not a theorem you get to cite. **Only measured, included covariates are balanced.** A covariate that was never recorded, or was recorded and left out of the score, is not balanced by any of this. Matching on observables handles confounding by observables and nothing more. The long-running LaLonde job-training literature is the standing illustration: non-experimental comparison pools drawn from national survey data produced estimates far from the experimental benchmark, and the Dehejia–Wahba reanalysis only approached it once rich pre-treatment earnings histories were included — with later work showing the result was sensitive to which sample and specification were chosen. **A highly predictive score is bad news, not good news.** People trained on prediction problems assume that a score which separates treated from untreated cleanly is a well-built score. It is the opposite. If most units have scores near 0 or near 1, the two populations barely overlap, few treated units have any plausible partner, and matching will either drop them or pair them with someone very different. The score's job is comparability, not accuracy. ## Common confusions The score is not the probability of a good outcome; it models assignment, not response. It is not a similarity metric between a unit and the treated group as a whole — two units can share a score for entirely different covariate reasons, which is why balance is checked in aggregate rather than pair by pair. And it must be built from pre-treatment variables only: anything measured after treatment can itself be affected by the treatment, so conditioning on it removes part of the very effect you are trying to estimate.

  • Does matching on the propensity score deal with unmeasured confounding?
    No. The balancing property covers only covariates that entered the score, so anything you never measured is free to differ between matched groups. That is the standing critique of the LaLonde job-training evaluations: non-experimental comparison pools gave estimates far from the experimental benchmark, and the Dehejia-Wahba reanalysis only came close once rich pre-treatment earnings histories were included, and even then the result proved sensitive to the sample and specification chosen.
  • Is a propensity score that predicts treatment almost perfectly a good sign?
    No, it is close to fatal. Scores near 0 or 1 for most units mean the treated and untreated populations barely overlap, so few treated units have a plausible partner and matching either discards them or pairs them with someone very different. The score is a nuisance quantity whose job is comparable groups, not accurate prediction.
  • Should anything measured after treatment go into the propensity score?
    No. The score is a function of pre-treatment covariates only. A variable measured after treatment can itself have been changed by the treatment, so conditioning on it strips out part of the effect you are trying to estimate. Including the outcome is worse still, because it turns the design stage into outcome-driven fishing.

It is like sorting job applicants by the single probability that they would have been shortlisted, then comparing hired and non-hired people who had the same shortlisting odds, instead of hunting for someone with an identical CV.

saying these in an interview costs you the question

  • Calls the propensity score the probability of a good outcome
  • Treats high predictive accuracy of the score as the goal
  • Claims matching on the score removes unmeasured confounding
  • Puts post-treatment variables into the score
  • Assumes the balancing theorem guarantees balance in the actual sample

context

open as a page

What does adding a confounder to a regression do to the treatment coefficient?

level: juniorimportance: must knowfreq 76%

basics

~20 s

Adding a measured confounder turns the treatment coefficient from a raw comparison into a within-strata one: it now compares treated and untreated units that share the same covariate value, removing the part of the gap that covariate explained.

open as a page

How does inverse probability of treatment weighting estimate an average treatment effect?

level: middleimportance: must knowfreq 68%

basics

~20 s

Each unit is weighted by the inverse of its probability of receiving the arm it actually got: treated units by 1/e(X), controls by 1/(1-e(X)). The reweighted sample mimics a population where treatment was assigned independently of the measured covariates.

open as a page

How do you check covariate balance after propensity score matching?

level: middleimportance: must knowfreq 82%

basics

~20 s

Compare standardized mean differences for each covariate before and after matching, with an absolute value under about 0.1 as the usual bar, and read them off a love plot. Check variance ratios and interactions too, not just means.

open as a page

How do you sign the omitted-variable bias when ability is left out of a wage-on-schooling regression?

level: middleimportance: must knowfreq 68%

basics

~20 s

Multiply two signs: the effect of the omitted variable on the outcome, times its correlation with the included regressor. Ability raises wages and is positively correlated with schooling, so the schooling coefficient is biased upward.

open as a page

In an IPW analysis one user has propensity score 0.01 and weight 100 — what do you do?

level: seniorimportance: must knowfreq 56%

basics

~20 s

Treat it as a diagnostic, not a nuisance. Check the largest weights and the effective sample size, see how far the estimate moves without that unit, then decide between trimming extreme-score units and capping weights.

open as a page

What do stabilized inverse probability weights change compared with plain 1/e(X) weights?

level: middleimportance: should knowfreq 42%

basics

~20 s

Stabilized weights put the marginal treatment probability in the numerator: P(T=1)/e(X) for treated units and P(T=0)/(1-e(X)) for controls. Weights then average about 1 and sum to roughly the sample size, which lowers variance without changing the estimand.

open as a page

What does a caliper do in propensity score matching, and how wide should it be?

level: middleimportance: should knowfreq 47%

basics

~20 s

A caliper caps how far apart a treated unit and its match may sit on the propensity score, commonly at 0.2 standard deviations of the score's logit. Treated units with no control inside that distance are left unmatched.

open as a page

How does standardisation, the g-formula, turn a fitted outcome model into an average treatment effect?

level: middleimportance: should knowfreq 44%

basics

~20 s

Fit one outcome model on treatment and covariates, then predict every unit twice, once with treatment set to 1 and once set to 0, and average the difference of those two predictions over the whole sample.

open as a page

What does an augmented IPW (doubly robust) estimator add over plain IPW?

level: seniorimportance: should knowfreq 38%

basics

~20 s

It combines a treatment model with an outcome model and stays consistent if either one is correct, instead of betting everything on the propensity model. It also reaches the best achievable precision when both models are right.

open as a page

In propensity score matching, when should you match with replacement rather than without?

level: seniorimportance: should knowfreq 44%

basics

~20 s

Match with replacement when the control pool is thin or overlaps poorly, because every treated unit then gets its closest available control instead of the leftovers. The price is that heavily reused controls shrink the effective sample size and widen uncertainty.

open as a page

In an adjusted regression table, why is only the exposure coefficient causally interpretable?

level: seniorimportance: should knowfreq 38%

basics

~20 s

The adjustment set was chosen to identify one exposure's effect. Each covariate has its own confounders, so its coefficient generally mixes a partial effect with leftover confounding. Reading every row as an effect is the Table 2 fallacy.

open as a page

Why can a logistic model's conditional odds ratio differ from the marginal odds ratio?

level: seniorimportance: nice to knowfreq 26%

basics

~20 s

The odds ratio is non-collapsible: adjusting for a covariate that predicts the outcome moves it away from 1 even with no confounding. Conditional and marginal odds ratios are different quantities, not a biased and an unbiased version.

open as a page

Why does a marginal structural model use IPW weights when a confounder is affected by prior treatment?

level: principalimportance: nice to knowfreq 22%

basics

~20 s

Because such a confounder is also a consequence of earlier treatment. Conditioning on it blocks part of that earlier effect; ignoring it leaves confounding. Weighting by the inverse probability of the observed treatment history breaks the link without conditioning.

open as a page

Propensity matching dropped your enterprise accounts for lack of common support - what do you report?

level: principalimportance: nice to knowfreq 29%

basics

~20 s

Report that the estimate now covers only the treated units that had comparable controls, not all of them. Quantify how many were dropped and how they differ, and state plainly that this design cannot answer the enterprise question.

open as a page