How does inverse probability of treatment weighting estimate an average treatment effect?
answer
- reweight the sample, do not pair units
- one over the chance of your own arm
- rare-in-their-arm units count for more
- cloning units into a pseudo-population
- normalize the weights within each arm
basics
~20 sEach unit is weighted by the inverse of its probability of receiving the arm it actually got: treated units by 1/e(X), controls by 1/(1-e(X)). The reweighted sample mimics a population where treatment was assigned independently of the measured covariates.
solid answer
~50 sYou first model the probability of treatment given measured covariates, e(X) = P(T=1 | X). Then every treated unit gets weight `1/e(X)` and every control gets weight `1/(1-e(X))`. Units that were unlikely to land in the arm they ended up in are rare in the data, so they are counted for more; units that were near-certain to land there are counted for less. The result is a pseudo-population in which each unit is effectively cloned to represent the covariate profiles that its arm is missing, so treated and control arms have the same covariate distribution. The ATE estimate is then the weighted mean outcome among the treated minus the weighted mean among the controls, normally with the weights normalized within each arm. Standard errors must account for weighting, not be read off the naive weighted mean.
go deeper
Recall the weight itself: treated units get one over their probability of treatment, controls get one over one minus it. Be able to say in a sentence that the point is to make the two groups look alike on measured covariates.
Explain the mechanics: where the score comes from, why a low-propensity treated unit gets a big weight, what the pseudo-population is, and why the arm-normalized weighted average is the usual estimator rather than the raw weighted sum.
Show production judgment: state which estimand your weights target, check covariate balance after weighting rather than assuming it, and use standard errors that account for estimated weights instead of the naive weighted-mean formula.
Own the choice between weighting, pairing and outcome modelling for the question at hand, and be explicit that all of them buy the same thing only under measured-confounding assumptions you should be arguing for on domain grounds, not statistical ones.
## The problem being solved In observational data, who gets treated depends on covariates. Sicker patients get the drug; heavier users get the feature. A raw difference in means between treated and untreated therefore mixes the causal effect with the fact that the two groups are different kinds of unit. Adjustment estimators fix this after the fact. Inverse probability weighting fixes it by *reweighting* rather than by pairing units or by modelling the outcome. ## The propensity score and the weight The propensity score is `e(X) = P(T = 1 | X)`: the probability that a unit with covariates `X` receives treatment. It is a single number per unit that summarises how treatment-prone that covariate profile is. The inverse-probability-of-treatment weight is one over the probability of the arm the unit actually landed in: - treated unit: `w = 1 / e(X)` - control unit: `w = 1 / (1 - e(X))` A treated unit with `e(X) = 0.5` gets weight 2; a treated unit with `e(X) = 0.1` gets weight 10. The second unit had covariates that rarely lead to treatment, so the few such treated units that exist have to stand in for all the similar units who went untreated. ## The pseudo-population picture The cleanest mental model is cloning. Imagine replacing each unit by `w` copies of itself. A treated unit with `e(X) = 0.1` becomes 10 copies; the corresponding control with `1 - e(X) = 0.9` becomes about 1.1 copies. Do this for everyone and you have built a synthetic population in which, within every level of `X`, the treated and untreated arms are the same size and the covariate distribution is identical across arms. In that pseudo-population treatment is unrelated to the measured covariates, so a simple contrast of arm means is an unconfounded contrast — under the usual causal assumptions and provided the treatment model is right. This is worth saying out loud in an interview: weighting does not remove confounding by variables you never measured. It removes the imbalance in the variables that entered `e(X)`, and nothing more. ## The two estimator forms The unnormalized (Horvitz-Thompson) form averages `T*Y/e(X)` over all n units for the treated potential outcome and `(1-T)*Y/(1-e(X))` for the control one, dividing by n in both cases. It is unbiased when the weights are correct, but the weights do not sum to n in any given sample, so the estimate can drift and even fall outside the range of the observed outcomes. The normalized (Hajek) form divides each arm's weighted sum of outcomes by that arm's sum of weights, i.e. it takes a proper weighted average within each arm. It is slightly biased in finite samples but bounded by the observed outcome range and almost always lower variance. In practice the normalized form is the default, and equivalently you can fit a weighted regression of the outcome on the treatment indicator using the weights. ## Choosing the estimand The weights above target the ATE: the effect if the whole population were treated versus the whole population untreated. If you want the ATT (the effect among those actually treated), the weights change: treated units get weight 1 and controls get `e(X) / (1 - e(X))`, which reweights the control arm to look like the treated arm. Being able to name which estimand your weights target is a standard follow-up. ## Practical cautions 1. **The propensity model matters.** The weights are only as good as `e(X)`. A misspecified treatment model gives systematically wrong weights and a biased estimate. 2. **Inference is not free.** The weights are estimated, and cloning units inflates the apparent sample size. A naive weighted standard error understates uncertainty; use robust (sandwich) standard errors that account for the weighting, or a resampling scheme that refits the treatment model each time. 3. **Overlap is the failure mode.** If some covariate region contains almost no treated units, the corresponding scores approach zero and the weights explode, which shows up as a handful of units dominating the estimate. 4. **Check what the weights did.** After weighting, the covariate distributions in the two arms should look alike; if they do not, the weights are not doing their job and the treatment model needs work. ## What a good answer sounds like Name the weight, give the pseudo-population picture in one sentence, say which estimand the weights target, and immediately flag that the whole thing rests on a correctly specified treatment model plus measured confounders only.
- How do the weights change if you target the ATT instead of the ATE?For the ATT you leave the treated arm alone (weight 1) and give each control the odds weight `e(X) / (1 - e(X))`, which reweights controls to match the covariate distribution of the treated. The estimand becomes the effect among those actually treated, which is a different number from the ATE whenever the effect varies with covariates.
- What is the difference between the unnormalized and the weight-normalized version of the estimator?The unnormalized form divides each arm's weighted outcome sum by n, so the weights need not sum to n and the estimate can fall outside the observed outcome range. The normalized form divides by that arm's sum of weights, giving a genuine weighted average: slightly biased in finite samples, bounded, and usually much lower variance. The normalized form is the default.
- Why can you not report the ordinary standard error of the weighted mean?Two reasons. The weights are themselves estimated from a fitted treatment model, which adds uncertainty the naive formula ignores, and the effective information in a weighted sample is smaller than its nominal size because a few units carry much of the weight. Use robust sandwich standard errors built for weighted estimation, or a resampling scheme that refits the treatment model each replicate.
It is exit polling with unequal response rates: if only one in ten young voters answers, you count each young respondent ten times so the poll reflects the electorate rather than the people who happened to answer.
saying these in an interview costs you the question
- Claims weighting removes confounding by unmeasured variables
- Weights treated units by e(X) rather than by 1/e(X)
- Treats the sum of weights as a real sample size
- Reports the naive weighted-mean standard error unadjusted
- Cannot say whether the weights target the ATE or the ATT