skip to content

What does an augmented IPW (doubly robust) estimator add over plain IPW?

level: seniorimportance: should knowfreq 38%

answer

  1. two models instead of one
  2. treatment model or outcome model, only one
  3. predictions plus weighted residual correction
  4. bias is the product of two errors
  5. efficient when both nuisance models are right

basics

~20 s

It combines a treatment model with an outcome model and stays consistent if either one is correct, instead of betting everything on the propensity model. It also reaches the best achievable precision when both models are right.

solid answer

~50 s

Plain inverse probability weighting is a single bet: get the treatment model wrong and the estimate is biased. Augmented IPW adds a second model for the outcome and combines them, so it is consistent if *either* model is correctly specified — hence 'doubly robust'. A convenient way to see it is that the estimate of the treated mean is the average of the outcome model's predictions plus an inverse-probability-weighted average of that model's residuals among the treated: if the outcome model is right the correction term has mean zero, and if the propensity model is right the weighted residual term repairs whatever the outcome model got wrong. When both are correct it also attains the smallest achievable asymptotic variance. It is not magic: with near-zero propensity scores it can be more unstable than either component, and if both models are wrong it can be worse than either.

go deeper

for a junior

Know the headline: the estimator uses a model for who gets treated and a model for the outcome, and it still works if just one of them is right. The name to remember is doubly robust.

for a middle

Explain the construction: outcome-model predictions for everyone plus an inverse-probability-weighted average of the residuals, and why each half of the guarantee holds when its own model is correct.

for a senior

Show operational judgment: check overlap before reaching for it, know that dividing by tiny scores makes the correction term volatile, and use cross-fitting when the nuisance models are flexible learners.

for a principal

Own the choice of estimator for the team. Weigh the extra robustness against a two-model pipeline that is harder to review, explain and monitor, and be clear that no estimator upgrades a weak identification argument.

## Two routes to the same estimand There are two obvious ways to adjust for measured confounders. You can model **treatment** — estimate `e(X) = P(T = 1 | X)` and reweight — or you can model the **outcome** — estimate `m1(X) = E[Y | T = 1, X]` and `m0(X) = E[Y | T = 0, X]`, predict both potential outcomes for everyone, and average. Each is consistent only if its own model is right. Choosing between them is a bet on which model you trust. Augmented inverse probability weighting refuses the bet by using both. ## The estimator For the mean outcome under treatment, AIPW averages over all n units the quantity `m1(X_i) + T_i * (Y_i - m1(X_i)) / e(X_i)` and symmetrically for the control mean, using `m0` and `1 - e(X)`. The effect estimate is the difference of the two means. Read the formula as **prediction plus weighted correction**. The first term is the outcome model's guess for every unit, treated or not. The second term takes the treated units' residuals — how wrong the outcome model was on the people you actually observe under treatment — and spreads them back over the population using inverse probability weights. ## Why it is doubly robust Suppose the **outcome model is correct**. Then the residuals `Y - m1(X)` have conditional mean zero among treated units at every `X`, so the correction term averages to zero no matter what weights you attach. The estimator collapses to the outcome-model answer and is consistent even if `e(X)` is nonsense. Now suppose the **propensity model is correct** but the outcome model is not. Rewrite the same expression as the ordinary IPW term `T*Y/e(X)` minus a term involving `(T - e(X)) * m1(X) / e(X)`. Because `T - e(X)` has conditional mean zero when the propensity model is right, that second term averages to zero, leaving the consistent IPW estimator. The wrong outcome model does not bias the answer; it only changes the variance. So the bias of AIPW is governed by a *product* of the two models' errors. Either factor being zero kills the bias. That is the whole idea, and it is the sentence to say in an interview: **you get two chances to be right and only need one.** ## Efficiency, not just robustness When both models are correct, AIPW attains the semiparametric efficiency bound — no other estimator relying on the same assumptions has smaller asymptotic variance. Even when only the propensity model is right, the outcome model usually acts as a variance-reducing control, so AIPW is typically tighter than plain IPW. ## Modern practice: machine-learned nuisances and cross-fitting Because the bias depends on the *product* of the two errors, each model may converge more slowly than the usual root-n rate as long as the product converges faster. That is what licenses flexible machine-learned models for `e(X)` and `m(X)` while still getting valid confidence intervals for the effect. Two conditions come with it. First, **cross-fitting**: fit the nuisance models on one split of the data and evaluate the estimating equation on the held-out split, rotating through folds, so that overfitting in a flexible model does not leak into the effect estimate. Second, the nuisance models must be regularised sensibly — a flexible treatment model that separates the arms manufactures near-zero scores and destroys stability. ## Where it disappoints 1. **Extreme weights still hurt.** The correction term divides by `e(X)`. If some scores are near zero the augmentation term can be wildly variable, and AIPW can have higher variance than a trimmed plain IPW estimate. 2. **Both models wrong is not covered.** Double robustness protects against one failure, not two. With both nuisances mildly misspecified and poor overlap, the doubly robust estimate can be worse than either single-model estimate. 3. **It does not buy an assumption.** No amount of augmentation addresses unmeasured confounding or a covariate region with essentially no treated units. Both routes rest on the same identification story. 4. **It is harder to explain.** For a stakeholder audience, a well-diagnosed weighted estimate with clear balance evidence may communicate better than a two-model estimator that nobody in the room can sanity-check. ## How to answer the question Name the two nuisance models, state the guarantee precisely (consistent if either is correct, efficient if both are), give the prediction-plus-weighted-residual reading of the formula, and finish with the honest limits: overlap problems and simultaneous misspecification are still fatal.

  • Give the precise statement of the double robustness guarantee.
    The estimator is consistent for the target effect if at least one of the two nuisance models — the propensity model for treatment or the outcome regression — is correctly specified, without knowing which. When both are correct it additionally attains the semiparametric efficiency bound. It offers no protection when both are misspecified.
  • Can a doubly robust estimator be worse than plain IPW in practice?
    Yes. The augmentation term divides by the propensity score, so with scores near zero it can be extremely variable, and a trimmed plain weighted estimate may have lower total error. When both nuisance models are mildly wrong and overlap is poor, the doubly robust estimate can also be more biased than either single-model estimate. It is a better default, not a guarantee.
  • Why is cross-fitting recommended when the nuisance models are flexible machine-learned models?
    Fitting a flexible model and evaluating the effect on the same rows lets overfitting in the nuisance leak into the effect estimate, biasing it and shrinking the interval. Cross-fitting estimates the nuisances on one fold and evaluates the estimating equation on the held-out fold, rotating through folds, which restores valid inference under mild rate conditions.

It is a belt-and-braces bet: you back two horses that pay out on the same race, and only one of them has to come in.

saying these in an interview costs you the question

  • Claims double robustness protects against unmeasured confounding
  • Says both nuisance models must be correctly specified
  • Thinks it removes the need for overlap between arms
  • Believes it is always lower variance than plain weighting
  • Fits flexible nuisance models on the same rows without splitting

context