skip to content

How does standardisation, the g-formula, turn a fitted outcome model into an average treatment effect?

level: middleimportance: should knowfreq 44%

answer

  1. predict twice, average the difference
  2. set treatment on for everyone
  3. average over covariates, not condition
  4. handles interactions and non-linear links
  5. bootstrap the whole procedure

basics

~20 s

Fit one outcome model on treatment and covariates, then predict every unit twice, once with treatment set to 1 and once set to 0, and average the difference of those two predictions over the whole sample.

solid answer

~50 s

Standardisation is a plug-in recipe. Fit a model for the outcome given treatment and covariates. Then, for every unit in the sample, produce two predictions: one with the treatment variable forced to treated, one forced to control, keeping that unit's real covariates. Average those two prediction sets and take the difference — that is the estimated ATE, averaged over the covariate distribution of your sample. It works because it estimates `E[Y | A=a, X]` from the model and then averages over `X` rather than conditioning on it. The advantages over reading a single coefficient: it handles interactions, nonlinearity and non-linear link functions correctly, and it gives a marginal effect on whatever scale you choose. The assumptions are the usual ones plus correct model specification, and the standard error comes from bootstrapping the whole procedure, not from the model's own coefficient errors.

go deeper

for a junior

Know the shape of the recipe: one outcome model, predict the whole sample as treated, predict it again as untreated, and take the mean difference of the two prediction sets.

for a middle

Explain why averaging predictions differs from reading a coefficient once interactions or a non-linear link are present, and which estimand you get by choosing which rows you average over.

for a senior

Show that you check covariate overlap before predicting counterfactuals, that you bootstrap the entire procedure for inference, and that you can say what breaks when the outcome model is misspecified.

for a principal

Own the estimand choice and the reporting: which population the average is taken over, what decision that answers, and how much of the result rests on one model that nobody else has scrutinised.

## The idea in one sentence A model gives you predictions conditional on covariates; a causal effect is a contrast between two whole worlds. Standardisation — also called the g-formula, or g-computation — is the bridge: predict everyone under each treatment value, then average. ## The recipe 1. Fit one model for the outcome as a function of treatment `A` and covariates `X`. Include interactions between `A` and `X` if the effect plausibly varies; the recipe accommodates them. 2. Take a copy of the whole dataset and set `A = 1` for every row, leaving covariates untouched. Predict. Call the mean of those predictions `m1`. 3. Take another copy with `A = 0` for every row. Predict. Call the mean `m0`. 4. The estimated average treatment effect is `m1 - m0`. The two averages estimate `E[Y^1]` and `E[Y^0]` — mean outcome if everyone were treated and if nobody were. The formula behind it is `E[Y^a] = average over X of E[Y | A=a, X]`, where the average uses the covariate distribution of the population you care about. ## Why not just read the treatment coefficient On a linear model with no treatment-covariate interaction and a correctly specified form, the coefficient and the standardised estimate agree, and the extra work buys nothing. Everywhere else they part company: - **Effect heterogeneity.** If the model contains `A*X` interactions, no single coefficient is the population effect. The `A` coefficient is the effect at `X = 0`, which may not even be a covariate value that exists. - **Non-linear links.** A logistic or log-link model's coefficient is a conditional contrast on the model's scale — a conditional odds ratio, say. Averaging predicted probabilities instead gives you a marginal risk difference or risk ratio, which is what most decisions actually need. - **Choice of estimand.** Averaging over the whole sample gives the ATE. Averaging the same predicted difference over the treated rows only gives the effect of treatment on the treated. Same fitted model, different question, and the recipe makes the choice explicit rather than implicit. ## What has to hold Standardisation is an estimation strategy, not an identification argument. It gives a causal quantity only if: - **No unmeasured confounding given `X`** — treatment is as good as randomised within covariate strata. - **Positivity** — every covariate pattern has some chance of both treatment values. This one deserves attention: step 2 asks the model to predict outcomes under treatment for units that were never treated. Where such units genuinely do not exist in the data, the prediction is extrapolation dressed up as an estimate, and nothing in the output flags it. Checking the overlap of covariate distributions between arms before running the recipe is part of the job. - **Consistency** — a well-defined intervention that matches the treatment as observed. - **Correct outcome model** — the whole estimate rides on the model being right, since every counterfactual mean comes from it. There is no second line of defence in this estimator. ## Inference The naive standard error from the fitted model does not apply, because the estimate is a nonlinear function of the fitted parameters averaged over the sample's own covariates. The standard approach is a bootstrap of the entire procedure: resample rows, refit the model, redo the two predictions and the averaging, and take the spread of the resulting estimates as the sampling distribution. Bootstrapping only the prediction step while holding the fit fixed understates the uncertainty, because it ignores the fact that the model itself was estimated. ## How to talk about it A crisp interview answer is: I fit an outcome model, then predict the whole sample twice with treatment forced on and forced off, and take the mean difference. That gives me a marginal effect on the scale I care about, it survives interactions and nonlinear links, and I bootstrap the whole thing for the interval. The assumption I lean on hardest is that the outcome model is right, and I check overlap first so the counterfactual predictions are interpolation rather than invention.

  • When is a single treatment coefficient already equal to the ATE?
    When the outcome model is linear in the outcome scale, correctly specified, and contains no treatment-covariate interaction, so the effect is constant across covariate values. Once effects vary, an additive least-squares fit returns a weighted average of stratum effects with weights driven by treatment variance within strata, which is not the population average.
  • How would you estimate the effect of treatment on the treated instead of the ATE?
    Keep the same fitted model and the same two predictions per unit, but average the predicted difference over the treated rows only. That reweights the effect toward the covariate mix of units who actually received treatment, which is often the more relevant estimand when you are evaluating a programme as it was rolled out.
  • How do you get a confidence interval for a standardised estimate?
    Bootstrap the whole pipeline: resample rows, refit the outcome model, redo both counterfactual prediction passes, re-average, and read the spread across replicates. The model's own coefficient standard errors do not apply, because the estimate is a nonlinear average of predictions rather than a single fitted parameter.
  • What is the biggest practical risk with this estimator?
    That the outcome model is wrong, since every counterfactual mean comes from it and there is no second line of defence. The close second is silent extrapolation: predicting treated outcomes for covariate regions where no treated unit exists produces numbers with no data behind them, so I check arm overlap on the covariates before trusting anything.

Instead of asking how two teams compared on the pitches they happened to play on, you replay the entire fixture list twice, once with each team, and compare the two full seasons.

saying these in an interview costs you the question

  • Predicts only the units that did not get that treatment
  • Reports the model coefficient standard error as the effect's error
  • Ignores overlap and extrapolates into empty covariate regions
  • Thinks the recipe removes the need for a no-unmeasured-confounding assumption
  • Averages predictions over the treated rows and calls it the ATE

context