skip to content

Why does a marginal structural model use IPW weights when a confounder is affected by prior treatment?

level: principalimportance: nice to knowfreq 22%

answer

  1. the confounder is also a consequence
  2. adjusting blocks part of the effect
  3. you cannot include and exclude one variable
  4. use it in the weight, not a regressor
  5. weights multiply across time periods

basics

~20 s

Because such a confounder is also a consequence of earlier treatment. Conditioning on it blocks part of that earlier effect; ignoring it leaves confounding. Weighting by the inverse probability of the observed treatment history breaks the link without conditioning.

solid answer

~50 s

With a treatment repeated over time, a variable like engagement can confound the next period's treatment while itself being a result of the previous period's treatment. That puts any single regression in a trap: adjust for it and you block the part of the earlier treatment's effect that flows through it, and can open a spurious path if it shares an unmeasured cause with the outcome; leave it out and the later treatment stays confounded. Weighting escapes the trap because it never conditions on the variable. Each unit is weighted by the inverse probability of its whole observed treatment sequence given its treatment and covariate history, so in the weighted pseudo-population treatment at each period is unrelated to the past. You then fit a simple marginal model of the outcome on the treatment history in that population — stabilized weights are effectively mandatory, since the product over periods otherwise explodes.

go deeper

for a junior

You are unlikely to be asked this, but know the shape of the trap: when a treatment repeats over time, a variable the treatment itself changed must not simply be dropped into the regression as a control.

for a middle

Be able to say why the variable is both a confounder of later treatment and a consequence of earlier treatment, and that weighting uses it to build a weight instead of conditioning on it.

for a senior

Show you could run it: per-period treatment models, a product weight in stabilized form, weight diagnostics per period, and a check that covariates were recorded before the decisions they inform.

for a principal

Own the decision to attempt it at all. Weigh a fragile many-period weighted estimate against a coarser regime comparison the organisation can audit, and set the reporting standard for weight concentration and sensitivity before any result is circulated.

## The setting A treatment is not applied once but repeatedly: a drug given or withheld each month, a discount applied each week, a policy toggled each quarter. Write the treatment sequence as `A_1, ..., A_T`, the covariates measured before each decision as `L_1, ..., L_T`, and the final outcome as `Y`. The interesting complication is **treatment-confounder feedback**: some `L_t` both influences the next treatment decision and is itself a consequence of earlier treatment. A clinician's dosing depends on a lab value that the earlier dose changed; a targeting rule depends on engagement that last week's discount raised. ## Why one regression cannot work Suppose you fit a single outcome regression on the treatment history and adjust for every `L_t`. **Adjusting for `L_t` blocks part of the effect you want.** Because `L_t` sits on the causal path from `A_{t-1}` to `Y`, conditioning on it removes exactly the portion of the earlier treatment's effect that travelled through it. The coefficient on earlier treatment now measures only the leftover direct path, which is not the total effect anyone asked about. This is over-adjustment, and it happens even with perfectly measured data. **Adjusting can also create bias.** `L_t` is a common effect of earlier treatment and of whatever else drives it. If some unmeasured factor influences both `L_t` and `Y`, conditioning on `L_t` opens a non-causal association between earlier treatment and the outcome — a collider path that did not exist before you adjusted. **Not adjusting is no better.** Leave `L_t` out and the later treatment decisions remain confounded by it, since it drives both `A_t` and `Y`. There is no way to include and exclude the same variable at once, so no single regression on the observed data returns the total effect of the treatment sequence. This is the classic argument for a different machinery. ## What weighting does instead Weighting never conditions on `L_t`; it uses it only to build a weight. For each unit compute, at each period, the probability of the treatment it actually received given its treatment history and covariate history, and take the product over periods. The weight is one over that product, and in practice the stabilized version is used: the numerator is the product of the probabilities of the same treatment history given only past treatment (and possibly fixed baseline variables). In the pseudo-population these weights create, the dependence of `A_t` on the time-varying covariates has been severed: at every period, treatment looks as if it were assigned without regard to the history. Because `L_t` was never conditioned on, the effect flowing through it is preserved. You then fit the **marginal structural model** itself — a deliberately simple model of the outcome as a function of the treatment history alone, for instance on cumulative exposure or on always-treat versus never-treat — in the weighted population. Its coefficients carry a causal interpretation about the counterfactual outcomes under those treatment regimes, which is precisely what the naive regression could not deliver. ## The judgment call a lead has to make This machinery is powerful and expensive, and deciding whether to spend it is the real principal-level question. **Costs.** Weights multiply across periods, so their spread grows quickly with the number of periods; by ten periods a few units can dominate entirely and the effective sample size collapses. Every period needs its own credible treatment model, and every period needs enough units of both kinds at each history for the weights to be meaningful. You also need the time-varying covariates to be measured well and at the right times; a covariate recorded after the decision it supposedly informed silently invalidates the whole construction. **Alternatives worth considering first.** Coarsen the treatment into a small number of regimes that the business actually cares about — started within the first month versus never started — so the weights involve far fewer periods. Shorten the horizon. Restrict to a subpopulation with real variation in treatment at every period. Or, when the intervention is genuinely under your control, run an experiment on the assignment rule rather than reconstructing it from observational history. **Communication.** A weighted time-varying analysis is hard for stakeholders to audit. If you commit to one, commit also to publishing the weight diagnostics per period, the effective sample size over time, and a sensitivity analysis over the truncation rule — otherwise the result is a number nobody can challenge, which is worse than a cruder number everyone can. ## The one-sentence answer When a confounder is also a consequence of earlier treatment, conditioning on it is both necessary and forbidden; weighting resolves the contradiction by using the variable to construct weights rather than as a regressor, and the marginal structural model is then fitted in the resulting pseudo-population.

  • What exactly goes wrong with the earlier treatment's coefficient if you adjust for the time-varying covariate in a regression?
    The covariate lies on the path from earlier treatment to the outcome, so conditioning on it removes that portion of the effect and the coefficient measures only the remaining direct path. Worse, because the covariate is a common effect of earlier treatment and possibly of unmeasured causes of the outcome, conditioning can also open a spurious association that was not there before.
  • What practical problem grows as the number of treatment periods increases?
    The weight is a product of per-period inverse probabilities, so its spread compounds. After several periods a small number of units can hold most of the weight, the effective sample size collapses, and the estimate becomes a function of a handful of histories. Stabilization, truncation and a coarser treatment definition are the usual defences, and each has to be reported.
  • Before committing a team to this analysis, what would you check about the data?
    That the time-varying covariates are measured before each treatment decision rather than after it, that both treatment values genuinely occur at every period within the relevant histories, and that the horizon and treatment definition can be coarsened to something stakeholders care about. If any of those fails, a simpler, well-defined regime comparison is the more defensible deliverable.

saying these in an interview costs you the question

  • Adjusts for the post-treatment covariate in one regression
  • Says the weights condition on the time-varying confounder
  • Ignores that weights multiply across periods
  • Uses a covariate measured after the treatment decision
  • Treats the method as a fix for unmeasured confounding

context