Why is controlling for a post-treatment mediator called a bad control?
answer
- on the path, not beside it
- holding the mechanism fixed
- total effect versus direct effect
- nothing measured after treatment
- vanished effect means fully mediated
basics
~20 sA mediator sits on the causal path from treatment to outcome, so adjusting for it removes the very effect you set out to measure. A real total effect shrinks toward zero, and new bias can appear.
solid answer
~50 sA mediator is a variable the treatment causes and which in turn causes the outcome. If a drug reduces strokes only by lowering blood pressure, then blood pressure is the mediator - and adding it as a control estimates the effect of the drug *holding blood pressure fixed*, which is close to zero by construction. The drug works; the specification hid it. That is over-control bias, and it is why the standard rule is never to control for anything measured after treatment. There is a second, less obvious problem: if the mediator and the outcome share an unmeasured common cause, conditioning on the mediator also turns it into a collider and creates bias in an unpredictable direction. So the honest reading of a coefficient that vanishes when a mediator enters the model is not 'no effect' but 'the effect flows through this variable'.
go deeper
Know that a mediator is caused by the treatment and in turn causes the outcome, and that controlling for it removes part of the effect you wanted to measure.
Explain the decomposition into total, direct and indirect effect, and show with a concrete mechanism why the coefficient collapses when the mediator enters.
Demonstrate that you check the timing of every covariate against treatment assignment, and that you can tell a stakeholder why the smaller coefficient is not the safer one.
Own the estimand decision: decide and document whether the organisation is buying a total or a direct effect, and what extra assumptions the direct-effect version commits you to.
## Total effect, direct effect, indirect effect Suppose a treatment `T` affects an outcome `Y` through some intermediate variable `M` that `T` causes and that in turn causes `Y`. `M` is called a **mediator**. Three quantities are worth naming: - The **total effect** of `T` on `Y`: everything the treatment does, through every route. - The **indirect effect**: the part that flows through `M`. - The **direct effect**: the part that does not flow through `M` - loosely, the effect of `T` on `Y` if `M` were held fixed. Almost every applied question - does this drug prevent strokes, does this policy raise earnings, does this change raise retention - is asking for the **total effect**. And the total effect is what you get from a well-specified model that does *not* condition on `M`. ## What over-control does Take a blood-pressure drug that reduces stroke risk entirely by lowering blood pressure. There is no other route. If you regress strokes on the drug with no post-treatment controls, you recover the real, useful, life-saving total effect. Now add blood pressure as a control 'to be careful'. You are now comparing treated and untreated patients *at the same blood pressure* - and among patients with identical blood pressure, the drug does nothing further. The coefficient collapses toward zero. This is **over-control bias**. Nothing went wrong numerically; the model answered a different question from the one asked. The dangerous part is the interpretation: a team that reads the shrunken coefficient as evidence the drug does not work has drawn the exact opposite of the truth. A vanished effect after adding a mediator is evidence that the effect is *fully mediated* by that variable. Partial mediation gives the partly-attenuated version. If the treatment also works through a second route, the coefficient shrinks but does not vanish, and the residue is the direct effect, not a weaker estimate of the total effect. Either way the number stops answering the original question. ## The second problem: the mediator can be a collider too Over-control is not the only cost. Suppose some unmeasured variable `U` affects both the mediator and the outcome - lifestyle affecting both blood pressure and stroke risk, say. The mediator is now a common effect of the treatment and of `U`. Conditioning on it makes treatment and `U` associated in the data, which contaminates the estimate with a bias whose sign you cannot predict from the model output. So conditioning on a mediator can be worse than merely answering the wrong question: it can answer *nothing* reliably. ## When conditioning on a mediator is legitimate Sometimes the direct effect genuinely is the question. Does the training programme raise wages through anything other than the extra certification it confers? Does a drug have a benefit beyond its blood-pressure route? These are **mediation analysis** questions, and they are respectable - but they require assumptions strictly stronger than the ones needed for a total effect. In particular you need no unmeasured confounding between the *mediator and the outcome*, which randomization of the treatment does *not* buy you: randomizing `T` says nothing about why patients ended up at the blood pressures they did. Mediation results should be presented with that assumption stated and probed, not slipped in as a robustness check. ## Practical rules - **Timestamp every covariate.** If it was measured, or could have changed, after treatment was assigned, it does not go in the model by default. - **Name the estimand first.** Write down whether you want the total or the direct effect *before* choosing controls; the two require different specifications and cannot both be read off one coefficient. - **Do not let fit statistics choose controls.** A mediator is by construction highly predictive of the outcome, so adding it improves fit while destroying the causal interpretation. Predictive quality and causal validity are different criteria. - **Report the drop, do not bury it.** If a coefficient moves sharply when a post-treatment variable enters, that is a finding about the mechanism - say so explicitly rather than presenting the smaller number as the more conservative one. The umbrella point is that 'conservative' is not a property of having more controls. Adding a variable is a causal claim, and for post-treatment variables that claim is usually wrong.
- A drug's effect vanishes once blood pressure is added as a control. What do you conclude?That the drug works entirely through lowering blood pressure, not that it fails. Conditioning on the mediator estimates the effect at fixed blood pressure, which is near zero by construction when the mechanism is fully mediated. The useful number for a treatment decision is the unadjusted total effect; the shrunken coefficient is a statement about mechanism.
- When is conditioning on a mediator the right analysis?When the direct effect really is the question - whether an intervention does anything beyond its known channel. That is mediation analysis, and it needs an extra assumption the total-effect estimate does not: no unmeasured confounding between mediator and outcome. Randomizing treatment does not deliver that, so state the assumption and test its sensitivity rather than assuming it.
- A colleague controls for a post-treatment variable they insist is not a mediator. Is that safe?Usually not. Even off the causal path, a variable affected by treatment is often a common effect of treatment and of unmeasured causes of the outcome, so conditioning on it induces bias of unknown sign. The timing rule exists precisely because 'it is not a mediator' is hard to establish and rarely rescues the adjustment.
saying these in an interview costs you the question
- Adds any strongly predictive variable to tighten the model
- Reads a vanished coefficient as proof of no effect
- Calls a post-treatment variable a confounder because it correlates with both
- Assumes randomizing treatment makes mediator adjustment valid
- Treats more controls as the more conservative specification