How would you set a team policy for which covariates may enter a causal model?
answer
- exclusion as the default posture
- timestamp every covariate
- prediction rules are not causal rules
- fix the list before the number
- more controls is not more conservative
basics
~10 sMake exclusion the default: ban anything measured after treatment, require a stated causal reason for every covariate that stays, and fix the list before seeing the estimate. More controls is not more conservative.
solid answer
~50 sThe policy I would write has five clauses. First, causal estimation and prediction obey different rules - a model built to forecast may use anything predictive, a model built to estimate an effect may not. Second, a hard ban on covariates measured or realised after treatment, with the timestamp recorded for each one; exceptions need a named estimand and an explicit argument. Third, every remaining covariate carries a one-line justification of its causal role, written by a human, and the list is fixed before anyone sees the coefficient. Fourth, no automated selection procedure and no fit statistic chooses controls, because a mediator is by construction the best predictor in the room. Fifth, publish a small set of specifications with the estimate's movement across them, rather than one blessed number. The cultural work is harder than the rules: 'we controlled for everything' has to stop reading as a rigor claim.
go deeper
Know that covariates in a causal model are chosen deliberately, not by an automatic procedure, and that variables measured after treatment are excluded by default.
Be able to argue the mechanics: why fit-based selection favours exactly the variables that ruin a causal interpretation, and why timing matters.
Show how you would operationalise this in real work - covariate timestamps, a written role per variable, specifications reported as a band rather than one number.
Own the tradeoffs out loud: the cost of excluding a real confounder, how much ceremony a decision's stakes justify, and how you shift a culture that reads more controls as more rigor.
## Why a policy is needed at all Every covariate added to a causal model is a substantive claim about how the world works. Add a common cause of treatment and outcome and you remove bias. Add a common effect and you create bias. Add something on the causal path and you erase part of the effect you were hired to measure. Nothing in a regression output tells you which of the three you just did, and the instinct most analysts arrive with - control for everything you have, it is more conservative - is wrong in a way that produces confident, defensible-looking, incorrect answers. A team that leaves this to individual judgment will produce a portfolio of estimates whose reliability varies with who ran them. ## The clauses worth writing down **1. Separate the two modes explicitly.** A predictive model is judged on out-of-sample accuracy and may use any variable available at prediction time. A causal estimate is judged on whether the comparison it makes is the right one, and predictive accuracy is nearly irrelevant to that. Most bad-control incidents start with a habit imported from the predictive side. Say which mode a piece of work is in, in the document, at the top. **2. Ban post-treatment variables by default.** Require every covariate to carry a timestamp or a stated position relative to treatment assignment. Anything realised afterwards is out unless someone writes down why - and 'it improves fit' is never why. This single rule catches the largest and most common class of bad controls, including mediators and variables the treatment itself influences. **3. Require a stated causal role per covariate, written before the estimate is seen.** One line per variable: what it causes and what causes it, and therefore why adjusting for it helps. This is deliberately cheap and deliberately human. It also creates the artefact a reviewer can argue with. Fixing the list before anyone looks at the coefficient removes the temptation to keep tuning until the number is agreeable. **4. No algorithmic selection of controls, and no fit-based justification.** Stepwise procedures, information criteria and variance-explained comparisons all reward variables that predict the outcome well - which is exactly the profile of a mediator. Any procedure that would happily include a post-treatment variable is disqualified from choosing controls. **5. Report a specification band, not a single number.** Show the estimate under the pre-registered specification plus a few deliberate variations, and treat large movement as a finding to explain rather than noise to suppress. Where a suspicious covariate is genuinely uncertain, run with and without and report both. **6. Name the estimand.** Total effect or direct effect - the two need different specifications and cannot be read off the same coefficient. Requiring the estimand in writing prevents the most common misreport, where a shrunken coefficient after adding a mechanism variable gets presented as the conservative estimate of the total effect. ## The tradeoffs a lead has to own The policy is not free, and pretending otherwise weakens it. - **Excluding a genuine confounder is also a bias.** The discipline is not 'fewer controls always'; it is 'each control argued'. When knowledge is thin, you sometimes accept the smaller risk of including a variable whose role you cannot fully certify, and you say so. - **Pre-registering a covariate list slows work down**, and there are decisions where a fast approximate answer beats a slow correct one. Scale the ceremony to the stakes: a directional read for a low-cost decision does not need the full apparatus, and saying which tier a piece of work is in is part of the policy. - **Human justification does not scale linearly.** If a team runs hundreds of estimates, the review capacity, not the rule, becomes the binding constraint. Templates, a reusable covariate registry for recurring domains, and review only above a stakes threshold are how this survives contact with volume. - **Structural knowledge is never complete.** Some variables genuinely cannot be classified from what the team knows. The policy should say what to do then - report the sensitivity, do not manufacture false confidence - rather than pretending a resolution exists. ## The cultural half The rules are the easy part. The harder change is that in most organisations 'we controlled for everything' functions as a rigor claim, and a shrinking coefficient reads as the honest, careful number. Both readings have to be reversed in public, repeatedly, using the team's own examples: here is the analysis where adding a plausible control destroyed the effect, and here is why the smaller number was the wrong one. Attaching the policy to a short review checklist that a peer signs, and to the template every analysis document starts from, does far more than a memo, because it puts the question in front of the analyst at the moment the choice is made.
- A stakeholder insists controlling for everything is the more rigorous choice. How do you answer?I show them a case from our own data where adding a plausible control killed a real effect, and explain that the added variable was on the causal path, so the smaller number answered a different question. Then I make the general point: each control is a causal claim, and a wrong claim biases the estimate in a direction the output cannot reveal. Rigor is argued controls, not more of them.
- How do you handle a covariate whose timing relative to treatment nobody can establish?Treat it as post-treatment until proven otherwise, since that is the costlier error. Run the estimate with and without it, report both, and make the divergence part of the writeup rather than picking the friendlier one. If the decision hinges on which is right, that is a signal to spend engineering effort on the timestamp instead of on more modelling.
- How does this differ from feature selection for a predictive model?Almost entirely. A predictive model is scored on out-of-sample accuracy, so any variable that helps is welcome, including ones caused by the treatment. A causal estimate is scored on whether the comparison is the right one, and the variables that most improve fit are often the ones that destroy the interpretation. The two modes should never share a default covariate list.
saying these in an interview costs you the question
- Uses stepwise selection or information criteria to choose controls
- Judges a causal specification by variance explained
- Presents controlled-for-everything as a rigor claim
- Adjusts the covariate set until the result is significant
- Reuses the predictive feature list for a causal estimate