Why is obesity a problematic exposure when you want to estimate its causal effect?
answer
- a state is not an action
- many routes to the same value
- what exactly would the intervention be
- observed outcome equals potential outcome received
- specify a protocol, accept a narrower claim
basics
~10 sObesity names a state, not an intervention. Losing weight through diet, exercise, illness or surgery are different treatments with different effects, so the causal contrast is undefined until you say which intervention you mean.
solid answer
~40 sThe consistency assumption says the outcome you observe equals the potential outcome under the treatment actually received: `Y_i = Y_i(T_i)`. That only makes sense if the label treated picks out one thing. An exposure like obesity, or an instruction like exercise more, hides many versions with genuinely different effects, so there is no single `Y(1)` for a unit to have. The repair is the question every causal reviewer asks: what exactly would the intervention be? Replace the state with a protocol, for example assignment to a 12-week supervised exercise programme, and the estimand becomes well defined and estimable. The cost is scope. A sharply specified intervention answers a narrower question and may hold for a narrower population, which is a real tradeoff rather than a defect to hide.
go deeper
Recognise that a causal question needs something someone could actually do, and that a characteristic like weight or engagement level is not by itself a treatment.
Be able to state that the observed outcome must equal the potential outcome under the treatment received, and explain why several routes to the same state break that link.
Show that you refuse ambiguous exposures and rewrite them as protocols with a named comparison and timing, while acknowledging the narrower scope that follows.
Own the reframing conversation with stakeholders: turn an unanswerable question about a state into a decision about an intervention the organisation could run, and set that expectation for the team.
## Consistency, stated Consistency links the observed data to the counterfactual framework: `Y_i = Y_i(T_i)` The outcome you actually record for a unit is that unit's potential outcome under the treatment it actually received. It sounds like bookkeeping, and it is trivially true when the treatment is a single, sharply defined act. It becomes substantive as soon as the treatment label covers several materially different things, because then it is ambiguous which potential outcome the observed value corresponds to. ## Why a state is not an intervention Obesity is a characteristic of a person, not something anyone does to them. To speak of its causal effect you must implicitly compare a world where the person is obese with a world where they are not, and worlds differ in how they got there: - Weight reduced by a sustained dietary change. - Weight reduced by an exercise programme. - Weight reduced by bariatric surgery. - Weight reduced by undiagnosed illness. These have different effects on almost any downstream outcome, and the last one reverses the sign of the association entirely, since the illness itself drives the outcome. Estimating the effect of obesity therefore estimates an average over an unknown mixture of routes, weighted by however people in this dataset happened to lose or gain weight. Move to another dataset with a different mixture and the number changes, with no confounding having been introduced. The same problem afflicts vaguer exposures generally: exercise more, eat healthily, be less stressed, use the product more. Each names an outcome of many possible actions rather than an action. ## The repair The standard discipline is to ask what exactly the intervention would be, and to answer it before any estimation: 1. **Name the action.** Assignment to a 12-week supervised walking programme of five sessions per week. Not exercising more. 2. **Name the comparison.** Assignment to usual activity, or to a stretching programme, or to a waitlist. The estimand is a contrast, so both arms need definition. 3. **Name the timing.** When does the intervention start, how long does it last, and when is the outcome measured. 4. **Check that the intervention is one a unit could plausibly have received.** This is where consistency meets positivity: an intervention nobody in your population could receive has no comparison group. What you get is a well-defined estimand: the effect of that protocol, in that population, on that outcome. What you lose is generality. The effect of a supervised programme does not automatically tell you the effect of a public-health message urging people to exercise, because the message is a different intervention that changes behaviour only partially. ## Relationship to the other assumptions Consistency and the treatment-variation half of SUTVA are two views of the same requirement. SUTVA states it as a property of the design (no hidden versions of the treatment); consistency states it as the link between potential outcomes and observed data. Different authors divide the labels differently, and interviewers rarely care which convention you use as long as you can articulate the requirement and spot a violation. It is also worth seeing that consistency is not a statistical problem to be fixed by more or better adjustment. No amount of covariate control repairs a treatment whose meaning is ambiguous, because the ambiguity is in the estimand, not in the comparison. ## Recognising the failure mode in a work setting Product and analytics work is full of state-shaped exposures: power users, engaged accounts, adopters of a feature. Asking what is the causal effect of being a power user has the same defect as asking about obesity, because becoming one has many routes. The productive reframe is to name an action the organisation could actually take, such as enrolling a user in an onboarding flow, and estimate the effect of that. This is often the single most useful contribution a causal thinker makes to a product discussion: replacing an unanswerable question with a nearby answerable one. ## What strong answers include A clear statement that the exposure must be an intervention, an example of two routes to the same state with different effects, the observation that the resulting estimate is an average over an unknown version mix, and the honest note that specifying a protocol narrows the claim.
- State the consistency assumption formally.`Y_i = Y_i(T_i)`: the outcome observed for a unit is exactly that unit's potential outcome under the treatment it actually received. It fails when the treatment label is ambiguous, because then several distinct potential outcomes could correspond to the same recorded exposure value and the estimand has no single meaning.
- Does specifying a precise protocol solve the problem outright?It makes the estimand well defined, which is the essential step, but it narrows the claim. A sharply specified intervention often has fewer people in the population who could plausibly have received it, so overlap can shrink, and the result speaks only to that protocol. You trade generality for a question the data can actually answer.
- How would you reframe the causal effect of being a power user?Replace the state with an action the organisation could take. Instead of asking what being a power user causes, estimate the effect of enrolling users in a specific onboarding flow, or of a particular feature prompt, measured over a stated horizon. That converts an undefined contrast into an intervention someone can actually decide to run.
Asking for the effect of being warm is not a question until you say whether the warmth came from a coat, a fever, or a fire. The number differs by route.
saying these in an interview costs you the question
- Says the effect of obesity is well defined once you control for confounders
- Treats any measured characteristic as a valid exposure
- Assumes different routes to the same state have similar effects
- Reports an effect without naming the intervention or the comparison
- Believes more covariates can fix an ambiguous treatment definition