What does SUTVA, the stable unit treatment value assumption, require in a causal study?
answer
- two parts, not one
- my outcome, my treatment only
- one version of the treatment
- spillover between units breaks it
- vaccination protects the unvaccinated too
basics
~10 sSUTVA has two parts: one unit's treatment does not affect another unit's outcome, and there is only one version of the treatment, so every unit labelled treated received effectively the same thing.
solid answer
~50 sThe first half is no interference. A unit's potential outcomes depend on its own assignment only, not on anyone else's. A vaccination campaign breaks this plainly: vaccinating some people lowers infection risk for the unvaccinated around them, so a control unit's outcome now depends on other units' treatments and `Y(0)` is no longer well defined by that unit's own assignment. The second half is no hidden versions of treatment: if the label treated pools people who received two doses with people who received one, the quantity you estimate is a mixture whose value depends on an unrecorded mix. When SUTVA fails, the practical consequence is that the individual-level contrast stops answering the question you meant. Under herd immunity the vaccinated-versus-unvaccinated difference understates the benefit of vaccinating the whole population, because the comparison group was partly protected too.
go deeper
Recall both halves of the assumption and give one concrete violation, such as vaccinated people protecting their unvaccinated neighbours, without needing the formal notation.
Explain why interference makes the number of potential outcomes explode, and why a treatment label that pools different doses turns the estimate into an average over an unknown mix.
Show that you look for shared households, teams, markets or communication paths before choosing a unit of analysis, and that you can reason about the direction of the resulting bias.
Own the decision of what the analysis unit and estimand should be when units clearly interact, and make clear to stakeholders which quantity the organisation is actually buying.
## The two components SUTVA is the assumption that makes the notation `Y_i(1)` and `Y_i(0)` meaningful in the first place. It bundles two requirements. **1. No interference.** Unit i's potential outcomes depend only on unit i's own treatment. Formally, if potential outcomes were written against the full assignment vector for the whole sample, `Y_i(t_1, t_2, ..., t_n)`, SUTVA is what collapses that to `Y_i(t_i)`. **2. No hidden versions of treatment**, sometimes called treatment-variation irrelevance. Every unit recorded as treated received the same intervention in whatever sense matters for the outcome, so the label treated picks out one well-defined thing. ## Why interference is not a technicality Without the first component the number of potential outcomes per unit explodes. With n units and binary treatment there are `2^n` assignment vectors, so each unit has `2^n` potential outcomes rather than two, and no sample identifies them all. Analysts who must handle interference usually restrict attention to a summary of others' treatments, for example the fraction of a unit's neighbours treated, and define the estimand against that summary instead. That is a change of estimand, not a repair of the original one. ## The vaccination example Herd immunity is the textbook violation. When a share of a community is vaccinated, transmission falls for everyone, including those who were not vaccinated. Now: - An unvaccinated person's infection risk depends on how many others were vaccinated, so their `Y(0)` is not a fixed property of that person. - Comparing vaccinated to unvaccinated people inside the same community measures only the extra protection of being vaccinated yourself, on top of the protection everybody already received. It understates the benefit of the campaign as a whole, because the comparison group was partly protected by the treatment being studied. Interference does not always attenuate. In settings where treating one unit takes something away from another, the untreated group can be made worse off, and the naive contrast then overstates the effect. The direction depends on the mechanism, which is why naming the mechanism comes before naming the bias. ## The second component in practice Hidden versions are easy to overlook because they hide behind a tidy binary column. Examples: - A vaccination variable that pools one-dose and two-dose recipients. - A treatment flag for a discount that was 10 percent for some users and 30 percent for others. - An intervention delivered by different clinics with different fidelity. When versions differ in their effects, the estimand becomes an average over the version mix that happened to occur, and it will not transfer to a setting with a different mix. The fix is definitional: describe the intervention precisely enough that the label means one thing, or estimate version-specific effects. ## Relationship to consistency The second half of SUTVA is closely tied to the consistency assumption, which says the observed outcome equals the potential outcome under the treatment actually received. If there are several materially different versions hiding under one label, it is ambiguous which potential outcome the observed value corresponds to. Different textbooks split the two ideas differently; what matters in an interview is that you can state both requirements and recognise a violation, not which name a particular author gives to which half. ## How to check it There is no test. You reason about the mechanism: can units talk to each other, share a resource, compete, infect each other, or observe each other's treatment? Are they nested in households, teams, classrooms, or markets? And separately: does the treatment column collapse things that differ in kind or in dose? A candidate who asks those two questions unprompted is doing the right work.
- How does interference change what a potential outcome even means?It forces the outcome to be indexed by the whole assignment vector rather than by the unit's own assignment, so each unit has one potential outcome per possible configuration of everyone else. Since no sample can identify that many quantities, analysts restrict to a summary of others' exposure, such as the fraction of neighbours treated, and redefine the estimand accordingly.
- In a vaccination campaign, which way does herd immunity bias the naive comparison?Downward for the campaign's value. Unvaccinated people in the same community are partly protected by everyone else's vaccination, so their infection risk is lower than it would be with no campaign at all. The vaccinated-minus-unvaccinated contrast therefore captures only the extra individual protection and understates the total benefit of vaccinating the population.
- What does the second half of SUTVA rule out?Hidden versions of the treatment. Everyone recorded as treated must have received the same intervention in the sense that matters for the outcome. If the treated label pools one dose with two doses, or a 10 percent discount with a 30 percent one, the estimate is an average over an unrecorded mix and will not transfer to a setting where that mix differs.
saying these in an interview costs you the question
- States only the no-interference half and forgets treatment versions
- Assumes randomization removes spillover
- Says spillover always makes the effect look larger
- Treats units in the same household or team as independent by default
- Pools clearly different treatment intensities under one label