skip to content

In causal inference, what does the positivity assumption require of every covariate stratum?

level: middleimportance: must knowfreq 60%

answer

  1. no comparison, no effect
  2. both conditions must occur in each stratum
  3. probability strictly between zero and one
  4. rule-based empty cell versus sparse cell
  5. narrow the estimand to where overlap exists

basics

~20 s

Positivity requires every covariate stratum in the target population to contain both treated and untreated units: 0 < P(T=1 | X=x) < 1. A stratum with only one condition supplies no comparison, so no effect is identified there.

solid answer

~50 s

Positivity, also called overlap or common support, is the assumption that for every value of X in the population you want an answer about, the probability of treatment is strictly between zero and one. It matters because identification is local: the causal contrast at `X = x` is built from treated and untreated units that share that x, and if one side is empty there is nothing to compare. The important distinction is structural versus random violations. If a guideline means no patient over 75 with kidney disease ever received the surgery, that cell is empty by rule and no volume of data fills it; any number a model reports there comes from the functional form, not from evidence. A finite-sample empty cell that could in principle be filled is a precision problem instead. The honest response to a structural violation is to redefine the estimand to the population where overlap actually exists and say so.

go deeper

for a junior

Know that a causal comparison needs both treated and untreated units with similar characteristics, and that a group where everyone got the same thing supports no comparison at all.

for a middle

Be ready to state the condition on the treatment probability, explain why an empty cell forces the model to extrapolate, and separate rule-based emptiness from small-sample sparsity.

for a senior

Demonstrate that you inspect overlap before estimating, and that you would rather narrow the estimand and name the population it covers than report a number from a region with no data.

for a principal

Own the tradeoff between a richer adjustment set and shrinking overlap, and set the expectation that analyses declare the subpopulation their causal claim applies to.

## The condition Positivity, sometimes called overlap or common support, requires `0 < P(T = 1 | X = x) < 1` for every x with positive probability in the target population. In words: whatever combination of covariates you look at, some units in that combination could have been treated and some could have been untreated. It is the assumption that pairs with conditional ignorability. Ignorability says a comparison inside a stratum would be fair; positivity says such a comparison exists at all. ## Why identification is local An average causal effect over a population is an average of stratum-level contrasts, weighted by how common each stratum is. If a stratum contributes weight to the average but contains only treated units, the contrast for that stratum has to come from somewhere other than data. A regression will happily supply one by extrapolating its functional form into a region it has never seen. The output looks like an estimate, has a standard error, and is entirely a statement about the model. ## Structural versus random violations **Structural (deterministic) violations** occur when a rule, contraindication or eligibility criterion makes the probability exactly 0 or 1 for some x. If clinical guidance forbids the surgery for patients over 75 who also have kidney disease, then in that cell `P(T = 1 | X = x) = 0` by construction. Collecting ten times the data changes nothing, because the mechanism generating the data will never place a treated unit there. The effect of surgery in that group is simply not identified from this population, and the only correct answers are to change the estimand or to change the data source. **Random (practical) violations** occur when the probability is genuinely between 0 and 1 but the sample happens to contain few or no units of one kind in a stratum, usually because X is finely sliced. More data, or a coarser stratification, can fix these. They show up as very large variance and as a handful of units dominating the estimate. The test that separates them is a subject-matter question, not a statistical one: could a unit with these characteristics ever have received the other condition? Ask the person who runs the process. ## How violations surface - Estimated propensities pile up near 0 or 1, so a few units carry enormous influence. - The estimate swings wildly when a small number of observations are dropped. - Confidence intervals widen dramatically as X grows richer, which is the price of demanding comparisons in ever-thinner cells. - Covariate distributions for treated and untreated barely overlap in one or more dimensions. ## Responses The repair that is always available is to change the question rather than the arithmetic. Restrict the estimand to the region of covariate space where both conditions occur, and report explicitly who the answer is about: not the effect for everyone, but the effect for the population where the treatment decision was genuinely in play. That is a smaller claim, and it is a claim the data can support. Coarsening X, or conditioning on fewer variables where the mechanism argument permits it, can also restore overlap, at the cost of leaning harder on ignorability with a smaller adjustment set. The tension is real: a richer X makes ignorability more credible and positivity less credible at the same time, and choosing where to sit on that curve is judgment, not a formula. ## Randomization and positivity In an experiment where everyone eligible is assigned with a probability strictly inside (0, 1), positivity holds by construction within the eligible population. But eligibility criteria still bound who the answer applies to: nobody excluded from the study has overlap, so no causal claim about them is identified either. Positivity is therefore not only an observational concern; it is also the formal statement of who an experiment can speak for. ## What interviewers listen for The distinction between structural and random violations, the willingness to say a quantity is not identified rather than reporting a number, and the recognition that the estimand can be narrowed to match the data instead of the data being tortured to match the estimand.

  • How do you tell a structural positivity violation from a random one?
    Ask whether a unit with those characteristics could ever have received the other condition. If a rule, contraindication or eligibility criterion forbids it, the probability is exactly zero and more data cannot help: the violation is structural. If it is simply an unlucky empty cell in a finely sliced covariate space, more data or coarser strata resolve it.
  • What happens to your estimate if you ignore a near-violation and fit the model anyway?
    The model extrapolates. Numbers for sparse cells come from the assumed functional form rather than from any observed comparison, variance inflates, and a few units drive the answer. The result is an estimate that looks precise enough to report but changes materially when a handful of rows are removed.
  • Why can a richer covariate set make positivity worse?
    Adding covariates slices the population into thinner cells, and thin cells are more likely to contain only one condition. That is the standing tension: more X strengthens the case for ignorability while weakening the case for overlap. The choice of adjustment set has to be made with both assumptions in view.

It is like asking whether a medicine helps left-handed pilots when no left-handed pilot has ever been given it. However careful your arithmetic, there is nobody to compare against.

saying these in an interview costs you the question

  • Says more data always fixes an empty stratum
  • Reports a population-wide effect despite zero overlap in some cells
  • Confuses positivity with equal treated and control group sizes
  • Accepts model extrapolation into unobserved cells as evidence
  • Checks overlap only marginally, never within covariate combinations

context