skip to content

Causal Inference

You will learn to estimate cause and effect without an experiment: recognize confounding, draw a DAG, and apply propensity scores, diff-in-diff, IVs, or regression discontinuity. Interviewers probe it with 'the metric moved — did the feature cause it?' scenarios, where correlation-only answers fail.

on this pageshow

explore

questions

68 · 5 sections

In causal inference, what does conditional ignorability require of treatment assignment?

level: middleimportance: must knowfreq 70%
basics
~10 s

Conditional ignorability requires that within every level of the measured covariates X, which units received treatment is independent of their potential outcomes. Treatment must be as good as randomly assigned inside each X stratum.

open as a page

In causal inference, what does the positivity assumption require of every covariate stratum?

level: middleimportance: must knowfreq 60%
basics
~20 s

Positivity requires every covariate stratum in the target population to contain both treated and untreated units: 0 < P(T=1 | X=x) < 1. A stratum with only one condition supplies no comparison, so no effect is identified there.

open as a page

What is the difference between ATE, ATT and ATU as treatment-effect estimands?

level: middleimportance: must knowfreq 70%
basics
~20 s

ATE averages the effect Y(1) - Y(0) over the whole population, ATT over only the units that actually got treated, and ATU over the untreated ones. ATE is their size-weighted mix, and the three can differ in sign.

open as a page

What is the fundamental problem of causal inference in the potential outcomes framework?

level: middleimportance: must knowfreq 78%
basics
~20 s

Each unit has two potential outcomes, one under treatment and one under control, but only the one matching the arm it actually received is ever observed. The other stays missing, so an individual causal effect is never measured directly.

open as a page

What does SUTVA, the stable unit treatment value assumption, require in a causal study?

level: juniorimportance: should knowfreq 48%
basics
~10 s

SUTVA has two parts: one unit's treatment does not affect another unit's outcome, and there is only one version of the treatment, so every unit labelled treated received effectively the same thing.

open as a page

What is Simpson's paradox, and why can the pooled rate reverse?

level: juniorimportance: must knowfreq 72%
basics
~20 s

Simpson's paradox is when a comparison that holds inside every subgroup reverses once the subgroups are pooled. It happens because the two groups sit in the subgroups in very different proportions, so the pooled rate weights those subgroups unequally.

open as a page

What is a backdoor path in a causal DAG, and what does the backdoor criterion require?

level: middleimportance: must knowfreq 80%
basics
~20 s

A backdoor path from treatment T to outcome Y is any path starting with an arrow into T; it carries non-causal association. A valid adjustment set contains no descendant of T and blocks every such path.

open as a page

Why does conditioning on a collider create an association between two independent causes?

level: middleimportance: must knowfreq 68%
basics
~20 s

A collider is a variable that two others both cause. Conditioning on it - filtering to a subgroup, stratifying, or adding it as a control - correlates those two causes in the data even when they are independent.

open as a page

Why is controlling for a post-treatment mediator called a bad control?

level: middleimportance: must knowfreq 58%
basics
~20 s

A mediator sits on the causal path from treatment to outcome, so adjusting for it removes the very effect you set out to measure. A real total effect shrinks toward zero, and new bias can appear.

open as a page

For an email campaign, do customer tenure, Black-Friday timing and clicking the email belong in the adjustment set for purchases?

level: seniorimportance: must knowfreq 58%
basics
~20 s

Tenure and Black-Friday timing each influence who was emailed and who buys, so both sit on backdoor paths and belong in the adjustment set. Clicking is caused by the campaign, so the criterion excludes it.

open as a page

What is a propensity score, and why match on it rather than on the raw covariates?

level: juniorimportance: must knowfreq 76%
basics
~20 s

A propensity score is a unit's probability of receiving the treatment given its observed covariates. Conditioning on that single number balances the covariates that went into it, so matching happens in one dimension instead of many.

open as a page

What does adding a confounder to a regression do to the treatment coefficient?

level: juniorimportance: must knowfreq 76%
basics
~20 s

Adding a measured confounder turns the treatment coefficient from a raw comparison into a within-strata one: it now compares treated and untreated units that share the same covariate value, removing the part of the gap that covariate explained.

open as a page

How does inverse probability of treatment weighting estimate an average treatment effect?

level: middleimportance: must knowfreq 68%
basics
~20 s

Each unit is weighted by the inverse of its probability of receiving the arm it actually got: treated units by 1/e(X), controls by 1/(1-e(X)). The reweighted sample mimics a population where treatment was assigned independently of the measured covariates.

open as a page

How do you check covariate balance after propensity score matching?

level: middleimportance: must knowfreq 82%
basics
~20 s

Compare standardized mean differences for each covariate before and after matching, with an absolute value under about 0.1 as the usual bar, and read them off a love plot. Check variance ratios and interactions too, not just means.

open as a page

How do you sign the omitted-variable bias when ability is left out of a wage-on-schooling regression?

level: middleimportance: must knowfreq 68%
basics
~20 s

Multiply two signs: the effect of the omitted variable on the outcome, times its correlation with the included regressor. Ability raises wages and is positively correlated with schooling, so the schooling coefficient is biased upward.

open as a page

What does a difference-in-differences estimate compute from a two-group, two-period panel?

level: juniorimportance: must knowfreq 70%
basics
~20 s

It subtracts the comparison group's before-to-after change from the treated group's before-to-after change. That double difference cancels the fixed level gap between the groups and any shock that moved both of them over the same window.

open as a page

How does two-stage least squares turn an instrument into a causal estimate?

level: middleimportance: must knowfreq 62%
basics
~20 s

Two-stage least squares first predicts the treatment from the instrument, then relates the outcome to that prediction. Only the instrument-driven part of the treatment's variation is used, and that part is assumed free of the confounding.

open as a page

What assumptions must an instrumental variable satisfy to identify a causal effect?

level: middleimportance: must knowfreq 72%
basics
~20 s

A valid instrument needs relevance, meaning it actually shifts the treatment, and the exclusion restriction, meaning it affects the outcome only through that treatment. It must also be as good as randomly assigned with respect to unmeasured confounders.

open as a page

In a sharp regression discontinuity, what assumption identifies the effect and what does it estimate?

level: middleimportance: must knowfreq 70%
basics
~20 s

Identification rests on continuity: average outcomes with and without treatment must vary smoothly through the cutoff, so nothing but treatment jumps there. The estimate is the average treatment effect for units sitting at the cutoff, not for everyone.

open as a page

Why can a treatment with an average effect of zero still help some users and harm others?

level: juniorimportance: must knowfreq 62%
basics
~20 s

An average pools opposite effects. A drug that raises recovery for patients under 50 and lowers it by a similar amount for patients over 70 posts an overall average near zero while both real effects remain untouched underneath it.

open as a page

What does a sensitivity analysis for unmeasured confounding tell you about a causal estimate?

level: juniorimportance: must knowfreq 55%
basics
~10 s

A sensitivity analysis says how strong an unmeasured confounder would have to be to overturn the estimate. It never shows confounding is absent; it prices how much hidden bias the finding can tolerate.

open as a page

How do negative-control outcomes and exposures expose residual confounding in an observational study?

level: seniorimportance: must knowfreq 48%
basics
~20 s

You rerun the analysis on a relationship that must be null: an outcome the treatment cannot cause, or an exposure that cannot cause the outcome, both sharing the study's confounding. Finding an effect where none can exist proves bias remains.

open as a page

How do the S-, T- and X-learner meta-learners differ when estimating conditional treatment effects?

level: middleimportance: should knowfreq 48%
basics
~20 s

The S-learner fits one outcome model with treatment as a feature and differences its predictions. The T-learner fits separate treated and control models and subtracts them. The X-learner adds a stage that models imputed per-unit effects from each side and blends them.

open as a page

Why does targeting a retention discount by churn risk differ from targeting by uplift?

level: middleimportance: should knowfreq 55%
basics
~20 s

A churn-risk model ranks who will leave; an uplift model ranks whose behaviour the discount changes. Those are different people: the highest-risk customers often leave regardless, and some contented ones cancel only because the offer reminded them the subscription exists.

open as a page