Causal Inference
You will learn to estimate cause and effect without an experiment: recognize confounding, draw a DAG, and apply propensity scores, diff-in-diff, IVs, or regression discontinuity. Interviewers probe it with 'the metric moved — did the feature cause it?' scenarios, where correlation-only answers fail.
on this pageshowhide
explore
- Potential Outcomes10 questions
- Treatment Effects5 questions
- Identification Assumptions5 questions
- Confounding Structures16 questions
- Simpson's Paradox5 questions
- Backdoor Criterion6 questions
- Colliders and Bad Controls5 questions
- Adjustment Estimators15 questions
- Propensity Score Matching5 questions
- Inverse Probability Weighting5 questions
- Regression Adjustment5 questions
- Quasi-Experimental Designs17 questions
- Difference-in-Differences6 questions
- Instrumental Variables5 questions
- Regression Discontinuity6 questions
- Heterogeneity and Credibility10 questions
- Heterogeneous Treatment Effects5 questions
- Sensitivity and Validity Checks5 questions
questions
68 · 5 sectionsIn causal inference, what does conditional ignorability require of treatment assignment?
basics
~10 sConditional ignorability requires that within every level of the measured covariates X, which units received treatment is independent of their potential outcomes. Treatment must be as good as randomly assigned inside each X stratum.
In causal inference, what does the positivity assumption require of every covariate stratum?
basics
~20 sPositivity requires every covariate stratum in the target population to contain both treated and untreated units: 0 < P(T=1 | X=x) < 1. A stratum with only one condition supplies no comparison, so no effect is identified there.
What is the difference between ATE, ATT and ATU as treatment-effect estimands?
basics
~20 sATE averages the effect Y(1) - Y(0) over the whole population, ATT over only the units that actually got treated, and ATU over the untreated ones. ATE is their size-weighted mix, and the three can differ in sign.
What is the fundamental problem of causal inference in the potential outcomes framework?
basics
~20 sEach unit has two potential outcomes, one under treatment and one under control, but only the one matching the arm it actually received is ever observed. The other stays missing, so an individual causal effect is never measured directly.
What does SUTVA, the stable unit treatment value assumption, require in a causal study?
basics
~10 sSUTVA has two parts: one unit's treatment does not affect another unit's outcome, and there is only one version of the treatment, so every unit labelled treated received effectively the same thing.
What is Simpson's paradox, and why can the pooled rate reverse?
basics
~20 sSimpson's paradox is when a comparison that holds inside every subgroup reverses once the subgroups are pooled. It happens because the two groups sit in the subgroups in very different proportions, so the pooled rate weights those subgroups unequally.
What is a backdoor path in a causal DAG, and what does the backdoor criterion require?
basics
~20 sA backdoor path from treatment T to outcome Y is any path starting with an arrow into T; it carries non-causal association. A valid adjustment set contains no descendant of T and blocks every such path.
Why does conditioning on a collider create an association between two independent causes?
basics
~20 sA collider is a variable that two others both cause. Conditioning on it - filtering to a subgroup, stratifying, or adding it as a control - correlates those two causes in the data even when they are independent.
Why is controlling for a post-treatment mediator called a bad control?
basics
~20 sA mediator sits on the causal path from treatment to outcome, so adjusting for it removes the very effect you set out to measure. A real total effect shrinks toward zero, and new bias can appear.
For an email campaign, do customer tenure, Black-Friday timing and clicking the email belong in the adjustment set for purchases?
basics
~20 sTenure and Black-Friday timing each influence who was emailed and who buys, so both sit on backdoor paths and belong in the adjustment set. Clicking is caused by the campaign, so the criterion excludes it.
What is a propensity score, and why match on it rather than on the raw covariates?
basics
~20 sA propensity score is a unit's probability of receiving the treatment given its observed covariates. Conditioning on that single number balances the covariates that went into it, so matching happens in one dimension instead of many.
What does adding a confounder to a regression do to the treatment coefficient?
basics
~20 sAdding a measured confounder turns the treatment coefficient from a raw comparison into a within-strata one: it now compares treated and untreated units that share the same covariate value, removing the part of the gap that covariate explained.
How does inverse probability of treatment weighting estimate an average treatment effect?
basics
~20 sEach unit is weighted by the inverse of its probability of receiving the arm it actually got: treated units by 1/e(X), controls by 1/(1-e(X)). The reweighted sample mimics a population where treatment was assigned independently of the measured covariates.
How do you check covariate balance after propensity score matching?
basics
~20 sCompare standardized mean differences for each covariate before and after matching, with an absolute value under about 0.1 as the usual bar, and read them off a love plot. Check variance ratios and interactions too, not just means.
How do you sign the omitted-variable bias when ability is left out of a wage-on-schooling regression?
basics
~20 sMultiply two signs: the effect of the omitted variable on the outcome, times its correlation with the included regressor. Ability raises wages and is positively correlated with schooling, so the schooling coefficient is biased upward.
What does a difference-in-differences estimate compute from a two-group, two-period panel?
basics
~20 sIt subtracts the comparison group's before-to-after change from the treated group's before-to-after change. That double difference cancels the fixed level gap between the groups and any shock that moved both of them over the same window.
What does the parallel-trends assumption in difference-in-differences actually require?
basics
~20 sParallel trends requires that, absent the treatment, the treated and comparison groups' average outcomes would have changed by the same amount. It is a claim about an unobserved counterfactual, so it can never be verified directly.
How does two-stage least squares turn an instrument into a causal estimate?
basics
~20 sTwo-stage least squares first predicts the treatment from the instrument, then relates the outcome to that prediction. Only the instrument-driven part of the treatment's variation is used, and that part is assumed free of the confounding.
What assumptions must an instrumental variable satisfy to identify a causal effect?
basics
~20 sA valid instrument needs relevance, meaning it actually shifts the treatment, and the exclusion restriction, meaning it affects the outcome only through that treatment. It must also be as good as randomly assigned with respect to unmeasured confounders.
In a sharp regression discontinuity, what assumption identifies the effect and what does it estimate?
basics
~20 sIdentification rests on continuity: average outcomes with and without treatment must vary smoothly through the cutoff, so nothing but treatment jumps there. The estimate is the average treatment effect for units sitting at the cutoff, not for everyone.
Why can a treatment with an average effect of zero still help some users and harm others?
basics
~20 sAn average pools opposite effects. A drug that raises recovery for patients under 50 and lowers it by a similar amount for patients over 70 posts an overall average near zero while both real effects remain untouched underneath it.
What does a sensitivity analysis for unmeasured confounding tell you about a causal estimate?
basics
~10 sA sensitivity analysis says how strong an unmeasured confounder would have to be to overturn the estimate. It never shows confounding is absent; it prices how much hidden bias the finding can tolerate.
How do negative-control outcomes and exposures expose residual confounding in an observational study?
basics
~20 sYou rerun the analysis on a relationship that must be null: an outcome the treatment cannot cause, or an exposure that cannot cause the outcome, both sharing the study's confounding. Finding an effect where none can exist proves bias remains.
How do the S-, T- and X-learner meta-learners differ when estimating conditional treatment effects?
basics
~20 sThe S-learner fits one outcome model with treatment as a feature and differences its predictions. The T-learner fits separate treated and control models and subtracts them. The X-learner adds a stage that models imputed per-unit effects from each side and blends them.
Why does targeting a retention discount by churn risk differ from targeting by uplift?
basics
~20 sA churn-risk model ranks who will leave; an uplift model ranks whose behaviour the discount changes. Those are different people: the highest-risk customers often leave regardless, and some contented ones cancel only because the offer reminded them the subscription exists.