skip to content

Confounding Structures

How bias arises structurally: a common cause behind treatment and outcome, a control you should never have conditioned on, a rate that flips when groups pool. Interviewers ask which controls belong.

on this pageshow

explore

questions

16

What is Simpson's paradox, and why can the pooled rate reverse?

level: juniorimportance: must knowfreq 72%

answer

  1. a comparison that flips
  2. true in every slice, false combined
  3. a pooled rate is a weighted average
  4. the groups have different subgroup mixes
  5. 1973 Berkeley admissions, by department

basics

~20 s

Simpson's paradox is when a comparison that holds inside every subgroup reverses once the subgroups are pooled. It happens because the two groups sit in the subgroups in very different proportions, so the pooled rate weights those subgroups unequally.

solid answer

~50 s

Simpson's paradox is the reversal of a comparison under aggregation: group X beats group Y in every stratum, yet loses in the combined table. The arithmetic cause is weighting. A pooled rate is a weighted average of stratum rates, and the weights are how each group's observations are spread across strata. If most of X's observations sit in a stratum where nobody does well, X's pooled rate is dragged down even though X wins that stratum. The historical case is 1973 graduate admissions at UC Berkeley: about 44% of male applicants and about 35% of female applicants were admitted overall, yet department by department women were admitted at rates equal to or higher than men, because women applied disproportionately to departments that admitted few applicants of either sex. Both tables are correct arithmetic; choosing which one answers your question is a causal judgment.

go deeper

for a junior

Be ready to state the definition in one sentence and sketch a small table where the reversal actually happens. The point interviewers listen for is that both tables are arithmetically correct.

for a middle

Explain the mechanism, not just the name: a pooled rate is a weighted average of subgroup rates, and the weights are each group's subgroup composition. Show how unequal composition flips the ordering.

for a senior

Expect a reversed table handed to you with a decision attached. Name the confounding variable, say why the stratified comparison carries the causal reading, and offer a rate standardised to a common composition as the single headline number.

for a principal

Own the reporting policy: which breakdowns ship by default, when a headline metric must be composition-adjusted, and how you stop teams from quietly picking whichever aggregation flatters the result.

## The definition **Simpson's paradox** is the reversal of a comparison when data are aggregated. A comparison holds inside every subgroup of the data — group X does better than group Y among the small cases, and again among the large cases — yet when the subgroups are combined into one table, Y comes out ahead. Nothing has been miscalculated. The stratified rates and the pooled rate are both correct arithmetic on the same rows. ## A worked table Take an illustrative admissions example with two departments (round numbers, chosen to make the effect obvious). | | Dept A applicants | admitted | rate | Dept B applicants | admitted | rate | |---|---|---|---|---|---|---| | Men | 500 | 310 | 62% | 100 | 18 | 18% | | Women | 100 | 65 | 65% | 500 | 105 | 21% | Women beat men in Department A (65% vs 62%) and in Department B (21% vs 18%). Now pool: - Men: 328 admitted out of 600 applicants = **54.7%** - Women: 170 admitted out of 600 applicants = **28.3%** Women are ahead in both departments and 26 points behind overall. Same rows, opposite conclusion. ## Why it happens: a pooled rate is a weighted average The pooled rate for a group is its stratum rates weighted by how many of its observations fall in each stratum: ``` pooled(men) = (500/600) * 0.62 + (100/600) * 0.18 = 0.547 pooled(women) = (100/600) * 0.65 + (500/600) * 0.21 = 0.283 ``` The two groups are averaged with *different weights*. Men's average is dominated by the easy department; women's by the hard one. That is the entire mechanism. Two conditions must both hold for a reversal to be possible: 1. **The strata must differ in their base outcome rate** — Department A admits far more freely than Department B. 2. **The groups must be distributed differently across the strata** — men concentrate in A, women in B. Remove either condition and the reversal cannot occur. If both groups had the same departmental spread, the weights would match and the pooled ordering would follow the stratified one. ## The historical case The canonical example is graduate admissions at the University of California, Berkeley for the autumn of 1973. Roughly 44% of men who applied were admitted against roughly 35% of women, a gap large enough that it looked like clear evidence of bias against women. When the admissions statisticians broke the data out by department, the department-level rates showed no such pattern: in most departments women were admitted at the same rate as men or a slightly higher one. Women applied in much larger numbers to departments with low admission rates for everybody, and men to departments with high ones. The overall gap was produced by *where* the applications went, not by how applications were treated once they arrived. One honest nuance is worth carrying into an interview. Department is chosen by the applicant, so it sits downstream of the applicant's sex in time. The department-specific comparison therefore answers a narrow question — "within a department, were women admitted at a lower rate?" — and answers it in the negative. It does not answer why women applied more heavily to the crowded, low-admission departments, which is a real question about the pipeline that the stratified table simply does not address. ## What the paradox is not - **Not a sample-size artifact.** Multiply every cell in the table by ten and every rate is unchanged, so the reversal survives. More data does not dissolve it; it is a property of the composition of the data, not of sampling noise. - **Not an arithmetic error.** Candidates often insist one of the two tables must be wrong. Both are right; they answer different questions. - **Not the same as an unweighted average of subgroup rates.** Averaging 65% and 21% to get 43% for women is a separate mistake — it silently reweights the departments to be equally important, which is a decision, not a calculation. ## Why interviewers ask it It is the cheapest way to find out whether a candidate distinguishes a number from a conclusion. Anyone can compute both tables. The signal is whether the candidate says "I cannot pick between these without knowing what the stratifying variable is doing causally" instead of confidently reading one of them off. In real analytics work the same shape shows up whenever a headline rate is compared between two populations with different composition — hospitals with different case mixes, cohorts with different tenure profiles, regions with different age structures. ## The takeaway The paradox is arithmetic; the resolution is causal. Report the stratified table when the stratifier is a common cause of group membership and of the outcome. Report the pooled table when it is not, or when the marginal number is genuinely what you were asked for. And when you must give one headline number for populations with different composition, reweight both groups' stratum rates to a single common composition and say which composition you used.

  • Is one of the two tables simply computed wrong?
    No. Both are correct arithmetic on the same rows. The stratified table and the pooled table answer different questions, and the paradox is that the answers can point opposite ways. Deciding which one you want requires knowing how the stratifying variable relates causally to the grouping and to the outcome. The numbers by themselves never settle it.
  • In the Berkeley admissions case, is department a confounder or something else?
    Applicants choose a department, so department sits after sex in time rather than being a prior common cause. The department-specific comparison therefore answers a narrow question — within a department, were women admitted less often? — and the answer was no. It leaves untouched the separate question of why applications from women concentrated in departments that admitted few people.
  • Does collecting more data make Simpson's paradox go away?
    No. Scale every cell in the table by ten and every rate, stratified and pooled, is identical, so the reversal is untouched. It reflects composition, not sampling error. What removes it is a change of design or reporting: randomising which group gets which condition, or publishing rates standardised to one common subgroup composition.

Two students each outscore a rival on the easy exam and on the hard exam. If one sat mostly the hard exam and the other mostly the easy one, the rival can still finish with the higher overall percentage.

saying these in an interview costs you the question

  • Claims one of the two tables must be arithmetically wrong
  • Says a larger sample makes the reversal disappear
  • Averages the subgroup rates unweighted and calls it the overall rate
  • Treats the paradox as a rare textbook curiosity

context

open as a page

What is a backdoor path in a causal DAG, and what does the backdoor criterion require?

level: middleimportance: must knowfreq 80%

basics

~20 s

A backdoor path from treatment T to outcome Y is any path starting with an arrow into T; it carries non-causal association. A valid adjustment set contains no descendant of T and blocks every such path.

open as a page

Why does conditioning on a collider create an association between two independent causes?

level: middleimportance: must knowfreq 68%

basics

~20 s

A collider is a variable that two others both cause. Conditioning on it - filtering to a subgroup, stratifying, or adding it as a control - correlates those two causes in the data even when they are independent.

open as a page

Why is controlling for a post-treatment mediator called a bad control?

level: middleimportance: must knowfreq 58%

basics

~20 s

A mediator sits on the causal path from treatment to outcome, so adjusting for it removes the very effect you set out to measure. A real total effect shrinks toward zero, and new bias can appear.

open as a page

For an email campaign, do customer tenure, Black-Friday timing and clicking the email belong in the adjustment set for purchases?

level: seniorimportance: must knowfreq 58%

basics

~20 s

Tenure and Black-Friday timing each influence who was emailed and who buys, so both sit on backdoor paths and belong in the adjustment set. Clicking is caused by the campaign, so the criterion excludes it.

open as a page

A treatment beats its rival on small kidney stones and on large ones but loses overall, so which result do you report?

level: seniorimportance: must knowfreq 56%

basics

~20 s

Report the stone-size-specific results. Stone size influenced both which treatment a patient received and how likely success was, so the pooled comparison blends the treatment effect with differences in case severity. Only the size-specific comparison supports a causal reading.

open as a page

In a causal DAG, what does an arrow assert, and what does a missing arrow assert?

level: juniorimportance: should knowfreq 55%

basics

~10 s

An arrow from X to Y asserts X may directly cause Y. A missing arrow is the stronger claim: no direct effect at all. Acyclicity forbids a variable causing itself through a loop.

open as a page

In the causal DAG A -> B -> C, why does conditioning on B make A and C independent?

level: middleimportance: should knowfreq 62%

basics

~10 s

In the chain A -> B -> C, all of A's influence on C travels through B. Fixing B leaves no route, so A and C become conditionally independent. This is d-separation's chain-blocking rule.

open as a page

How can a batter beat a rival in each of two seasons yet trail on the combined average?

level: middleimportance: should knowfreq 50%

basics

~20 s

A combined batting average is total hits over total at-bats, so each season is weighted by its at-bats rather than counted equally. If a batter's strong season carries few at-bats and the rival's strong season carries many, the ordering flips.

open as a page

Why does maternal smoking look protective among low-birth-weight infants?

level: seniorimportance: should knowfreq 34%

basics

~20 s

Birth weight is a common effect of smoking and of other, more dangerous unmeasured causes. Restricting the analysis to low-birth-weight infants makes those causes trade off, so within that group smoking appears protective even though it raises risk overall.

open as a page

Is the stratified table always the right answer when a rate reverses after stratifying?

level: seniorimportance: should knowfreq 38%

basics

~20 s

No. Stratify only when the stratifier is a common cause of group and outcome, fixed before treatment. If treatment itself changes that variable, splitting on it hides part of the effect and the pooled table is right.

open as a page

How do you defend an adjustment set when the DAG behind it cannot be verified from data?

level: principalimportance: should knowfreq 38%

basics

~20 s

Defend it as a stated assumption, not a finding. Elicit the graph from whoever owns the assignment mechanism, test the independencies it implies, estimate under rival graphs, and report how strong a hidden confounder must be.

open as a page

How would you set a team policy for which covariates may enter a causal model?

level: principalimportance: should knowfreq 41%

basics

~10 s

Make exclusion the default: ban anything measured after treatment, require a stated causal reason for every covariate that stays, and fix the list before seeing the estimate. More controls is not more conservative.

open as a page

How does the front-door criterion identify a causal effect when the confounder is unmeasured?

level: seniorimportance: nice to knowfreq 32%

basics

~20 s

The front-door criterion uses a fully observed mediator carrying the entire effect. Estimate the treatment's effect on the mediator and the mediator's effect on the outcome, then chain them; the unmeasured confounder is never adjusted for.

open as a page

How can M-bias make a covariate measured before treatment a harmful control?

level: seniorimportance: nice to knowfreq 19%

basics

~20 s

M-bias arises when a pre-treatment variable is a common effect of two unmeasured causes: one also drives treatment, the other the outcome. Nothing is wrong until you adjust for it, and adjusting links those hidden causes.

open as a page

In Lord's paradox, why do a gain-score analyst and a baseline-adjusted analyst disagree?

level: principalimportance: nice to knowfreq 22%

basics

~20 s

They estimate different quantities under different untestable assumptions. One treats each person's own starting value as the fair baseline; the other compares people who started alike. With groups that differ at baseline, those comparisons can point opposite ways.

open as a page