skip to content

Why does conditioning on a collider create an association between two independent causes?

level: middleimportance: must knowfreq 68%

answer

  1. two arrows in, not out
  2. explaining away competing causes
  3. the filter creates the correlation
  4. hospital inpatients, dating pool
  5. bias made by the analyst

basics

~20 s

A collider is a variable that two others both cause. Conditioning on it - filtering to a subgroup, stratifying, or adding it as a control - correlates those two causes in the data even when they are independent.

solid answer

~50 s

A collider is a variable with two arrows pointing into it: both `X` and `Y` cause `C`. `X` and `Y` can be independent in the population, but once you condition on `C` - restrict the rows, stratify, or put `C` in a regression - knowing `X` tells you something about `Y`, because a high value of `C` has to be explained by one cause or the other. This is the explaining-away effect. Berkson's paradox is the classic case: two unrelated diseases show a negative association among hospital inpatients, because either disease alone can get you admitted, so an inpatient who does not have one is more likely to have the other. The same structure answers why the attractive people you meet seem so dull: if either trait can get someone into your dating pool, the two trade off inside the pool. The bias is manufactured by the analyst, not present in nature.

code

python · 18 lines
python
import random, statistics

def corr(xs, ys):
    mx, my = statistics.fmean(xs), statistics.fmean(ys)
    cov = sum((x - mx) * (y - my) for x, y in zip(xs, ys))
    dx = sum((x - mx) ** 2 for x in xs) ** 0.5
    dy = sum((y - my) ** 2 for y in ys) ** 0.5
    return cov / (dx * dy)

random.seed(0)
looks = [random.gauss(0, 1) for _ in range(20000)]
charm = [random.gauss(0, 1) for _ in range(20000)]
print("population:", round(corr(looks, charm), 3))          # about 0.005

pool = [(a, b) for a, b in zip(looks, charm) if a + b > 1.5]
print("pool size:", len(pool))                              # about 2959
print("inside pool:", round(corr([a for a, _ in pool],
                                 [b for _, b in pool]), 3))  # about -0.67

go deeper

for a junior

Be able to say what a collider is - a variable that two others both cause - and give one concrete example where filtering on it invents a correlation.

for a middle

Explain the explaining-away mechanism and work a small numeric case out loud. Interviewers at this level want the arithmetic, not just the slogan.

for a senior

Show you catch it in real analyses: name the point in a pipeline where a filter, a cohort definition, or a strongly predictive control silently conditions on a common effect.

for a principal

Own the tradeoff that no data-driven diagnostic separates a collider from a confounder, and describe how your team encodes that structural knowledge before analysis rather than after a surprising result.

## The structure Call a variable a **collider** when two arrows point into it - that is, when it is a *common effect* of two other variables. If smoking causes lung damage and asbestos exposure causes lung damage, then lung damage is a collider of smoking and asbestos. The word describes a position in a causal structure, not a property of the variable itself: the same variable is a collider on one triple and an ordinary cause on another. The key fact is counter-intuitive and is the reason interviewers ask about it. Two causes of a common effect may be completely independent of each other. **Conditioning on their shared effect makes them dependent.** Conditioning means any of: filtering the dataset to rows where the collider takes some value, stratifying and reading results within a stratum, adding the collider as a control in a regression, or - the sneakiest version - having a sample that only contains units selected on the collider in the first place. ## Why the dependence appears The intuition is *explaining away*. If an alarm rings when either a burglar or an earthquake occurs, and you are told the alarm rang, then learning that an earthquake occurred lowers your belief in a burglar. Neither event causes the other; the observed alarm forces them to compete as explanations. Here is the arithmetic, with nothing hidden. Let `X` and `Y` be independent fair coin flips, each 1 with probability 0.5. Keep a row whenever `X = 1` or `Y = 1`. In the full population the four cells (1,1), (1,0), (0,1), (0,0) each have probability 0.25, and `P(Y = 1 | X = 1) = P(Y = 1 | X = 0) = 0.5` - independence. Among kept rows only three cells survive, each with probability 1/3. Now `P(Y = 1 | X = 1)` uses the cells (1,1) and (1,0), giving 0.5, while `P(Y = 1 | X = 0)` can only be the cell (0,1), giving 1.0. Inside the selected data, `X` and `Y` are strongly negatively associated. Nothing about the world changed; the filter did all the work. ## Berkson's paradox Joseph Berkson's original point was about hospital data. Suppose two diseases are unrelated in the general population, and each one independently raises the chance of being admitted to hospital. Study only inpatients and the two diseases will appear negatively associated: an inpatient who does not have the first disease is admitted for some other reason, which makes the second disease more likely. A researcher reading only the inpatient file could easily publish the finding that disease A protects against disease B. The dating-pool version is the same shape in everyday clothing. Suppose looks and personality are independent in the population, and suppose someone enters your dating pool if they are high enough on the sum of the two. Inside the pool, the two traits are negatively correlated - the very attractive people you meet are disproportionately the ones who got in on looks alone. A simulation makes this vivid: independent traits with a population correlation near zero can show a correlation around -0.6 once you keep only the people above a threshold on their sum. ## Why this matters for controls The practical damage is that a collider looks empirically identical to a variable you *should* adjust for. It is correlated with one cause, correlated with the other, and adding it often improves model fit. But a **confounder** is a common *cause* of two variables and conditioning on it removes bias, whereas a **collider** is a common *effect* and conditioning on it creates bias. No amount of staring at correlations distinguishes the two - that distinction lives in subject-matter knowledge about which variable causes which. A few practical consequences: - **More controls is not safer.** Every added control is a decision that can create bias as easily as remove it. - **Sample selection is silent conditioning.** If units entered your dataset through a process that both variables affect, you are conditioning on a collider before you write a single line of analysis. - **The induced sign is usually, but not always, negative** when both causes push the collider in the same direction. The magnitude grows with how strongly the two causes drive the collider and how severe the selection is. The defence is structural rather than statistical: ask, for every variable you filter on, stratify by, or control for, whether the treatment and the outcome (or their causes) both affect it. If they do, leaving it alone is the correct move.

  • How does a collider differ from a confounder in what you should do with it?
    A confounder is a common cause of both variables, and conditioning on it removes bias. A collider is a common effect, and conditioning on it creates bias that was not there. The two can look identical in the correlation table - both are associated with treatment and outcome - so only knowledge of which variable causes which tells you whether adjustment helps or hurts.
  • Can collider bias enter through how data was collected rather than through a control variable?
    Yes, and that is the harder version to spot. If units enter the dataset only when some variable crosses a threshold, and both quantities you are studying push that variable, the dataset is already conditioned on a collider before any model runs. Volunteer panels, files of people who reached a certain state, and any 'we only kept rows where...' filter all qualify.
  • Is the induced association always negative?
    No. Negative is the common case when both causes raise the collider and you select on high values, because the causes then compete as explanations. Flip the direction of one cause, or select on low values, and the induced association can be positive. The sign follows from the structure, so reason it out rather than assuming.

An alarm rings if either a burglar or an earthquake occurs. Burglars and earthquakes are unrelated, but once you know the alarm rang, hearing that there was an earthquake makes you far less worried about a burglar.

saying these in an interview costs you the question

  • Says adding more control variables always reduces bias
  • Calls anything correlated with treatment and outcome a confounder
  • Assumes randomization makes any later conditioning safe
  • Thinks the association among selected units is a real population fact
  • Believes model fit improving proves the control belongs

context