What does it mean for two variables to be conditionally independent given a third?
answer
- independence inside a stratum
- fix the third variable, then check
- P(A given B and C) equals P(A given C)
- a common cause can manufacture association
- conditioning on a shared effect works the other way
basics
~10 sTwo variables are conditionally independent given a third when, once that third is fixed, they carry no information about each other: P(A and B given C) = P(A given C) times P(B given C).
solid answer
~50 s`A` and `B` are conditionally independent given `C` when `P(A and B | C) = P(A|C) * P(B|C)` for every value of `C` with positive probability, equivalently `P(A | B and C) = P(A|C)`: once you know `C`, learning `B` changes nothing about `A`. The textbook case is shoe size and reading ability in schoolchildren, strongly associated across a whole primary school because older children have bigger feet and read better, yet unrelated within any single age group. Age is a common cause of both. The crucial point is that conditional and unconditional independence are logically separate, and neither implies the other. Conditioning on a shared cause can remove dependence; conditioning on a shared effect can create it, as when two independent coin flips become perfectly dependent once you are told their parity bit says they differ.
go deeper
Recall the equation P(A and B given C) = P(A given C) times P(B given C), and one example where an association vanishes once you look within a single group.
Explain the equivalent form P(A given B and C) = P(A given C) and why a common cause produces association in pooled data. Be able to work through a stratified table.
Show that neither implication holds in general, including the case where conditioning on a shared effect creates dependence, and discuss stratum size and binning when testing the assumption on real data.
Own the decision of which conditional-independence assumptions a modelling approach rests on, how they are documented, and what evidence the organisation should require before treating one as settled.
## The definition For events `A`, `B` and a conditioning event `C` with `P(C) > 0`, conditional independence of `A` and `B` given `C` means ``` P(A and B | C) = P(A|C) * P(B|C) ``` Equivalently, whenever `P(B and C) > 0`, ``` P(A | B and C) = P(A|C) ``` That second form is the one to say out loud: **given `C`, the event `B` carries no additional information about `A`.** For random variables the same statement is made for every joint value: the conditional joint distribution factorises into the product of the conditional marginals, at each level of the conditioning variable. Notice that this is the ordinary definition of independence applied inside the conditional probability measure obtained by restricting the sample space to `C`. Everything you know about independence transfers, one stratum at a time. ## The common-cause picture Measure shoe size and reading ability for every child in a primary school. The scatter shows a clear positive association: children with bigger feet read better. Nothing about feet causes literacy. Age drives both — older children have larger feet and have had more years of reading instruction. Split the data by age and look within, say, the seven-year-olds: the association disappears. Shoe size and reading ability are conditionally independent given age, and dependent unconditionally. This is the structural reason stratified analysis exists. An association observed in a pooled population can be entirely manufactured by a variable that varies across the pool and drives both measurements. Conditioning on that variable is what removes it. ## Neither direction implies the other This is the part interviewers probe, and it is worth being explicit about both failures. **Conditional independence does not imply unconditional independence.** The shoe-size example is exactly that: independent within each age, dependent overall. Averaging the strata together reintroduces the association, because the conditioning variable itself moves both quantities. **Unconditional independence does not imply conditional independence.** Conditioning on a common *effect* creates dependence between previously independent causes. Flip two fair coins; the flips are independent. Now condition on the parity bit — you are told the two flips differ. Within that conditioning event, the first flip determines the second exactly: heads on the first forces tails on the second. Two independent variables became perfectly dependent purely by conditioning on something they jointly produce. So 'controlling for more variables' is not automatically a route to cleaner estimates. Conditioning on a shared cause removes a spurious association; conditioning on a shared effect creates one. ## Checking it against data To assess conditional independence empirically you stratify by the conditioning variable and compare the within-stratum association to the pooled one. Practical considerations dominate: - **Stratum size.** Slicing finely leaves each stratum with little data, so within-stratum associations get noisy and 'no association' can simply mean 'no power'. Absence of a detected association is not evidence of independence. - **Continuous conditioners.** Age in days cannot be stratified exactly; you bin it or model it. Coarse bins leave residual variation inside each bin, which leaves residual association — an artefact of the binning, not of the underlying structure. - **Multiple conditioners.** Conditional independence is always *given a specified set*. Independence given age may fail given age and school, or vice versa. Always name the conditioning set; an unqualified claim of conditional independence is meaningless. - **Direction of the check.** The two forms of the definition give two equivalent tests — compare the conditional joint to the product of conditional marginals, or compare `P(A | B and C)` to `P(A|C)`. The second is often easier to read off a table. ## Why it matters in practice Conditional independence is the assumption that makes large joint distributions tractable. Instead of estimating a probability for every combination of many variables, you assert that a variable depends only on a small conditioning set and estimate a manageable set of conditionals. Every such simplification is a conditional-independence claim, and every one of them can be wrong in a specific, checkable way. The professional habit is to write the assumption down explicitly, name the conditioning set, and say what observable pattern would falsify it. ## Answering it well Give the equation, give the equivalent 'no extra information' phrasing, give one common-cause example where dependence disappears under conditioning, and then state clearly that the implication runs in neither direction, with the conditioning-on-an-effect case as the counterexample. That combination — definition, example, and both failure directions — is a complete senior answer.
- Does conditional independence given C imply that A and B are independent overall?No. Shoe size and reading ability are independent within each age group yet clearly associated across a whole school, because age shifts both. Pooling strata reintroduces the association whenever the conditioning variable itself varies and moves both quantities. The two properties are logically separate.
- Can conditioning ever create dependence between independent events?Yes, when you condition on a common effect. Two fair coin flips are independent, but if you are told the parity bit says the flips differ, the first flip determines the second exactly. Adding a conditioning variable can therefore make an analysis worse, not better — it depends on what that variable is.
- How would you check conditional independence in observed data?Stratify by the conditioning variable and compare the within-stratum association to the pooled one. Watch stratum sizes, since thin strata give no power and can fake independence, and remember that binning a continuous conditioner leaves residual variation inside each bin, which leaves residual association.
- Why is an unqualified claim of conditional independence meaningless?Because the property is always relative to a named conditioning set. Two variables can be conditionally independent given age but dependent given age and school, or the reverse. Stating the set is part of stating the assumption, and it is what makes the claim falsifiable against data.
Two clocks across town appear to move together until you notice both are set by the same radio signal; hold that signal fixed and each ticks on its own.
saying these in an interview costs you the question
- Treats conditional independence as ordinary independence
- Assumes a pooled association must hold within every subgroup
- Says conditioning on more variables is always safer
- Claims conditional independence without naming the conditioning set
- Reads a noisy within-stratum result as proof of independence