A teammate reads a high mutual information score as evidence that the signal causes fraud — what is wrong?
answer
- same number in both directions
- dependence priced, direction not
- confounder and leakage score alike
- top scorer sits at the ceiling
- ask when the value is written
basics
~20 sMutual information is symmetric and carries no direction: swapping its two arguments returns the same number. A high score is evidence of shared bits only, and is equally consistent with influence, a shared upstream driver, or a column written after the outcome.
solid answer
~40 sThe definition gives `I(X;Y) = I(Y;X)`, so nothing in the measure marks one variable as the earlier or the source. A high score has at least three explanations: the signal genuinely influences the outcome; a third factor moves both; or the column is recorded downstream of the outcome, which is leakage. The third case is the one that bites, because a field written by the same process that assigns the label carries almost all of it and tends to top the ranking outright. The correction is not a bigger information quantity — conditioning on a third variable leaves the measure symmetric in the first two. It is a timeline check (when is this value known, relative to the label?) plus, where you can get it, an intervention or a time-ordered holdout.
go deeper
Remember that a dependence score is symmetric: it reports that two columns share bits, never which one came first. Treat any causal reading of it as an assumption you have added, not something the number told you.
Explain the symmetry from the definition itself, and name the three patterns that all produce a high value: genuine influence, a common upstream driver, and a column derived from the outcome.
Demonstrate the reflex a practitioner has: the suspiciously top-scoring candidate gets a timeline check before anything else, and the evaluation is time-ordered with every field frozen to its value at the decision moment.
Own the process question. Decide what evidence a team must present before a signal is allowed into production — timing provenance, an intervention where affordable — and who is accountable when an upstream writer quietly changes when a field is populated.
## The measure has no arrow in it **Mutual information** is defined as `I(X;Y) = H(X) - H(X|Y)`, the bits of uncertainty about one variable that observing the other removes. The algebra immediately gives the same value from the other side: `H(X) - H(X|Y) = H(Y) - H(Y|X)`. Equivalently, the symmetric form `I(X;Y) = sum over x,y of p(x,y) * log2( p(x,y) / (p(x) * p(y)) )` treats the two variables identically — swap their roles and every term is unchanged. Whatever a large score is telling you, it is telling you the same thing read in either direction. There is no computation you can perform on that number alone that recovers which variable came first. This is not fixed by reaching for a richer quantity. **Conditional mutual information** `I(X;Y|Z)`, which prices the bits `X` and `Y` share once a third variable is held fixed, is still symmetric in its first two arguments. It can tell you that a dependence disappears once you condition on something else — a genuinely useful fact — but it never labels one side cause and the other effect. ## Three patterns that all score high 1. **Direct influence.** The signal really is part of what produces the outcome. This is the case the teammate has in mind, and it is one of several. 2. **A shared upstream driver.** Some third factor moves both the signal and the label. Both columns then carry bits about that factor, so they carry bits about each other, with no influence running between them at all. Remove or hold the driver fixed and the dependence collapses. 3. **The reverse direction, recorded as a column.** A field that is written after the outcome is decided — a status, a queue assignment, a note added during a review — is downstream of the label. It shares bits with it precisely because it was derived from it. | Pattern | Typical score | What swapping the arguments reveals | |---|---|---| | Signal influences outcome | Moderate to high | Nothing; the number is identical | | Shared upstream driver | Can be very high | Nothing; the number is identical | | Column derived from the outcome | Near the label's entropy ceiling | Nothing; the number is identical | The third column is the whole point: the measure cannot separate the rows. ## Why the top of the ranking is the suspicious end The score is capped: `I(X;Y)` can never exceed `H(X)` or `H(Y)`, so its ceiling against a label is the label's own entropy. A column that was computed from the outcome sits right at that ceiling, because knowing it removes essentially all the label's uncertainty. Genuine predictive signals almost never do that. So a candidate that scores dramatically above everything else on the list is, more often than not, a record of the answer rather than a clue to it — and the ranking presents it as the best thing you have. The practical check is a **timeline check**, and it is a question about the data-producing process rather than about the numbers: - At the moment the score has to be produced, does this field already have a value, or is it still null? - Who writes it, and is that writer upstream or downstream of whatever decides the label? - Does the field's value ever change after the label is assigned? A time-ordered holdout — fit on earlier rows, score on later ones, with every field frozen to the value it had at the decision moment — turns those questions into a measurement. ## What does establish direction Nothing inside information theory does it alone. Direction comes from assumptions you bring: - **Time.** A value that provably exists before the outcome cannot have been derived from it. This is the cheapest and most commonly available assumption. - **Intervention.** Change the signal deliberately, at random, and see whether the outcome distribution moves. This is the strongest evidence and usually the most expensive to obtain. - **Domain structure.** Knowing how the pipeline is wired often rules out one direction outright. A dependence score is a filter that proposes candidates for those checks. It is not itself one of them. ## What an interviewer is listening for The expected answer is short and specific: the measure is symmetric, so it prices dependence and not direction; a high value is consistent with influence, confounding, or leakage; and the practically important case in a signal-ranking exercise is leakage, caught by asking when the value becomes available rather than by computing something further. A candidate who says "correlation is not causation" and stops has stated the slogan without the mechanism — what earns the point is naming the symmetry in the definition and the leakage signature at the entropy ceiling.
- Which candidate usually tops such a ranking, and why is that a warning sign?Often a field written by the same process that settles the outcome — a review status, a closing note, a downstream flag. It shares nearly all the label's bits, so its score approaches the label's entropy, which is the measure's ceiling. Genuine predictors rarely come close. Treat a candidate far above the rest as a leakage suspect and check when its value first exists.
- Does conditioning on a third variable settle the direction question?No. Conditional mutual information is still symmetric in its first two arguments, so it reports shared bits given a third variable and never which side is the source. It is genuinely useful for showing that a dependence vanishes once a common driver is held fixed, but direction has to come from timing, an intervention, or known pipeline structure.
saying these in an interview costs you the question
- Reads the top-ranked column as the strongest cause
- Thinks swapping the arguments reveals which way influence runs
- Never asks whether the column is written after the outcome
- Denies that a shared upstream driver can produce a high score
- Assumes only a causal link can make the score large