A candidate fraud signal shows zero mutual information with the fraud label — what does that tell you?
answer
- bits removed, not correlation
- zero means the joint factorises
- symmetric: same number both ways
- estimated from counts, rarely exact
- pairwise zero, still useful jointly
basics
~20 sZero mutual information means the signal and the label are independent in the distribution you measured: observing one removes none of the other's uncertainty. The reading is symmetric, it is an estimate, and it says nothing about the signal's value alongside other signals.
solid answer
~40 sMutual information is `I(X;Y) = H(X) - H(X|Y)`: the bits of uncertainty about one variable that observing the other removes. It is never negative, and it is zero exactly when the joint distribution factorises, `p(x,y) = p(x)p(y)` — so zero means independence over the distribution measured, nothing weaker and nothing stronger. Because `H(X) - H(X|Y)` equals `H(Y) - H(Y|X)`, the statement is symmetric: the label says as little about the signal as the signal says about the label. Two caveats matter in practice. The figure is estimated from a finite sample, so an exact zero is rare and a small value is often noise. And a signal with zero pairwise score can still be informative jointly with another, so zero alone is not a reason to throw it away.
code
pseudocode · 10 lines# counts[x][y] = rows with signal x and label y, N rows total
I = 0
for each x in values_of_signal:
for each y in values_of_label:
p_xy = counts[x][y] / N
if p_xy > 0: # skips the 0 * log 0 terms
p_x = row_total(x) / N
p_y = col_total(y) / N
I = I + p_xy * log2( p_xy / (p_x * p_y) )
return I # bits; exactly 0 when p_xy = p_x * p_y in every cellgo deeper
Recall the one-line meaning: mutual information counts the bits that observing one variable removes from the uncertainty about another, it is measured in the same unit as entropy, and zero means the two are independent.
Be able to write it both ways — entropy minus conditional entropy, from either side — and explain why that symmetry makes the score directionless, why it can never go negative, and why zero is exactly the factorising condition.
Show that you treat the figure as an estimate: say how sample size and binning move it, and why you would not drop a zero-scoring column without checking it alongside the columns already chosen.
Frame where a dependence score belongs in a selection policy at all. It prices shared bits and nothing else — not redundancy between candidates, not the cost of obtaining one at scoring time, not whether the dependence will still hold next quarter.
## What the number counts **Mutual information**, written `I(X;Y)`, is measured in **bits** — the same unit as entropy. It answers one question: before you observe `Y`, how much of your uncertainty about `X` do you expect that observation to remove? The definition states exactly that: `I(X;Y) = H(X) - H(X|Y)` Here `H(X)` is the **entropy** of `X`, the average number of bits needed to pin down its value, and `H(X|Y)` is the uncertainty that survives once `Y` is known. The difference is the part the two variables share. An equivalent form drops the conditioning entirely: `I(X;Y) = sum over x,y of p(x,y) * log2( p(x,y) / (p(x) * p(y)) )` This compares the **joint distribution** `p(x,y)` against the table you would get if the two were unrelated, `p(x) * p(y)`. When those two tables agree in every cell, each log term is `log2(1) = 0` and the whole sum collapses to zero. ## Three properties that follow - **It is never negative.** The sum above is a relative entropy between the joint and the product of the marginals, and that quantity is zero when the two agree and positive otherwise. There is no sign to carry: the score reports how much dependence there is, not whether the two variables rise and fall together. - **It is symmetric.** `H(X) - H(X|Y)` and `H(Y) - H(Y|X)` are the same number, so `I(X;Y) = I(Y;X)`. A signal tells you exactly as many bits about a label as the label tells you about the signal. - **It is zero exactly when the pair is independent** in the distribution it was computed over. Not merely uncorrelated — independent. A relation of any shape, including one that no straight line would pick up, shows up as a positive score. That is the main reason a dependence screen reaches for this measure rather than a linear coefficient. ## What a reported zero does and does not rule out | The zero says | The zero does not say | |---|---| | The joint table factorises over the sample measured | That the two are independent in data you have not seen | | Observing one leaves the other's uncertainty untouched | That the signal is worthless alongside other signals | | The same holds read in either direction | Anything about which variable came first, or caused what | | No relation was detected at this binning resolution | That a finer binning would also find none | Three practical caveats sit behind that right-hand column. 1. **It is an estimate.** The figure comes from counts on a finite sample. The straightforward plug-in estimator is biased upward, so an exact zero is rare in practice and a small positive value is frequently noise rather than a faint relation. 2. **Binning changes it.** A continuous quantity must be bucketed before it can be counted. Buckets that are too wide hide a dependence living inside one of them; buckets that are too narrow inflate the estimate, because each bucket then holds very few rows. 3. **Pairwise is not joint.** A signal can be independent of the label on its own and still determine it in combination. If the label is the parity of two fair binary signals, each signal alone leaves the label equally likely to take either value — zero bits — while the two together fix it exactly, one bit. Screening candidates one at a time discards precisely this kind of signal. ## Reading a non-zero score The score is bounded: `I(X;Y)` can never exceed `H(X)` or `H(Y)`, so its ceiling is the smaller of the two entropies. That matters when comparing columns. A signal with two values cannot score above one bit no matter how tightly it tracks the label, while a many-valued column has a much higher ceiling and will tend to sit higher in a raw ranking for that reason alone. Comparing raw scores across columns of very different cardinality is therefore an apples-to-oranges comparison; dividing by the ceiling gives a bounded ratio, at the cost of discarding the plain bit interpretation. ## Using it to rank candidate signals In a ranking exercise — scoring a pile of candidate inputs against an outcome label before anyone trains anything — the honest reading of a zero is "this column, on its own, at this resolution, on this sample, moves nothing". That is a reason to deprioritise, not a proof of uselessness. The honest reading of a large value is equally narrow: the column shares bits with the label. It does not say the column influences the outcome. ## What an interviewer is listening for A complete answer names the unit (bits), the definition through conditional entropy, the independence condition, and the symmetry — then adds that the number is an estimate and that pairwise independence is weaker than joint irrelevance. Candidates who describe it as "a better correlation" usually lose the two properties that matter most: that it carries no sign, and that it detects dependence of any shape rather than co-movement.
- Why can mutual information never be negative, while a correlation coefficient can?It is an average of `log2(p(x,y)/(p(x)p(y)))` weighted by the joint distribution — a relative entropy between the joint and the product of the marginals, which is zero when they agree and positive otherwise. It has no sign to carry: it measures how much dependence exists, not whether the variables move together or in opposite directions. A correlation coefficient encodes direction in its sign and only sees linear co-movement.
- How can a signal score zero against the label on its own and still matter?Dependence can live only in a combination. If the label is the parity of two fair binary signals, each signal alone leaves the label's uncertainty untouched — zero bits — while the pair determines it exactly. A one-at-a-time screen discards such a signal. Guard against it by scoring candidates in small groups, or by scoring each against the label conditioned on the signals already selected.
- Does a score of zero mean the two variables are unrelated in reality?No. It means the joint counts you measured factorise, at the binning you chose, on the sample you had. A dependence confined to a rare region, or visible only at a finer resolution, can be invisible to that estimate. Treat zero as a measurement about a sample rather than a statement about the world, and re-check on more data before acting on it.
saying these in an interview costs you the question
- Says a zero score proves the signal is useless
- Treats it as a correlation coefficient that can go negative
- Claims reading it label-to-signal gives a different number
- Believes mutual information only detects linear relationships
- Takes a sample estimate of zero as exact independence