skip to content

What does mutual information between acquisition channel and conversion actually measure?

level: middleimportance: should knowfreq 36%

answer

  1. how much one variable tells you about another
  2. uncertainty before minus uncertainty after
  3. divergence between joint and product of marginals
  4. zero if and only if independent
  5. H(Y) - H(Y|X)

basics

~10 s

It measures how many bits knowing the channel removes from the uncertainty about conversion: I(X;Y) = H(Y) - H(Y|X). It is symmetric, never negative, and exactly zero when channel and conversion are independent.

solid answer

~50 s

Mutual information is `I(X;Y) = H(Y) - H(Y|X)`, the average reduction in uncertainty about conversion once the acquisition channel is known, in bits. Equivalently it is `KL( P(X,Y) || P(X)P(Y) )` — the divergence between the true joint distribution and what it would be if the two were independent. That form makes two properties obvious: it is never negative, and it is exactly zero if and only if `X` and `Y` are independent. It is also symmetric, `I(X;Y) = I(Y;X)`, and it captures any kind of dependence, not just monotone or linear association. The practical catch is scale. With a 6% base conversion rate split into a 10% channel and a 2% channel, `H(Y) ≈ 0.327` bits and `H(Y|X) ≈ 0.305` bits, so `I ≈ 0.022` bits per visit — a tiny number describing a five-fold difference in rates.

go deeper

for a junior

Be ready to say it is the reduction in uncertainty about one variable from knowing the other, that it is never negative, and that it is zero when they are independent.

for a middle

Explain the equivalent forms, especially I(X;Y) = KL(joint || product of marginals), and use that to justify non-negativity and the exact-zero-under-independence property.

for a senior

Show you handle estimation: plug-in values are biased upward with cardinality, so use a permutation reference, collapse long tails, and report the value relative to the marginal entropy.

for a principal

Own how dependence is screened at scale: set the convention for cardinality control and null comparison, and keep a symmetric association statistic from being presented to stakeholders as a causal claim.

## Definition and three equivalent forms Let `X` be the acquisition channel of a visit and `Y` an indicator of whether the visit converts. Mutual information is ``` I(X;Y) = H(Y) - H(Y|X) = H(X) - H(X|Y) = H(X) + H(Y) - H(X,Y) ``` where `H(Y|X) = sum_x P(x) H(Y | X = x)` is the conditional entropy — the entropy of `Y` within each channel, averaged over channels. A fourth, and the most revealing, form is ``` I(X;Y) = KL( P(X,Y) || P(X)P(Y) ) ``` the divergence between the actual joint distribution and the product of the marginals, which is what the joint would be if the two were independent. ## What the properties follow from - **Non-negative.** Because it is a KL divergence, `I(X;Y) >= 0` always. Learning the channel can never, on average, increase your uncertainty about conversion — though it certainly can for a particular channel, since `H(Y | X = x)` may exceed `H(Y)` for one value of `x`; only the average is guaranteed to drop. - **Zero exactly under independence.** KL is zero only when its two arguments coincide, so `I(X;Y) = 0` if and only if `P(X,Y) = P(X)P(Y)` — exactly the definition of independence. This is a genuine if-and-only-if, which is the reason mutual information is quoted as a dependence measure at all. - **Symmetric.** `I(X;Y) = I(Y;X)`. It is a property of the pair, not a directed effect of one on the other. - **Any dependence counts.** Because it works on the joint distribution directly, it detects non-monotone and non-linear structure — a channel whose conversion rate is high at both extremes of some ordering and low in the middle still registers. - **Bounded by the entropies.** `I(X;Y) <= min( H(X), H(Y) )`, with equality when one variable is a deterministic function of the other. Uncertainty that is not there cannot be removed. ## The scale problem, with numbers Suppose visits split evenly between two channels, one converting at 10% and one at 2%, so the overall rate is 6%. ``` H(Y) = H(0.06) ≈ 0.327 bits H(Y|X) = 0.5*H(0.10) + 0.5*H(0.02) ≈ 0.5*0.469 + 0.5*0.141 ≈ 0.305 bits I(X;Y) ≈ 0.022 bits per visit ``` A five-fold difference in conversion rate — enormous commercially — is worth about 0.02 bits. The reason is that conversion is a rare event: `H(Y)` is only 0.327 bits to begin with, so there is very little uncertainty available to remove. Two lessons follow. First, never judge a mutual information value against an absolute intuition; compare it to `H(Y)`, or report the normalised ratio `I(X;Y) / H(Y)`, which here is about 7%. Second, a small absolute MI does not mean a small business effect, and a comparison of MI values across variables with different marginal entropies is not apples to apples. ## Estimation pitfalls **Plug-in MI is biased upward.** Computing `I` from observed counts systematically overstates dependence, and the bias grows with the number of categories relative to the sample size. In the extreme, an identifier column with one distinct value per row gives `H(Y|X) = 0` and therefore maximal apparent MI, while carrying no generalisable signal at all. Any ranking of features by raw MI will put high-cardinality columns on top for this reason. **Two defences.** Compare the observed value against a permutation reference: shuffle one variable relative to the other many times, recompute, and see where the real value falls in that null spread. And keep the category count sane — collapse a long tail into an explicit "other" bucket rather than letting hundreds of near-empty channels each contribute their own spurious certainty. **Binning drives continuous MI.** For a continuous variable, the value you get depends on the binning you chose, so the binning must be fixed and disclosed alongside the number. ## What it does not tell you Mutual information is symmetric and therefore says nothing about direction, let alone causation: channel and conversion may share a common cause such as intent, and MI is identical either way round. It also gives no sign — you learn that channel is informative about conversion, not which channels are good. For that you need the per-channel rates themselves, or the per-cell contributions to the sum. Treat MI as a screening statistic that ranks candidate relationships worth looking at, and follow it with something that carries direction and magnitude.

  • Why is mutual information exactly zero under independence, rather than just small?
    Because `I(X;Y) = KL( P(X,Y) || P(X)P(Y) )`, and a KL divergence is zero if and only if its two arguments are the same distribution. Independence is precisely the statement that the joint equals the product of the marginals, so the two conditions coincide. That makes it an if-and-only-if characterisation of independence for the true distributions — an estimate from finite data will of course be small but non-zero.
  • Ranking candidate variables by mutual information puts a near-unique identifier at the top. What went wrong?
    Plug-in mutual information is biased upward, and the bias grows with cardinality. With one distinct value per row the conditional entropy collapses to zero, so the estimate is maximal while the generalisable signal is nil. Defend against it by comparing each value against a permutation null built by shuffling one variable, and by collapsing long tails into an explicit other bucket before estimating.
  • Mutual information between channel and conversion is 0.02 bits. Is that a lot?
    You cannot tell without the marginal entropy. Compare it to `H(Y)`: if conversion runs at 6%, `H(Y) ≈ 0.327` bits, so 0.02 bits is roughly 7% of the available uncertainty — a real signal, and consistent with one channel converting several times better than another. Rare outcomes have little entropy to remove, so absolute MI values on them are always small and only the ratio is interpretable.
  • Does a non-zero mutual information tell you the channel drives conversion?
    No. Mutual information is symmetric, so it carries no direction, and it cannot distinguish a causal path from a shared common cause such as visitor intent. It also has no sign, so it says the two are associated without saying which channels are better. Use it to screen which relationships are worth investigating, then reach for the per-channel rates and a design that supports a causal claim.

It is the number of yes/no questions about conversion you no longer have to ask once someone tells you the channel.

saying these in an interview costs you the question

  • Says mutual information can be negative for opposing relationships
  • Reads a symmetric quantity as evidence of causal direction
  • Judges an absolute MI value without comparing it to the marginal entropy
  • Ranks high-cardinality columns highest and trusts the ranking
  • Thinks a small MI value proves independence in a finite sample

context