How do entropy, cross-entropy and KL divergence relate to one another?
answer
- one is the sum of the other two
- irreducible cost plus avoidable surcharge
- the floor is not zero
- the difference is exactly the divergence
- H(P,Q) = H(P) + KL(P||Q)
basics
~20 sCross-entropy splits into entropy plus divergence: H(P,Q) = H(P) + KL(P||Q). Entropy is the unavoidable cost of describing draws from P, KL is the extra cost of using Q instead, and cross-entropy is the total.
solid answer
~50 sEntropy `H(P) = -sum P log P` is the average cost, in bits, of describing draws from `P` using a scheme built for `P`. Cross-entropy `H(P,Q) = -sum P log Q` is the average cost of describing those same draws using a scheme built for `Q`. The gap between them is exactly the KL divergence: `H(P,Q) = H(P) + KL(P||Q)`. Since `KL >= 0`, cross-entropy is always at least the entropy, with equality only when `Q = P`. Two consequences follow. First, cross-entropy is not zero at a perfect match — it bottoms out at `H(P)`, so its floor depends on how uncertain `P` itself is. Second, when `P` is held fixed, `H(P)` is a constant, so comparing cross-entropies across candidate `Q` is the same ranking as comparing KL divergences. Note the argument order: all three are expectations under `P`.
go deeper
Be ready to write H(P,Q) = H(P) + KL(P||Q) and say which term is the irreducible part and which is the penalty for using the wrong distribution.
Explain the one-line derivation by splitting the log inside KL, and say why cross-entropy's minimum is H(P) rather than zero.
Demonstrate the practical consequence: absolute cross-entropy numbers are not comparable across populations whose own entropy differs, so quote the divergence when you mean discrepancy.
Own the measurement convention across teams: decide whether discrepancy is reported as a divergence or a raw cost, and make sure the base, the direction and the reference distribution are fixed and documented.
## Three quantities, one identity Fix two distributions on the same discrete outcome space: `P`, the one draws actually come from, and `Q`, the one you are describing them with. ``` H(P) = -sum_x P(x) log P(x) entropy H(P,Q) = -sum_x P(x) log Q(x) cross-entropy KL(P||Q)= sum_x P(x) log( P(x)/Q(x) ) divergence ``` All three are expectations under `P` — the weights are always `P(x)`. Split the log inside KL and the identity falls straight out: ``` KL(P||Q) = sum_x P(x) log P(x) - sum_x P(x) log Q(x) = -H(P) + H(P,Q) ``` so ``` H(P,Q) = H(P) + KL(P||Q) ``` ## Reading it as a cost decomposition The coding picture makes the identity intuitive. If you must communicate one draw from `P`: - Using a description scheme matched to `P` costs `H(P)` bits on average. This is irreducible — no scheme does better, by Shannon's source coding bound. - Using a scheme matched to `Q` instead costs `H(P,Q)` bits on average. - The surcharge for the mismatch is `KL(P||Q)` bits. So cross-entropy = unavoidable cost + avoidable regret. The unavoidable part depends only on `P`; the regret depends on both, and is zero only when `Q = P`. ## Consequences that get asked about **Cross-entropy does not bottom out at zero.** Its minimum over all `Q` is `H(P)`, attained at `Q = P`. So the absolute value of a cross-entropy is not interpretable on its own — a cross-entropy of 0.3 bits could be a perfect description of a highly predictable `P`, or a poor description of a near-certain one. KL is the quantity that is zero at a perfect match, which is why it is the one to quote when you want to say "how wrong is Q". **With P fixed, cross-entropy and KL rank candidates identically.** `H(P)` is then an additive constant, so `H(P,Q1) < H(P,Q2)` if and only if `KL(P||Q1) < KL(P||Q2)`. Comparing cross-entropies across *different* `P` is what breaks: the constant changes, so the numbers are not on a common scale. **The order of arguments is not decoration.** `H(P,Q)` weights `log Q` by `P`. Swapping to `H(Q,P) = -sum Q log P` is a different number answering a different question. The convention to remember is: the first argument supplies the weights (the truth), the second supplies the logs (the belief). **Support still matters.** If some outcome has `P(x) > 0` and `Q(x) = 0`, then `log Q(x) = -infinity` and the cross-entropy is infinite — the same failure mode as KL, for the same reason. ## A worked pair Let `P = (0.5, 0.5)` and `Q = (0.9, 0.1)` on a two-outcome space, in bits. - `H(P) = 1` bit — a fair binary outcome is maximally uncertain. - `KL(P||Q) ≈ 0.737` bits. - Therefore `H(P,Q) ≈ 1.737` bits, and you can check that directly: `-0.5*log2(0.9) - 0.5*log2(0.1) ≈ 0.5*0.152 + 0.5*3.322 ≈ 1.737`. The 1 bit is what any scheme must pay for a fair coin. The 0.74 bits is what believing in a 90/10 coin costs on top, and it is driven almost entirely by the half of the draws that `Q` assigns probability 0.1. ## Special cases and relatives - If `P` is a **point mass** on one outcome `x*` — probability 1 there, 0 elsewhere — then `H(P) = 0`, so `H(P,Q) = KL(P||Q) = -log Q(x*)`. The two quantities coincide, which is why the distinction is easy to lose sight of in settings where the truth is a single known outcome. - If `Q` is **uniform** over `k` outcomes, `H(P,Q) = log2(k)` for every `P`, so `KL(P||uniform) = log2(k) - H(P)`. The divergence from uniform is exactly the entropy deficit — a tidy way to see that low-entropy distributions are far from uniform. - **Conditional entropy** `H(Y|X)` and **mutual information** are built from the same ingredients: `I(X;Y) = H(Y) - H(Y|X)`, and equivalently the KL between the joint and the product of marginals. ## What a strong answer sounds like State the identity `H(P,Q) = H(P) + KL(P||Q)`, say in words that cross-entropy is the total cost while KL is the excess over the irreducible entropy, note that cross-entropy's floor is `H(P)` and not zero, and point out that with `P` fixed the two rank candidates the same way. Naming the argument-order convention shows you actually understand the expectation rather than having memorised three formulas.
- Why does cross-entropy not reach zero even when Q equals P?Because its floor is the entropy of `P` itself. `H(P,Q) = H(P) + KL(P||Q)`, and at `Q = P` the divergence term vanishes but `H(P)` remains. Describing draws from an uncertain source costs bits no matter how good the description is. Only KL, which strips out that irreducible term, is zero at a perfect match — which is why an absolute cross-entropy value is not interpretable without knowing `H(P)`.
- If cross-entropy and KL differ only by a constant, when does the distinction actually matter?It matters whenever `P` changes between the numbers you are comparing. `H(P)` is only a constant when `P` is fixed, so cross-entropies computed against different truths — different segments, different periods, different label mixes — sit on different scales and cannot be compared directly. KL removes that offset. It also matters when someone reads an absolute cross-entropy as a quality score.
- What is KL(P || uniform) in terms of entropy?For a uniform `Q` over `k` outcomes, `log Q(x) = -log2(k)` everywhere, so the cross-entropy is `log2(k)` for any `P`. The identity then gives `KL(P||uniform) = log2(k) - H(P)`: the divergence from uniform is exactly the entropy deficit. A distribution is far from uniform precisely when it is low-entropy, and the maximum-entropy distribution is the one at zero divergence.
Entropy is the fare you must pay anyway, KL is the penalty for the wrong ticket, and cross-entropy is what shows up on the bill.
saying these in an interview costs you the question
- Treats cross-entropy and KL divergence as the same quantity
- Expects cross-entropy to reach zero on a perfect match
- Compares absolute cross-entropies across different underlying distributions
- Gets the argument order backwards in H(P,Q)
- Forgets that cross-entropy is infinite when Q assigns zero probability