skip to content

How do entropy, cross-entropy and KL divergence relate to one another?

level: middleimportance: should knowfreq 52%

answer

  1. one is the sum of the other two
  2. irreducible cost plus avoidable surcharge
  3. the floor is not zero
  4. the difference is exactly the divergence
  5. H(P,Q) = H(P) + KL(P||Q)

basics

~20 s

Cross-entropy splits into entropy plus divergence: H(P,Q) = H(P) + KL(P||Q). Entropy is the unavoidable cost of describing draws from P, KL is the extra cost of using Q instead, and cross-entropy is the total.

solid answer

~50 s

Entropy `H(P) = -sum P log P` is the average cost, in bits, of describing draws from `P` using a scheme built for `P`. Cross-entropy `H(P,Q) = -sum P log Q` is the average cost of describing those same draws using a scheme built for `Q`. The gap between them is exactly the KL divergence: `H(P,Q) = H(P) + KL(P||Q)`. Since `KL >= 0`, cross-entropy is always at least the entropy, with equality only when `Q = P`. Two consequences follow. First, cross-entropy is not zero at a perfect match — it bottoms out at `H(P)`, so its floor depends on how uncertain `P` itself is. Second, when `P` is held fixed, `H(P)` is a constant, so comparing cross-entropies across candidate `Q` is the same ranking as comparing KL divergences. Note the argument order: all three are expectations under `P`.

go deeper

for a junior

Be ready to write H(P,Q) = H(P) + KL(P||Q) and say which term is the irreducible part and which is the penalty for using the wrong distribution.

for a middle

Explain the one-line derivation by splitting the log inside KL, and say why cross-entropy's minimum is H(P) rather than zero.

for a senior

Demonstrate the practical consequence: absolute cross-entropy numbers are not comparable across populations whose own entropy differs, so quote the divergence when you mean discrepancy.

for a principal

Own the measurement convention across teams: decide whether discrepancy is reported as a divergence or a raw cost, and make sure the base, the direction and the reference distribution are fixed and documented.

## Three quantities, one identity Fix two distributions on the same discrete outcome space: `P`, the one draws actually come from, and `Q`, the one you are describing them with. ``` H(P) = -sum_x P(x) log P(x) entropy H(P,Q) = -sum_x P(x) log Q(x) cross-entropy KL(P||Q)= sum_x P(x) log( P(x)/Q(x) ) divergence ``` All three are expectations under `P` — the weights are always `P(x)`. Split the log inside KL and the identity falls straight out: ``` KL(P||Q) = sum_x P(x) log P(x) - sum_x P(x) log Q(x) = -H(P) + H(P,Q) ``` so ``` H(P,Q) = H(P) + KL(P||Q) ``` ## Reading it as a cost decomposition The coding picture makes the identity intuitive. If you must communicate one draw from `P`: - Using a description scheme matched to `P` costs `H(P)` bits on average. This is irreducible — no scheme does better, by Shannon's source coding bound. - Using a scheme matched to `Q` instead costs `H(P,Q)` bits on average. - The surcharge for the mismatch is `KL(P||Q)` bits. So cross-entropy = unavoidable cost + avoidable regret. The unavoidable part depends only on `P`; the regret depends on both, and is zero only when `Q = P`. ## Consequences that get asked about **Cross-entropy does not bottom out at zero.** Its minimum over all `Q` is `H(P)`, attained at `Q = P`. So the absolute value of a cross-entropy is not interpretable on its own — a cross-entropy of 0.3 bits could be a perfect description of a highly predictable `P`, or a poor description of a near-certain one. KL is the quantity that is zero at a perfect match, which is why it is the one to quote when you want to say "how wrong is Q". **With P fixed, cross-entropy and KL rank candidates identically.** `H(P)` is then an additive constant, so `H(P,Q1) < H(P,Q2)` if and only if `KL(P||Q1) < KL(P||Q2)`. Comparing cross-entropies across *different* `P` is what breaks: the constant changes, so the numbers are not on a common scale. **The order of arguments is not decoration.** `H(P,Q)` weights `log Q` by `P`. Swapping to `H(Q,P) = -sum Q log P` is a different number answering a different question. The convention to remember is: the first argument supplies the weights (the truth), the second supplies the logs (the belief). **Support still matters.** If some outcome has `P(x) > 0` and `Q(x) = 0`, then `log Q(x) = -infinity` and the cross-entropy is infinite — the same failure mode as KL, for the same reason. ## A worked pair Let `P = (0.5, 0.5)` and `Q = (0.9, 0.1)` on a two-outcome space, in bits. - `H(P) = 1` bit — a fair binary outcome is maximally uncertain. - `KL(P||Q) ≈ 0.737` bits. - Therefore `H(P,Q) ≈ 1.737` bits, and you can check that directly: `-0.5*log2(0.9) - 0.5*log2(0.1) ≈ 0.5*0.152 + 0.5*3.322 ≈ 1.737`. The 1 bit is what any scheme must pay for a fair coin. The 0.74 bits is what believing in a 90/10 coin costs on top, and it is driven almost entirely by the half of the draws that `Q` assigns probability 0.1. ## Special cases and relatives - If `P` is a **point mass** on one outcome `x*` — probability 1 there, 0 elsewhere — then `H(P) = 0`, so `H(P,Q) = KL(P||Q) = -log Q(x*)`. The two quantities coincide, which is why the distinction is easy to lose sight of in settings where the truth is a single known outcome. - If `Q` is **uniform** over `k` outcomes, `H(P,Q) = log2(k)` for every `P`, so `KL(P||uniform) = log2(k) - H(P)`. The divergence from uniform is exactly the entropy deficit — a tidy way to see that low-entropy distributions are far from uniform. - **Conditional entropy** `H(Y|X)` and **mutual information** are built from the same ingredients: `I(X;Y) = H(Y) - H(Y|X)`, and equivalently the KL between the joint and the product of marginals. ## What a strong answer sounds like State the identity `H(P,Q) = H(P) + KL(P||Q)`, say in words that cross-entropy is the total cost while KL is the excess over the irreducible entropy, note that cross-entropy's floor is `H(P)` and not zero, and point out that with `P` fixed the two rank candidates the same way. Naming the argument-order convention shows you actually understand the expectation rather than having memorised three formulas.

  • Why does cross-entropy not reach zero even when Q equals P?
    Because its floor is the entropy of `P` itself. `H(P,Q) = H(P) + KL(P||Q)`, and at `Q = P` the divergence term vanishes but `H(P)` remains. Describing draws from an uncertain source costs bits no matter how good the description is. Only KL, which strips out that irreducible term, is zero at a perfect match — which is why an absolute cross-entropy value is not interpretable without knowing `H(P)`.
  • If cross-entropy and KL differ only by a constant, when does the distinction actually matter?
    It matters whenever `P` changes between the numbers you are comparing. `H(P)` is only a constant when `P` is fixed, so cross-entropies computed against different truths — different segments, different periods, different label mixes — sit on different scales and cannot be compared directly. KL removes that offset. It also matters when someone reads an absolute cross-entropy as a quality score.
  • What is KL(P || uniform) in terms of entropy?
    For a uniform `Q` over `k` outcomes, `log Q(x) = -log2(k)` everywhere, so the cross-entropy is `log2(k)` for any `P`. The identity then gives `KL(P||uniform) = log2(k) - H(P)`: the divergence from uniform is exactly the entropy deficit. A distribution is far from uniform precisely when it is low-entropy, and the maximum-entropy distribution is the one at zero divergence.

Entropy is the fare you must pay anyway, KL is the penalty for the wrong ticket, and cross-entropy is what shows up on the bill.

saying these in an interview costs you the question

  • Treats cross-entropy and KL divergence as the same quantity
  • Expects cross-entropy to reach zero on a perfect match
  • Compares absolute cross-entropies across different underlying distributions
  • Gets the argument order backwards in H(P,Q)
  • Forgets that cross-entropy is infinite when Q assigns zero probability

context