skip to content

Why is relative entropy not a distance between two probability distributions?

level: seniorimportance: nice to knowfreq 26%

answer

  1. a directed penalty, not a metric
  2. the first argument sets the weights
  3. swap the arguments, different number
  4. zero model probability, unbounded cost
  5. triangle inequality fails as well

basics

~20 s

Relative entropy is asymmetric: the bits wasted coding a source with a model differ from the reverse, because each direction weights the same log-ratios by a different distribution. It also fails the triangle inequality, so it is a directed penalty rather than a metric.

solid answer

~40 s

The definition `D(p||q) = sum p(s) log2(p(s)/q(s))` weights every log-ratio by the **first** argument, so the two arguments do not play the same role. Concretely, with true frequencies `(0.9, 0.1)` and a model at `(0.5, 0.5)`, `D(p||q)` is about 0.531 bits while `D(q||p)` is about 0.737 bits — same pair, different answers. It is also unbounded: if the model gives probability zero to a symbol the source does emit, that term grows without limit and the divergence is infinite. A metric would additionally satisfy the triangle inequality, and this quantity does not. The practical rule that falls out: when you know the real frequencies, put them **first**, because that is the direction whose value is the bits actually wasted.

go deeper

for a junior

Remember the headline: this quantity has a direction. Reporting it without saying which distribution came first leaves the number uninterpretable, and it is not a distance despite being used to compare distributions.

for a middle

Explain the asymmetry from the formula: the log-ratio flips sign under a swap but the weights come from the first argument only, so the two directions have no reason to agree.

for a senior

Show the operational habits: real frequencies first, a floor probability reserved for unseen symbols so the value stays finite, and no geometric reasoning about closeness since the triangle inequality does not hold.

for a principal

Decide what a team's comparison metric should actually be. A directed coding penalty and a symmetric summary answer different questions, and letting both circulate under one name produces arguments where two people are quoting different quantities.

## The definition is directed by construction **Relative entropy**, the **Kullback-Leibler divergence**, is `D(p||q) = sum over s of p(s) * log2( p(s) / q(s) )` Read the two arguments in their roles. The ratio inside the logarithm is antisymmetric — swapping `p` and `q` flips its sign — but the **weights** outside are not: every term is weighted by `p(s)`, the first argument alone. Swapping the arguments therefore changes both the signs inside and the weights outside, and there is no reason the two effects should cancel. They generally do not. The coding reading makes the asymmetry concrete. `D(p||q)` prices one specific situation: symbols arrive at frequencies `p`, a coder built on `q` assigns them lengths, and the answer is the excess bits per symbol. `D(q||p)` prices a different, hypothetical situation with the roles of source and model exchanged. There is no reason two different situations should cost the same. ## The numbers on a two-symbol pair Take `p = (0.9, 0.1)` and `q = (0.5, 0.5)`. | Direction | Computation | Value | |---|---|---| | `D(p||q)` | `0.9*log2(0.9/0.5) + 0.1*log2(0.1/0.5)` | about 0.531 bits | | `D(q||p)` | `0.5*log2(0.5/0.9) + 0.5*log2(0.5/0.1)` | about 0.737 bits | Same two distributions, two different answers, and the second is nearly 40% larger. Any procedure that silently picks an argument order is silently picking one of these. ## Three ways it fails the definition of a metric A metric on distributions would have to be symmetric, satisfy the triangle inequality, and take finite values. This quantity keeps only the first half of the identity condition — it is zero exactly when the two distributions agree — and drops the rest: 1. **Asymmetry.** Shown above: `D(p||q)` and `D(q||p)` are different numbers in general. 2. **No triangle inequality.** There are triples of distributions for which going directly costs more than the sum of two legs through an intermediate, which a distance can never allow. 3. **Unboundedness.** If `q(s) = 0` while `p(s) > 0`, the term `p(s) * log2(p(s)/q(s))` grows without limit and the divergence is infinite. In coding terms the model has assigned no finite codeword to a symbol that will actually arrive, so the stream cannot be encoded at all. That third point is the one that shows up in practice rather than in a proof. It is why a probability table built from observed counts is normally given a small floor probability for every symbol the alphabet allows: the cost is a fraction of a bit on the common symbols, and the benefit is that the worst case stays finite when something unseen finally turns up. ## Which order to use The argument order is a modelling decision with a right answer whenever you know which distribution is real: - **Put the true or observed frequencies first** when you are pricing waste. `D(true||model)` is the bits actually thrown away, because the definition weights by the first argument and reality supplies the weights. - The reverse order penalises the two kinds of error differently. Weighting by the first argument means the first distribution's high-probability regions dominate, so a model is punished hardest for putting little mass where the weighting distribution puts a lot. - Where no direction is natural — comparing two empirical tables with neither designated as truth — people use a **symmetrised** construction instead, typically averaging the two divergences against a mixture of the pair. That restores symmetry and boundedness at the cost of the plain "extra bits paid" reading. ## Practical consequences - **Never report a bare divergence without saying which way round it was computed.** The number is not interpretable without it. - **Never average the two directions and call the result the divergence.** It may be a reasonable symmetric summary, but it is a different quantity and should be named as one. - **Guard the zero case.** Any pipeline that estimates a probability table from counts should reserve mass for unseen symbols, or it will produce an infinity the first time reality exceeds the sample. - **Do not reason with it geometrically.** Because the triangle inequality fails, "A is close to B and B is close to C, so A is close to C" is not an argument you may make with these values. ## What an interviewer is listening for The complete answer names asymmetry first, explains it from the weighting term rather than by assertion, and then adds the two other metric failures: no triangle inequality, and an infinite value when the model has a zero where the source has mass. The strongest candidates finish with the operational rule — real frequencies go first — because that is the part that changes what someone types.

  • What happens when the model gives probability zero to a symbol the source emits?
    The term `p(s) * log2(p(s)/q(s))` grows without bound, so the divergence is infinite: the model implies no finite codeword for a symbol that will actually occur, and the stream cannot be coded at all. The standard repair is to reserve a small probability for every symbol the alphabet allows, paying a fraction of a bit on the common symbols to keep the worst case finite.
  • Which argument order should you use when the true frequencies are known?
    Put the true frequencies first. The definition weights each log-ratio by the first argument, and the coding reading charges those real frequencies at the model's code lengths, so that direction is the bits genuinely wasted. The reverse order prices a different, hypothetical arrangement and will generally give another number.

saying these in an interview costs you the question

  • Assumes swapping the two distributions leaves the value unchanged
  • Believes it obeys the triangle inequality like a metric
  • Thinks a zero model probability costs only a large finite penalty
  • Picks the argument order arbitrarily when the real frequencies are known
  • Claims the reverse direction measures the same modelling error