skip to content

How do you decide whether to minimise KL(P||Q) or KL(Q||P) when approximating a distribution?

level: principalimportance: nice to knowfreq 24%

answer

  1. which mistake costs you more
  2. look at whose mass supplies the weights
  3. one spreads out, the other locks on
  4. zero-avoiding versus zero-forcing
  5. a single fit to two modes reveals the difference

basics

~20 s

Pick the direction by which error you can least afford. Minimising KL(P||Q) is mass-covering: the approximation stretches to cover everything the target produces. Minimising KL(Q||P) is mode-seeking: it locks onto one region and ignores the rest.

solid answer

~50 s

The two directions penalise opposite mistakes. Forward KL, `KL(P||Q)` with expectation under the target `P`, charges you wherever `P` has mass and `Q` does not — it is zero-avoiding, so a simple `Q` fitted this way spreads out to cover every region `P` reaches, including the low-density gap between two modes. Reverse KL, `KL(Q||P)` with expectation under the approximation, only charges where `Q` puts mass, so `Q` can safely ignore parts of `P` entirely — it is zero-forcing and collapses onto a single mode. Choose forward when missing a region is the expensive failure: risk tails, rare-but-costly segments, anything where you must not declare a real outcome impossible. Choose reverse when a confident, sharp answer matters more than coverage, and a fit straddling a low-density valley would be worse than useless. Say which one you mean, always.

go deeper

for a junior

Be ready to say the two KL directions are different numbers, and that the expectation is taken under whichever distribution is written first.

for a middle

Explain why the weighting drives the behaviour: forward KL is penalised where the target has mass and the fit does not, reverse KL only where the fit itself has mass.

for a senior

Diagnose a fit from its symptoms: an over-wide approximation placing mass in an empty valley, or an over-confident one that has quietly dropped a whole sub-population.

for a principal

Own the choice and the convention: tie the direction to which error the business cannot absorb, document why it was picked, and require the argument order to be named wherever a divergence is reported.

## Two directions, two different objectives Let `P` be the target distribution — complicated, possibly multi-modal — and `Q` the simpler family you are fitting, say a single unimodal distribution. There are two ways to point KL at the problem: ``` forward: minimise KL(P||Q) = sum_x P(x) log( P(x)/Q(x) ) expectation under P reverse: minimise KL(Q||P) = sum_x Q(x) log( Q(x)/P(x) ) expectation under Q ``` The weights differ, and that is the whole story. The first averages over where the *target* has mass; the second averages over where the *approximation* has mass. ## Forward KL is zero-avoiding In `KL(P||Q)`, every region where `P(x) > 0` contributes, weighted by `P(x)`. If `Q(x)` is near zero there, `log(P/Q)` is large and the penalty is severe — infinite in the limit `Q(x) -> 0`. There is no corresponding penalty for `Q` putting mass where `P` has none: those terms carry weight `P(x) = 0` and vanish. The optimiser therefore does whatever it takes to keep `Q` positive everywhere `P` is positive. Fitting a single unimodal `Q` to a two-mode `P`, forward KL produces a wide fit centred between the modes, assigning real probability to the low-density valley that separates them — a region the target almost never visits. The fit is safe in the sense that it never says "impossible" about something that happens, and wrong in the sense that its typical draws are atypical of `P`. When `Q` is fit to samples, forward KL is also the direction that only needs draws from `P`, since the expectation is under `P`. ## Reverse KL is zero-forcing In `KL(Q||P)`, terms are weighted by `Q(x)`. Where `Q` puts no mass, nothing is charged, no matter what `P` does there. But where `Q` does put mass, `P` had better agree: `log(Q/P)` blows up if `Q` places mass where `P` is near zero. So the optimiser shrinks `Q` onto a region where `P` is comfortably large, and abandons the rest. A unimodal `Q` fitted to a two-mode `P` by reverse KL picks one mode, fits it snugly, and pretends the other does not exist. Draws from that `Q` look like plausible draws from `P` — they just do not represent the full picture, and the fit is typically over-confident, understating spread. ## Choosing, as a judgment call There is no universally right direction; there is a right one for the cost structure you face. **Prefer mass-covering when omission is the expensive error.** Anything where declaring a real outcome impossible is catastrophic — tail risk, safety-relevant rare events, a segment you cannot afford to be blind to, a proposal distribution that must not have holes — argues for forward KL. The price you accept is a diffuse, under-confident approximation that will assign probability to things that never happen. **Prefer mode-seeking when a sharp, self-consistent answer is what gets used.** If the downstream consumer takes a single representative value or a small set of draws and acts on them, a fit straddling a valley between two modes produces answers that are typical of *neither* mode and therefore useless in practice. The price you accept is under-stated uncertainty and total blindness to whatever `Q` walked away from. **Watch what is computable.** Forward KL needs expectations under `P`, which is natural when you have samples from the target and awkward when you only have an unnormalised description of it. Reverse KL needs expectations under `Q`, which you control and can sample from freely. That asymmetry, not aesthetics, often decides the matter — and it is worth being honest about when it does, because a direction chosen for tractability should not be defended afterwards as a modelling preference. **Know that neither is symmetric, and neither is a distance.** If you genuinely need a symmetric, bounded comparison rather than a fitting objective, the Jensen-Shannon divergence, `0.5*KL(P||M) + 0.5*KL(Q||M)` with `M` the mixture, is the standard answer. ## Organisational discipline The practical failure is not choosing wrong; it is not writing the choice down. "The KL between the model and the data" is ambiguous, and two teams reporting "KL" from opposite directions will produce numbers that disagree for reasons nobody can reconstruct months later. The convention to enforce is simple: name the arguments in order every time, state which is the target, and record the reason the direction was picked. If the fitted approximation is under-dispersed, that should be a documented consequence of a deliberate choice, not a surprise discovered when a downstream interval turns out to be too narrow. ## Diagnostics worth running Whichever direction you pick, check the failure mode it invites. For a mass-covering fit, ask what fraction of the approximation's mass sits in regions the target essentially never produces. For a mode-seeking fit, ask what the approximation assigns to regions the target does produce, and whether an entire sub-population has been dropped. A one-number objective will not tell you either of these; a comparison of the two distributions over the regions you care about will.

  • Fitting one unimodal distribution to a two-mode target, what does each direction produce?
    Forward KL, `KL(P||Q)`, gives a wide fit centred between the modes, putting real probability on the low-density valley the target rarely visits, because it is penalised for any region where the target has mass and the fit does not. Reverse KL, `KL(Q||P)`, picks one mode, fits it tightly, and ignores the other, because it is only charged where the fit itself puts mass.
  • Which direction tends to understate uncertainty, and why does that matter?
    Reverse KL. Because it is only penalised where the approximation has mass, it shrinks onto a high-density region and reports a narrower spread than the target actually has. That matters when the fit feeds an interval, a risk number, or any decision that hinges on how bad things could be — the answer will look more confident than the evidence supports, and the omission is invisible from the objective value alone.
  • Is one direction simply better, if computation were free?
    No — they answer different questions. Mass-covering is right when failing to represent a real region is the expensive error; mode-seeking is right when a sharp, internally consistent answer is what gets acted on and a fit between two modes would be typical of neither. What is not defensible is picking a direction for tractability and then defending it afterwards as a modelling preference.

Forward KL packs for every city on the itinerary and is over-prepared everywhere; reverse KL packs perfectly for one city and ignores the rest of the trip.

saying these in an interview costs you the question

  • Treats the two KL directions as interchangeable
  • Says KL is symmetric so the direction does not matter
  • Reports a KL number without naming which distribution is the target
  • Assumes a mode-seeking fit gives calibrated uncertainty
  • Picks a direction for convenience and never records why

context