Why does variational inference minimise KL(q||p) rather than KL(p||q)?
answer
- expectations under the one you control
- the unknown normaliser drops out
- mode-seeking versus mass-covering
- zero-forcing punishes mass on empty ground
basics
~20 sKL(q||p) takes expectations under the approximation you control, so it is computable up to a constant; the reverse direction needs expectations under the unknown posterior. The price is mode-seeking behaviour that ignores parts of a multimodal target.
solid answer
~50 sThe choice is dictated by tractability. `KL(q || p(z|x)) = E_q[log q(z) - log p(z|x)]` averages under `q`, which you can sample and evaluate, and the unknown normaliser drops out as a constant — that is exactly what makes the ELBO optimisable. The reverse, `KL(p(z|x) || q)`, averages under the posterior itself, so computing it presupposes the thing you are trying to approximate. The consequence is behavioural: reverse KL is mode-seeking, because putting `q` mass where the target density is near zero is heavily penalised while missing target mass is nearly free. On a bimodal posterior a unimodal `q` will lock onto one mode and ignore the other. Forward KL is mass-covering — it would smear a single `q` across both modes and the low-density valley between them. Neither is right in the abstract; the question is which failure you can live with.
go deeper
Remember that KL is asymmetric and that variational inference uses the direction whose expectation is taken under the approximating distribution rather than the posterior.
Be ready to expand both directions and show that only the reverse one leaves the intractable normaliser as a constant that drops out of the optimisation.
Show that you anticipate mode-seeking in practice: multiple initialisations, checks for multimodality, and honesty that a single confident fit may describe only one region of the posterior.
Own the decision of whether confident-but-partial uncertainty is acceptable for the decisions at stake, and when the risk justifies a costlier method or a richer approximating family.
## The two directions The Kullback-Leibler divergence is not symmetric. For a target `p` and an approximation `q` there are two distinct objectives: - **Reverse KL** (what variational inference uses): `KL(q || p) = E_q[log q(z) - log p(z)]`, averaged under `q`. - **Forward KL**: `KL(p || q) = E_p[log p(z) - log q(z)]`, averaged under `p`. Both are non-negative and both are zero only when the distributions agree. They disagree everywhere else, and which one you pick changes the answer you get. ## The computational reason for the choice Suppose the target is a posterior `p(z|x) = p(x,z) / p(x)`, where the joint `p(x,z)` is easy to evaluate pointwise and the normaliser `p(x)` is intractable. With **reverse KL**, expand: `KL(q || p(z|x)) = E_q[log q(z)] - E_q[log p(x,z)] + log p(x)`. Every expectation is under `q` — a distribution you chose, whose density you can evaluate and from which you can draw at will. The intractable `log p(x)` appears only as an additive constant that does not depend on `q`, so it can be dropped. What remains, negated, is the ELBO. The objective is computable. With **forward KL**, expand: `KL(p(z|x) || q) = E_{p(z|x)}[log p(z|x)] - E_{p(z|x)}[log q(z)]`. Now the expectations are under the posterior. To estimate them you would need draws from the posterior — the very thing that was out of reach. Forward KL is perfectly usable when you *can* sample the target (for example when fitting a tractable approximation to an empirical distribution, which is what maximum likelihood does), but it is circular here. So the direction is not chosen for its statistical virtues. It is chosen because it is the one that can be optimised without the answer. ## The behavioural consequence Each direction fails in a characteristic way, and the asymmetry is visible in where the integrand explodes. **Reverse KL is mode-seeking (zero-forcing).** The integrand is weighted by `q`. Where `q` has mass and `p` is near zero, `log q - log p` is enormous — a severe penalty. Where `p` has mass but `q` does not, the weight `q` is near zero and the contribution vanishes — missing mass is nearly free. So `q` is pushed to stay inside the target's high-density region, even if that means abandoning most of it. On a **bimodal target** — two well-separated humps with a deep valley between them — a unimodal Gaussian `q` minimising reverse KL will typically sit on one hump and ignore the other entirely. Straddling both would place substantial `q` mass in the valley where `p` is near zero, and that is exactly the configuration the objective punishes hardest. It is a locally sensible answer that is globally incomplete: the reported uncertainty describes one mode as if the other did not exist. **Forward KL is mass-covering (zero-avoiding).** The integrand is weighted by `p`. Any region where `p` has mass and `q` is near zero makes `-log q` blow up, so `q` is forced to cover everything the target covers. On the same bimodal target, a unimodal `q` minimising forward KL stretches across both humps, placing considerable mass in the empty valley between them — a distribution that is broad, over-dispersed, and assigns real probability to parameter values the data effectively rules out. Neither behaviour is uniformly better. Mode-seeking gives you a confident description of one region that may be the wrong region; mass-covering gives you a diffuse description that includes regions nothing supports. ## Practical implications - **Multimodality is a genuine risk, not a curiosity.** Mixture models with label-switching symmetry, non-identified parameterisations and some hierarchical models all produce multimodal posteriors. A variational fit will silently pick a mode; run multiple initialisations and check whether they land in different places before believing any single fit. - **Optimisation and modes interact.** Because the objective is non-convex, which mode you find depends on initialisation — the method has no obligation to find the dominant one. - **Enrich the family when modes matter.** A mixture approximating family can represent several modes at once and removes the forced choice, at the cost of a harder optimisation. - **Match the direction to the deliverable.** If you need to be right about where the bulk of the posterior is and can tolerate being confident about one region, reverse KL is fine. If false confidence about a single mode would drive a bad decision, that is an argument for a sampling-based approach or a richer family rather than for switching the divergence, which you cannot do without the posterior anyway. The complete interview answer has both halves: the direction is forced by what is computable, and the behaviour that follows is mode-seeking under-coverage rather than over-coverage.
- In what setting is forward KL actually optimisable?Whenever you can draw from the target. Fitting a model to an empirical data distribution is exactly a forward-KL minimisation, since the samples supply the expectation. For a posterior you cannot sample, the expectation is unavailable, which is why variational inference cannot use that direction.
- How would you detect that a variational fit has locked onto one mode?Refit from several well-separated initialisations and compare both the fitted parameters and the converged objective. Landing in materially different places is direct evidence of multimodality. Posterior predictive checks against held-out data can also expose a fit that describes only part of the plausible parameter space.
- Does a richer approximating family remove the mode-seeking behaviour?It relaxes it rather than removing it. A mixture family can place components on several modes at once, so the forced choice disappears if the optimiser finds them. The zero-forcing pressure remains, so any mode no component reaches is still effectively assigned no probability.
Reverse KL is a cautious tenant who only occupies rooms known to be safe, leaving half the house unused; forward KL is one who insists on covering every room, including the ones with no floor.
saying these in an interview costs you the question
- Says KL is symmetric so the direction does not matter
- Claims the reverse direction was chosen for accuracy
- Thinks reverse KL over-covers a multimodal target
- Cannot say which distribution the expectation is taken under
- Assumes a unimodal fit proves the posterior is unimodal