What does maximising the ELBO achieve in variational inference?
answer
- optimisation instead of integration
- a bound, not the evidence itself
- the gap is a KL divergence
- expected log joint plus entropy of q
basics
~20 sMaximising the evidence lower bound (ELBO) picks the distribution in a chosen tractable family that is closest to the true posterior. The bound's gap to the log evidence is exactly a KL divergence, so raising the ELBO shrinks that gap.
solid answer
~40 sVariational inference replaces integration with optimisation. You choose a tractable family of distributions and search inside it for the `q(z)` closest to the posterior `p(z|x)`. The key identity is `log p(x) = ELBO(q) + KL(q || p(z|x))`, with `ELBO(q) = E_q[log p(x,z)] - E_q[log q(z)]`. Since `log p(x)` does not depend on `q` and KL is never negative, the ELBO is a lower bound on the log evidence, and maximising it is exactly equivalent to minimising `KL(q || p(z|x))` — which you cannot minimise directly, because writing it out needs the unknown normalising constant `p(x)`. A useful regrouping is `ELBO = E_q[log p(x|z)] - KL(q(z) || p(z))`: fit the data, but do not drift far from the prior. What you get back is a fitted distribution, not draws from the true posterior.
go deeper
Be ready to say what the ELBO is in one line: a lower bound on the log marginal likelihood that variational methods maximise instead of computing the posterior exactly.
Expect to derive log p(x) = ELBO + KL(q || p(z|x)) on a whiteboard and explain why a fixed log p(x) makes bound-maximisation and divergence-minimisation the same problem.
Show that you monitor the ELBO like a training curve but never treat convergence as accuracy: the gap you cannot measure is exactly the error, so accuracy claims need an external check.
Own the argument about when an unquantified approximation gap is acceptable for the decisions the model feeds, and set the team's policy for validating variational fits before they reach production.
## The problem Bayes' rule gives the posterior over unknown parameters or latent variables `z` after seeing data `x`: `p(z|x) = p(x|z) p(z) / p(x)`, where `p(x) = integral of p(x|z) p(z) dz`. The numerator (the joint density `p(x,z) = p(x|z)p(z)`) is usually easy to evaluate for any given `z`. The denominator `p(x)` — the marginal likelihood, also called the evidence — is an integral over the whole parameter space and is almost never available in closed form. That single missing constant is why the posterior is hard. Variational inference (VI) attacks this by **turning inference into optimisation**. Instead of trying to characterise `p(z|x)` exactly, you fix a family `Q` of distributions that you can write down, sample from and integrate against — for example all Gaussians with diagonal covariance — and you search inside `Q` for the member closest to the posterior. ## Closeness measured by KL "Closest" is defined by the Kullback-Leibler divergence, `KL(q || p) = E_q[log q(z) - log p(z)]`, a non-negative quantity that is zero only when the two densities agree almost everywhere. VI minimises `KL(q(z) || p(z|x))`: the expectation is taken under the approximating distribution `q`, which is the one you control. You cannot minimise that objective directly, because expanding it gives `E_q[log q(z)] - E_q[log p(x,z)] + log p(x)`, and the last term is the intractable evidence. The trick is that `log p(x)` is a constant with respect to `q`, so you can drop it and optimise what is left, with the sign flipped. ## The ELBO and the decomposition Define the **evidence lower bound**: `ELBO(q) = E_q[log p(x,z)] - E_q[log q(z)]` Then for any `q`: `log p(x) = ELBO(q) + KL(q || p(z|x))` This is an exact identity, not an approximation. Two consequences follow immediately: 1. Because `KL >= 0`, `ELBO(q) <= log p(x)` for every `q` — hence "lower bound on the evidence". 2. Because `log p(x)` is fixed, **every unit of ELBO you gain is a unit of KL you remove**. Maximising the bound and minimising the divergence are the same optimisation. Two regroupings of the same quantity are worth memorising: - `ELBO = E_q[log p(x|z)] - KL(q(z) || p(z))` — an expected log-likelihood (fit the data) minus a penalty for straying from the prior. This form makes the regularisation interpretation obvious. - `ELBO = E_q[log p(x,z)] + H[q]`, where `H[q] = -E_q[log q(z)]` is the entropy of `q`. The entropy term is what stops the optimiser from collapsing `q` onto a single point: a point mass would maximise the expected log joint but has no spread at all. ## How it is optimised in practice With a factorised family and a conditionally conjugate model, the optimal update for each factor has a closed form and you can cycle through the factors, raising the ELBO monotonically until it plateaus. For models without that structure, the ELBO and its gradient are estimated from Monte Carlo draws of `q` and maximised by stochastic gradient ascent, which lets you subsample the data and scale to datasets far beyond what exact methods handle. ## Reading the converged number The ELBO is monitored during fitting the way a training loss is: it should rise and flatten. But three cautions matter in interviews. - **A converged ELBO is not a certificate of accuracy.** You know the bound's value, not the size of the gap, because the gap is exactly the quantity you could not compute. - **Comparing ELBOs across models compares bounds, not evidences.** A model with a looser bound can lose to a worse model with a tighter one. - **The objective is non-convex in general**, so different initialisations can converge to different optima; restarts are cheap insurance. ## What you actually get back The output of VI is a fitted distribution with known parameters — means, variances, mixing weights — from which you can compute summaries analytically or by drawing cheaply. Those draws come from `q`, not from the posterior. Everything downstream inherits whatever error `q` carries, which is why the shape of the chosen family, and the direction of the KL being minimised, are the things an interviewer will probe next.
- Why can't you minimise the KL divergence to the posterior directly?Writing `KL(q || p(z|x))` out gives `E_q[log q(z)] - E_q[log p(x,z)] + log p(x)`, and the last term is the intractable marginal likelihood you were avoiding in the first place. Since it does not depend on `q`, you can drop it and optimise the remaining two terms — that is exactly the ELBO, negated. The bound exists to sidestep the missing normaliser.
- What does the entropy term of the ELBO contribute?Written as `E_q[log p(x,z)] + H[q]`, the ELBO trades fit against spread. The expected log joint alone is maximised by concentrating all mass at the highest-density point; the entropy term rewards `q` for staying spread out. Removing it would collapse the approximation to a point and destroy any notion of uncertainty.
- Is a higher ELBO always a better posterior approximation?Within one family and one model, a higher ELBO does mean smaller KL, so yes. Across different approximating families or different models it does not: you are comparing two lower bounds whose gaps to their own evidences differ, so the ranking of bounds need not match the ranking of the quantities they bound.
It is like fitting the best-matching stock template over an irregular shape: you slide and stretch the template to overlap as much as you can, and the leftover mismatch is error you never see reported, only the quality of the fit you achieved.
saying these in an interview costs you the question
- Calls the ELBO an upper bound on the log evidence
- Says maximising the ELBO recovers the exact posterior
- Confuses the ELBO with the log-likelihood of the data
- Thinks variational inference draws samples from the true posterior
- Treats a plateaued ELBO as proof the approximation is accurate