skip to content

Why does a mean-field variational approximation understate posterior variance?

level: middleimportance: should knowfreq 52%

answer

  1. the family cannot tilt
  2. independence assumed, correlation discarded
  3. conditional spread, not marginal spread
  4. overreach is punished, under-reach is cheap

basics

~20 s

A mean-field family factorises the approximation into independent pieces, so it cannot represent correlation between parameters. On a correlated posterior it fits inside the narrow direction of the ridge rather than along it, producing spreads that are too small.

solid answer

~40 s

Mean-field variational inference assumes the approximation factorises, `q(z1, z2) = q(z1) q(z2)`, so by construction it has zero correlation between the blocks. Take a strongly correlated bivariate posterior — a long, thin, tilted ridge. An axis-aligned factorised Gaussian cannot tilt, so the best it can do is sit at the centre and shrink to roughly the width of the ridge's narrow axis; stretching along either coordinate axis would push mass into the low-density corners. The objective reinforces this: minimising `KL(q || p)` charges heavily for putting `q` mass where the posterior density is near zero, but charges little for missing regions the posterior covers. The practical result is that marginal means are usually reasonable while posterior standard deviations come out too small, so downstream uncertainty is overconfident.

go deeper

for a junior

Know the headline: a factorised approximation assumes the parameters are independent, so it reports uncertainty that is too small when they are not.

for a middle

Be able to draw the tilted ridge and show why an axis-aligned ellipse must shrink to fit inside it, and connect that picture to the conditional-versus-marginal spread distinction.

for a senior

Demonstrate that you check the damage before shipping — comparing against a reference fit, blocking correlated parameters, or reparameterising — rather than quoting variational standard deviations as-is.

for a principal

Own the call on when understated uncertainty is tolerable for the decisions the model drives, and make the accepted approximation explicit wherever those numbers are consumed.

## What "mean-field" means Variational inference approximates a posterior `p(z|x)` by the closest member `q(z)` of some family you choose. The **mean-field** family is the simplest useful choice: partition the unknowns into blocks and force the approximation to factorise across them, `q(z) = q1(z1) q2(z2) ... qK(zK)`. Each factor gets its own parameters and is optimised independently of the others' shapes (they still interact through the objective). With Gaussian factors and scalar blocks, this is a Gaussian with a **diagonal covariance matrix** — axis-aligned, no off-diagonal terms. The attraction is computational: each factor is low-dimensional, updates are often closed-form for conditionally conjugate models, and the cost grows linearly rather than quadratically in the number of parameters. The cost is representational: **whatever correlation the true posterior has is unrepresentable**. ## The geometry of the failure Picture a posterior over two parameters that is strongly positively correlated — a long, thin, diagonally tilted ellipse, a ridge. Its marginal spread along each coordinate is wide, because the ridge extends far in both coordinates. But conditionally, once you fix one coordinate, the other is pinned down to a narrow slice across the ridge. Now fit an axis-aligned two-dimensional Gaussian. It has no ability to tilt. If it stretched to match the wide marginals, it would become a large axis-aligned blob covering the two off-ridge corners where the true posterior density is essentially zero. If it shrinks to roughly the width of the ridge's narrow cross-section, it sits entirely inside the region of high posterior density but reports far too little uncertainty. The optimiser picks the second option, so **the fitted variances approximate conditional spread rather than marginal spread**. The general statement: for a jointly Gaussian target, the mean-field solution recovers the means exactly and produces variances equal to the reciprocals of the diagonal entries of the precision matrix — the conditional variances — which are less than or equal to the true marginal variances, with equality only when the parameters are already uncorrelated. Correlation strength is exactly what drives the shortfall. ## Why the objective pushes that way The divergence being minimised is `KL(q || p(z|x)) = E_q[log q(z) - log p(z|x)]`, with the expectation under `q`. Look at where each region contributes: - Where `q` has mass and the posterior density is near zero, `log p(z|x)` is very negative and the integrand blows up. This is a **large penalty** — the objective is intolerant of overreach. - Where the posterior has mass but `q` has essentially none, the region is weighted by `q` itself, so it contributes almost nothing. Missing posterior mass is **nearly free**. The asymmetry is one-sided in the shrinking direction: `q` is pushed to stay strictly inside the posterior's support and high-density region. Combined with a family that cannot tilt, the result is systematic under-dispersion. This behaviour is often described as zero-forcing: wherever the target is near zero, `q` must be near zero too. ## What is and is not damaged - **Point summaries survive reasonably well.** Posterior means from a mean-field fit are often close enough for prediction and ranking, which is why the method is used at all. - **Uncertainty is optimistic.** Posterior standard deviations, tail probabilities, and any decision rule sensitive to spread inherit the shrinkage. Probabilities like "chance the effect exceeds zero" are pushed towards 0 or 1. - **Joint statements are worst.** Any question about two parameters simultaneously — the probability both exceed a threshold, the spread of a sum or difference — depends on the correlation the family threw away, and can be badly wrong in either direction. ## What to do about it 1. **Enrich the family.** Allow a full covariance matrix, or group tightly coupled parameters into a single block so their dependence lives inside one factor rather than across the factorisation. This costs more per iteration but removes the structural constraint. 2. **Reparameterise to decorrelate.** Centring predictors, standardising scales, or a non-centred parameterisation of a hierarchical model can turn a tilted ridge into something close to axis-aligned, at which point the mean-field assumption becomes almost harmless. 3. **Check against a reference.** Fit an exact or sampling-based posterior on a subsample or a simplified version of the model and compare spreads; a consistent factor-of-two shrinkage tells you how far to distrust the reported uncertainty. 4. **Match the method to the use.** If the deliverable is a point prediction, the shrinkage may not matter at all. If the deliverable is a risk statement, it matters a great deal. The answer an interviewer wants is the causal chain: factorisation removes correlation, the divergence penalises overreach far more than under-reach, and the two together produce a fit that hugs the high-density core and reports too little uncertainty.

  • Are the posterior means from a mean-field fit biased in the same way as the variances?
    Usually much less. For a Gaussian target the mean-field solution recovers the means exactly and only the covariance is wrong. With skewed or multimodal posteriors the means can shift too, but as a rule the location is the part you can lean on and the spread is the part you cannot.
  • How would you fix an under-dispersed variational fit without abandoning the approach?
    Enrich the family — a full covariance Gaussian, or blocking tightly coupled parameters into one factor so their dependence is modelled inside it. Alternatively reparameterise the model to decorrelate the parameters, since an already axis-aligned posterior makes the factorisation assumption nearly free.
  • Which downstream quantities are hurt most by the discarded correlation?
    Anything joint. The spread of a sum or difference of two parameters, the probability that two effects are simultaneously positive, or a prediction that combines several coefficients all depend on covariance terms that a factorised approximation set to zero, so their uncertainty can be wrong in either direction.

Fitting an axis-aligned rectangle inside a long diagonal corridor: the largest rectangle that stays entirely inside is barely wider than the corridor, even though the corridor stretches a long way in both directions.

saying these in an interview costs you the question

  • Says mean-field assumes the true posterior is independent
  • Claims the fitted variances are too large, not too small
  • Thinks more optimisation iterations remove the shrinkage
  • Believes correlation can be recovered from the fitted factors
  • Reports mean-field uncertainty without any caveat

context