skip to content

How does the Laplace approximation build a Gaussian approximation to a posterior?

level: middleimportance: nice to knowfreq 24%

answer

  1. a quadratic fit at the peak
  2. gradient vanishes at a maximum
  3. curvature becomes the precision
  4. second derivatives inverted give the covariance

basics

~20 s

It expands the log posterior to second order around its peak. The linear term vanishes there, leaving a quadratic, which is the log of a Gaussian centred at the peak whose covariance is the inverse of the negative second-derivative matrix.

solid answer

~50 s

Take the unnormalised log posterior, find the parameter value that maximises it, and Taylor-expand around that point. The constant term is absorbed into normalisation and the gradient is zero at a maximum, so the expansion reduces to `log p(z|x) ≈ const - 0.5 (z - m)' H (z - m)`, where `m` is the maximiser and `H` is the negative Hessian — the matrix of second derivatives of the log posterior, sign-flipped — evaluated at `m`. That quadratic form is exactly the log density of a Gaussian with mean `m` and covariance `H` inverse. So the curvature of the log posterior at its peak supplies the uncertainty: sharp curvature gives a small covariance, flat curvature a large one. It is cheap, needs no sampling and no optimisation over a family, and is a reasonable approximation when the posterior is smooth, unimodal and roughly symmetric.

go deeper

for a junior

Know the one-line description: fit a Gaussian at the peak of the log posterior, with its width set by how sharply the surface curves there.

for a middle

Be able to write the second-order expansion, explain why the linear term drops out at the peak, and state that the covariance is the inverse of the negative second-derivative matrix.

for a senior

Show judgment about when it misleads — bounded, skewed, heavy-tailed or multimodal posteriors — and reach for an unconstrained reparameterisation before applying it.

for a principal

Own the tradeoff between a nearly free curvature-based approximation and costlier methods, and decide where in the stack a locally-fitted Gaussian is good enough.

## The construction Write the log of the unnormalised posterior as `f(z) = log p(x|z) + log p(z)`, which you can evaluate without knowing the intractable normalising constant. Let `m` be the value of `z` that maximises `f`. Expand `f` in a second-order Taylor series about `m`: `f(z) ≈ f(m) + g'(z - m) - 0.5 (z - m)' H (z - m)` where `g` is the gradient of `f` at `m` and `H` is the negative of the matrix of second partial derivatives of `f` at `m`. Two simplifications apply. First, `g = 0`, because `m` is an interior maximum, so the linear term disappears — this is why the expansion is taken at the peak rather than anywhere else. Second, `f(m)` is a constant, which is swallowed by normalisation. What is left is a pure quadratic in `z`, and exponentiating a negative-definite quadratic gives an unnormalised Gaussian density. Reading off the parameters: **Mean = `m`, covariance = `H` inverse.** The covariance comes from curvature alone. If the log posterior falls away steeply from its peak, `H` is large, its inverse is small, and the approximation reports tight uncertainty. If the surface is flat near the peak, `H` is small, its inverse is large, and the approximation reports diffuse uncertainty. In more than one dimension `H` also carries the off-diagonal terms, so unlike a factorised approximation the Laplace fit **does** capture local correlation between parameters — the fitted ellipse can tilt. ## Where it sits relative to optimisation-based approximation Both Laplace and the divergence-minimising approach replace a hard integral with a cheap deterministic computation, but they differ in how the approximation is chosen. Laplace does no search over a family: it matches the target locally at a single point, using the value and curvature there. The optimisation-based approach searches a family for a global compromise under a divergence, which weighs the whole distribution rather than one point. Consequences: Laplace is faster and has fewer knobs, but nothing about it optimises global fit; the divergence-based fit can be better where the target is asymmetric, at a higher cost. ## Why it works when it works The justification is asymptotic. With increasing data and a well-behaved, identified model, the posterior concentrates and becomes increasingly Gaussian around its peak — the log-likelihood surface becomes locally quadratic and the prior's influence recedes. In that regime the second-order expansion is capturing almost everything there is, and the approximation is very good for very little work. ## Where it breaks - **Skewness.** A quadratic is symmetric; a skewed posterior is not. Variance components and rate parameters bounded below at zero are the classic offenders: the true posterior has a long right tail and a hard left boundary, and a symmetric Gaussian both understates the tail and puts probability on impossible negative values. - **Boundaries.** If the maximum sits at or near an edge of the parameter space — a proportion near 0 or 1, a variance near 0 — the gradient need not vanish there and the whole construction is invalid. - **Multimodality.** The expansion sees one peak. Other modes contribute nothing, and their probability mass is silently discarded. - **Heavy tails.** Gaussian tails decay very fast; a posterior with heavier tails will have its extreme quantiles badly understated, which matters precisely for the risk questions people care about. - **Small samples.** The asymptotic argument has no force when the data are few and the prior still shapes the posterior strongly. - **Dimension.** Forming, storing and inverting the second-derivative matrix costs on the order of the squared and cubed parameter count respectively, which becomes prohibitive for large models and often forces a diagonal approximation — at which point the local correlation advantage is given up. ## Practical handling The standard remedy for skew and boundaries is to **transform to an unconstrained scale** before expanding: take logs of positive parameters, use a log-odds scale for probabilities, expand there, and transform draws back afterwards. The posterior on the transformed scale is typically far more symmetric, which is exactly the regime the method assumes. It is also worth checking that the fitted point really is a maximum — the second-derivative matrix must be positive definite after the sign flip — and comparing the approximation against a more faithful method on a smaller version of the problem before trusting it in production. The interview-ready summary: expand the log posterior to second order at its peak, the linear term dies, the curvature becomes the precision, and the result is a Gaussian that is excellent for smooth well-identified problems with plenty of data and misleading for skewed, bounded, multimodal or heavy-tailed ones.

  • Why does the Laplace approximation capture correlation when a factorised approximation cannot?
    The second-derivative matrix has off-diagonal entries, and inverting the full matrix produces a covariance with non-zero off-diagonal terms. The fitted ellipse can therefore tilt to follow a correlated ridge locally. A factorised family forbids those terms by construction, whatever the curvature says.
  • What would you do before applying it to a positive variance parameter?
    Move to an unconstrained scale first — expand on the log of the parameter rather than the parameter itself. The posterior there is far more symmetric, the maximum is interior rather than pressed against zero, and back-transforming gives a skewed approximation on the original scale instead of a symmetric one that allows negative values.
  • How does its cost scale with the number of parameters?
    Forming the second-derivative matrix is quadratic in the parameter count and inverting it is roughly cubic, so it becomes impractical for large models. Practitioners then fall back to a diagonal or otherwise structured curvature approximation, which restores tractability but throws away the correlation the full matrix captured.

saying these in an interview costs you the question

  • Says the covariance comes from the gradient rather than the curvature
  • Uses the curvature directly as the covariance without inverting
  • Applies it to a bounded parameter sitting at its boundary
  • Assumes it is accurate for skewed or multimodal posteriors
  • Thinks it captures every mode of the posterior

context