skip to content

Posterior Computation

Most posteriors have no closed form, so you approximate them by sampling with Metropolis-Hastings, Gibbs or Hamiltonian moves, or by optimising a bound. Interviewers ask what breaks.

on this pageshow

explore

questions

15

What does MCMC give you when a posterior has no closed-form normalising constant?

level: juniorimportance: must knowfreq 72%

answer

  1. the denominator is the problem
  2. a high-dimensional integral, no closed form
  3. simulate instead of integrate
  4. summaries become averages over draws
  5. dependent draws, not independent ones

basics

~20 s

MCMC produces a stream of parameter draws that visit each region in proportion to posterior probability. You never compute the normalising integral: every posterior summary, such as a mean or a tail probability, becomes an average over those draws.

solid answer

~50 s

Bayes rule gives `p(theta | y) = p(y | theta) p(theta) / Z`, where the normaliser `Z` is the integral of likelihood times prior over the whole parameter space. For something like a 12-parameter logistic-regression posterior with no conjugate form, that is a 12-dimensional integral with no closed form and no practical grid approximation. MCMC sidesteps it: you construct a Markov chain that visits parameter values in proportion to the unnormalised posterior, run it, and keep the states it passes through. Summaries then come from the draws — `E[theta_j]` is the sample average of `theta_j` across draws, `P(theta_j > 0.5)` is the fraction of draws above 0.5, and a credible interval is a pair of empirical quantiles. The price is that consecutive draws are dependent, so a given number of draws carries less information than the same number of independent ones.

go deeper

for a junior

Be ready to state Bayes rule, point at the denominator as the intractable piece, and say that MCMC gives you draws from which means, probabilities and intervals are computed as sample summaries.

for a middle

Explain why the denominator is an integral over the whole parameter space, why grid methods fail as dimension grows, and how a specific summary such as P(theta > 0.5) is read off the draws.

for a senior

Show the operating judgment: how many draws a reported number needs, why dependence between draws makes raw counts misleading, and what a chain that has not reached the bulk of the posterior does to your conclusions.

for a principal

Own the framing decision: when a full posterior is worth its compute budget versus when a point estimate answers the business question, and what you commit to when a nightly job must produce posterior summaries on a deadline.

## The problem MCMC exists to solve Bayesian inference states the answer in one line: ``` p(theta | y) = p(y | theta) * p(theta) / Z, Z = integral over theta of p(y | theta) p(theta) d theta ``` Here `theta` is the vector of unknown parameters, `y` is the observed data, `p(y | theta)` is the likelihood, `p(theta)` is the prior, and `Z` (the marginal likelihood, or evidence) is the constant that makes the posterior integrate to one. The numerator is cheap: for any candidate `theta` you can evaluate likelihood times prior directly, usually as a sum of log terms. The denominator is the hard part, because it is an integral over the entire parameter space. In a handful of textbook cases the integral has a closed form — a conjugate prior paired with its likelihood returns a posterior in the same family, and you can read the answer off. Realistic models rarely cooperate. Take a logistic regression with 12 coefficients and a normal prior on each: the likelihood involves a product of logistic terms, the prior is Gaussian, and the product integrates to nothing you can write down. Numerical quadrature does not rescue you either, because grid cost grows exponentially in dimension: a coarse 20-point grid per coordinate already means 20^12, roughly 4 * 10^15 evaluations, for a model most people would call small. ## The reframing: you rarely want the density itself The key observation is that a density function is almost never what a question actually needs. What you report is a summary: - a posterior mean or median for a coefficient, - a probability such as `P(theta_3 > 0.5)`, - a credible interval, - the posterior mean of some derived quantity, such as a predicted probability. Every one of those is an expectation or a quantile under the posterior. And expectations can be estimated from draws without ever knowing the density's scale. If `theta^(1), ..., theta^(S)` are draws whose long-run distribution is the posterior, then for any function `g`: ``` E[g(theta) | y] is approximated by (1/S) * sum over s of g(theta^(s)) ``` Set `g(theta) = theta_j` and you get the posterior mean of coefficient `j`. Set `g(theta) = 1 if theta_j > 0.5 else 0` and the average is exactly the fraction of draws above 0.5 — an estimate of `P(theta_j > 0.5 | y)`. Sort the draws of one coordinate and take the 5th and 95th percentiles and you have a 90 percent credible interval. Nothing in any of these calculations refers to `Z`. ## Where the draws come from That leaves the question of how to draw from a distribution you can only evaluate up to an unknown constant. Markov chain Monte Carlo answers it by giving up on independent draws. Instead of sampling from the posterior directly, you build a rule for stepping from the current parameter vector to the next one, designed so that the chain spends time in each region of parameter space in proportion to that region's posterior probability. Run it long enough and the sequence of visited states behaves, for averaging purposes, like a sample from the posterior. Metropolis-Hastings, Gibbs sampling and Hamiltonian Monte Carlo are three such stepping rules; they differ in how a move is proposed and whether it can be rejected, not in what they target. Crucially, every one of these rules only ever needs the unnormalised posterior — likelihood times prior — because the moves depend on comparing two candidate values, and the unknown constant scales both sides identically. ## What you give up MCMC output is an approximation with two distinct error sources, and a candidate who says otherwise is glossing over the hard part. First, Monte Carlo error: with a finite number of draws, the sample average is not the true posterior mean. More draws shrink this error, and you should report summaries with enough draws that the estimate is stable to the precision you plan to quote. Second, dependence: successive states of the chain are correlated by construction, since each is a step away from the last. Ten thousand MCMC draws therefore carry strictly less information than ten thousand independent draws, sometimes dramatically less. That is why raw draw counts are a poor measure of how much you have actually learned. Third, and easy to forget: the guarantee is asymptotic. A chain started far from the bulk of the posterior needs time to reach it, and a chain that never finds a second mode will produce confident, wrong summaries. Deciding whether a particular run has actually got there is a separate discipline from choosing a sampler. ## The practical takeaway MCMC converts an intractable integration problem into a simulation problem. You give up closed-form answers and independent samples; you get the ability to work with any posterior you can evaluate up to a constant, at whatever dimension the model requires. For most applied Bayesian work, that trade is the whole reason the method exists.

  • Given posterior draws for one coefficient, how do you report a 90 percent credible interval?
    Sort that coefficient's draws and take the 5th and 95th empirical percentiles; the interval between them is the central 90 percent credible interval. No density evaluation and no normalising constant is involved — it is a quantile of the sample you already have. Report enough draws that the endpoints are stable to the precision you quote.
  • Why not just evaluate the unnormalised posterior on a grid instead of sampling?
    Grid cost grows exponentially with the number of parameters. Even a coarse 20 points per coordinate over 12 parameters is 20^12, around 4 * 10^15 evaluations, and most of that grid sits in regions with negligible posterior mass. Sampling spends effort where the mass is, which is why it scales and quadrature does not.
  • Are MCMC summaries exact?
    No. They carry Monte Carlo error that shrinks as you take more draws, and because consecutive draws are dependent, a given number of draws is worth less than the same number of independent ones. Treat every reported posterior mean or probability as an estimate with its own uncertainty, not an exact number.

saying these in an interview costs you the question

  • Says MCMC computes the exact posterior density
  • Claims you must evaluate the normalising constant first
  • Treats consecutive MCMC draws as independent samples
  • Confuses a posterior distribution with a single point estimate
  • Thinks a fine grid would work fine in 12 dimensions

context

open as a page

Why does an MCMC run with 10,000 stored draws report an effective sample size of only 200?

level: middleimportance: must knowfreq 64%

basics

~20 s

Because MCMC draws are dependent. When each draw is nearly a repeat of the previous one, 10,000 correlated draws carry about as much information as 200 independent ones, and effective sample size reports that count rather than the stored count.

open as a page

What does the Gelman-Rubin R-hat statistic compare across MCMC chains?

level: middleimportance: must knowfreq 76%

basics

~20 s

R-hat compares the variation between several MCMC chains with the variation within each chain. If the chains have converged to the same distribution the two agree and R-hat sits near 1; values above roughly 1.01 mean the chains still disagree.

open as a page

In Metropolis-Hastings, why does the unknown normalising constant cancel out of the acceptance ratio?

level: middleimportance: must knowfreq 66%

basics

~20 s

The acceptance rule uses the target only as a ratio of its values at the proposed and current points. Both carry the same constant factor, so it cancels and the sampler needs nothing beyond likelihood times prior.

open as a page

What does maximising the ELBO achieve in variational inference?

level: middleimportance: must knowfreq 62%

basics

~20 s

Maximising the evidence lower bound (ELBO) picks the distribution in a chosen tractable family that is closest to the true posterior. The bound's gap to the log evidence is exactly a KL divergence, so raising the ELBO shrinks that gap.

open as a page

Why are the warm-up (burn-in) draws of an MCMC chain discarded before summarising the posterior?

level: juniorimportance: should knowfreq 58%

basics

~20 s

Early MCMC draws still reflect where the chain was started rather than the posterior, and in adaptive samplers the tuning is still changing. Those draws are dropped so the summary uses only draws from the stationary distribution.

open as a page

Why does a Gibbs sampler need no accept-reject step when it draws from full conditionals?

level: middleimportance: should knowfreq 48%

basics

~20 s

A Gibbs sampler draws each block from its exact full conditional, its distribution given the current values of all other parameters and the data. Because the proposal is already exact, the Metropolis acceptance probability equals one.

open as a page

Why does a mean-field variational approximation understate posterior variance?

level: middleimportance: should knowfreq 52%

basics

~20 s

A mean-field family factorises the approximation into independent pieces, so it cannot represent correlation between parameters. On a correlated posterior it fits inside the narrow direction of the ridge rather than along it, producing spreads that are too small.

open as a page

A hierarchical model's sampler reports divergent transitions clustered where the group-level scale is near zero. What do you do?

level: seniorimportance: should knowfreq 43%

basics

~10 s

Treat the fit as biased, not noisy. Divergences near a near-zero scale parameter signal the funnel geometry a centred hierarchical parameterisation creates; the fix is a non-centred parameterisation, then a smaller step size.

open as a page

Your random-walk Metropolis sampler accepts 2% of proposals — what is wrong and how do you fix it?

level: seniorimportance: should knowfreq 55%

basics

~20 s

The proposal steps are far too large, so almost every candidate lands in a low-density region and is rejected and the chain sits still for long stretches. Shrink the proposal scale until acceptance rises to roughly a quarter.

open as a page

Why does variational inference minimise KL(q||p) rather than KL(p||q)?

level: seniorimportance: should knowfreq 38%

basics

~20 s

KL(q||p) takes expectations under the approximation you control, so it is computable up to a constant; the reverse direction needs expectations under the unknown posterior. The price is mode-seeking behaviour that ignores parts of a multimodal target.

open as a page

How do you decide between variational inference and MCMC for a production Bayesian model?

level: principalimportance: should knowfreq 35%

basics

~20 s

Decide by which error you can afford. Variational inference is fast and scales, but its approximation error is biased and unmeasured; sampling is asymptotically exact but slow. Match the choice to how sensitive the downstream decision is to uncertainty.

open as a page

Does thinning an MCMC chain to every 10th draw improve the posterior estimates?

level: middleimportance: nice to knowfreq 33%

basics

~20 s

No. For a fixed number of iterations, keeping every tenth draw throws away information, so estimates are no better and usually slightly noisier than using all draws. Thinning is justified by storage or downstream cost, not accuracy.

open as a page

How does the Laplace approximation build a Gaussian approximation to a posterior?

level: middleimportance: nice to knowfreq 24%

basics

~20 s

It expands the log posterior to second order around its peak. The linear term vanishes there, leaving a quadratic, which is the log of a Gaussian centred at the peak whose covariance is the inverse of the negative second-derivative matrix.

open as a page

What does Hamiltonian Monte Carlo's momentum variable buy over random-walk Metropolis proposals?

level: seniorimportance: nice to knowfreq 33%

basics

~20 s

Momentum lets a proposal travel a long way while following the posterior's shape, so moves are coherent strides rather than a random walk. Distant proposals still get accepted, which is why exploration holds up as the number of parameters grows.

open as a page