What does MCMC give you when a posterior has no closed-form normalising constant?
answer
- the denominator is the problem
- a high-dimensional integral, no closed form
- simulate instead of integrate
- summaries become averages over draws
- dependent draws, not independent ones
basics
~20 sMCMC produces a stream of parameter draws that visit each region in proportion to posterior probability. You never compute the normalising integral: every posterior summary, such as a mean or a tail probability, becomes an average over those draws.
solid answer
~50 sBayes rule gives `p(theta | y) = p(y | theta) p(theta) / Z`, where the normaliser `Z` is the integral of likelihood times prior over the whole parameter space. For something like a 12-parameter logistic-regression posterior with no conjugate form, that is a 12-dimensional integral with no closed form and no practical grid approximation. MCMC sidesteps it: you construct a Markov chain that visits parameter values in proportion to the unnormalised posterior, run it, and keep the states it passes through. Summaries then come from the draws — `E[theta_j]` is the sample average of `theta_j` across draws, `P(theta_j > 0.5)` is the fraction of draws above 0.5, and a credible interval is a pair of empirical quantiles. The price is that consecutive draws are dependent, so a given number of draws carries less information than the same number of independent ones.
go deeper
Be ready to state Bayes rule, point at the denominator as the intractable piece, and say that MCMC gives you draws from which means, probabilities and intervals are computed as sample summaries.
Explain why the denominator is an integral over the whole parameter space, why grid methods fail as dimension grows, and how a specific summary such as P(theta > 0.5) is read off the draws.
Show the operating judgment: how many draws a reported number needs, why dependence between draws makes raw counts misleading, and what a chain that has not reached the bulk of the posterior does to your conclusions.
Own the framing decision: when a full posterior is worth its compute budget versus when a point estimate answers the business question, and what you commit to when a nightly job must produce posterior summaries on a deadline.
## The problem MCMC exists to solve Bayesian inference states the answer in one line: ``` p(theta | y) = p(y | theta) * p(theta) / Z, Z = integral over theta of p(y | theta) p(theta) d theta ``` Here `theta` is the vector of unknown parameters, `y` is the observed data, `p(y | theta)` is the likelihood, `p(theta)` is the prior, and `Z` (the marginal likelihood, or evidence) is the constant that makes the posterior integrate to one. The numerator is cheap: for any candidate `theta` you can evaluate likelihood times prior directly, usually as a sum of log terms. The denominator is the hard part, because it is an integral over the entire parameter space. In a handful of textbook cases the integral has a closed form — a conjugate prior paired with its likelihood returns a posterior in the same family, and you can read the answer off. Realistic models rarely cooperate. Take a logistic regression with 12 coefficients and a normal prior on each: the likelihood involves a product of logistic terms, the prior is Gaussian, and the product integrates to nothing you can write down. Numerical quadrature does not rescue you either, because grid cost grows exponentially in dimension: a coarse 20-point grid per coordinate already means 20^12, roughly 4 * 10^15 evaluations, for a model most people would call small. ## The reframing: you rarely want the density itself The key observation is that a density function is almost never what a question actually needs. What you report is a summary: - a posterior mean or median for a coefficient, - a probability such as `P(theta_3 > 0.5)`, - a credible interval, - the posterior mean of some derived quantity, such as a predicted probability. Every one of those is an expectation or a quantile under the posterior. And expectations can be estimated from draws without ever knowing the density's scale. If `theta^(1), ..., theta^(S)` are draws whose long-run distribution is the posterior, then for any function `g`: ``` E[g(theta) | y] is approximated by (1/S) * sum over s of g(theta^(s)) ``` Set `g(theta) = theta_j` and you get the posterior mean of coefficient `j`. Set `g(theta) = 1 if theta_j > 0.5 else 0` and the average is exactly the fraction of draws above 0.5 — an estimate of `P(theta_j > 0.5 | y)`. Sort the draws of one coordinate and take the 5th and 95th percentiles and you have a 90 percent credible interval. Nothing in any of these calculations refers to `Z`. ## Where the draws come from That leaves the question of how to draw from a distribution you can only evaluate up to an unknown constant. Markov chain Monte Carlo answers it by giving up on independent draws. Instead of sampling from the posterior directly, you build a rule for stepping from the current parameter vector to the next one, designed so that the chain spends time in each region of parameter space in proportion to that region's posterior probability. Run it long enough and the sequence of visited states behaves, for averaging purposes, like a sample from the posterior. Metropolis-Hastings, Gibbs sampling and Hamiltonian Monte Carlo are three such stepping rules; they differ in how a move is proposed and whether it can be rejected, not in what they target. Crucially, every one of these rules only ever needs the unnormalised posterior — likelihood times prior — because the moves depend on comparing two candidate values, and the unknown constant scales both sides identically. ## What you give up MCMC output is an approximation with two distinct error sources, and a candidate who says otherwise is glossing over the hard part. First, Monte Carlo error: with a finite number of draws, the sample average is not the true posterior mean. More draws shrink this error, and you should report summaries with enough draws that the estimate is stable to the precision you plan to quote. Second, dependence: successive states of the chain are correlated by construction, since each is a step away from the last. Ten thousand MCMC draws therefore carry strictly less information than ten thousand independent draws, sometimes dramatically less. That is why raw draw counts are a poor measure of how much you have actually learned. Third, and easy to forget: the guarantee is asymptotic. A chain started far from the bulk of the posterior needs time to reach it, and a chain that never finds a second mode will produce confident, wrong summaries. Deciding whether a particular run has actually got there is a separate discipline from choosing a sampler. ## The practical takeaway MCMC converts an intractable integration problem into a simulation problem. You give up closed-form answers and independent samples; you get the ability to work with any posterior you can evaluate up to a constant, at whatever dimension the model requires. For most applied Bayesian work, that trade is the whole reason the method exists.
- Given posterior draws for one coefficient, how do you report a 90 percent credible interval?Sort that coefficient's draws and take the 5th and 95th empirical percentiles; the interval between them is the central 90 percent credible interval. No density evaluation and no normalising constant is involved — it is a quantile of the sample you already have. Report enough draws that the endpoints are stable to the precision you quote.
- Why not just evaluate the unnormalised posterior on a grid instead of sampling?Grid cost grows exponentially with the number of parameters. Even a coarse 20 points per coordinate over 12 parameters is 20^12, around 4 * 10^15 evaluations, and most of that grid sits in regions with negligible posterior mass. Sampling spends effort where the mass is, which is why it scales and quadrature does not.
- Are MCMC summaries exact?No. They carry Monte Carlo error that shrinks as you take more draws, and because consecutive draws are dependent, a given number of draws is worth less than the same number of independent ones. Treat every reported posterior mean or probability as an estimate with its own uncertainty, not an exact number.
saying these in an interview costs you the question
- Says MCMC computes the exact posterior density
- Claims you must evaluate the normalising constant first
- Treats consecutive MCMC draws as independent samples
- Confuses a posterior distribution with a single point estimate
- Thinks a fine grid would work fine in 12 dimensions