skip to content

In Metropolis-Hastings, why does the unknown normalising constant cancel out of the acceptance ratio?

level: middleimportance: must knowfreq 66%

answer

  1. the rule compares two points
  2. same constant top and bottom
  3. unnormalised posterior is enough
  4. symmetric proposal drops the q terms
  5. log difference, not raw densities

basics

~20 s

The acceptance rule uses the target only as a ratio of its values at the proposed and current points. Both carry the same constant factor, so it cancels and the sampler needs nothing beyond likelihood times prior.

solid answer

~50 s

Metropolis-Hastings proposes `theta'` from a proposal density `q(theta' | theta)` and accepts it with probability `min(1, [p(theta') q(theta | theta')] / [p(theta) q(theta' | theta)])`. With a symmetric proposal, such as a Gaussian random walk, the two `q` terms are equal and this collapses to `min(1, p(theta') / p(theta))`. Write the true posterior as `p_unnorm(theta) / Z`: the `Z` appears in numerator and denominator identically and divides out, so `p` can be the unnormalised posterior — likelihood times prior. That is the whole trick, and it is why an intractable evidence integral never blocks the sampler. In practice you work in logs, comparing `log(u)` for `u` uniform on (0, 1) against `log p_unnorm(theta') - log p_unnorm(theta)`, because raw densities underflow. On rejection you do not skip the iteration: you record the current state again as the next draw.

code

python · 19 lines
python
import math, random

def log_target(t):
    if t <= 0:
        return -math.inf
    return -((t - 1.0) ** 2) / 8.0

random.seed(7)
theta, draws, accepted = 1.0, [], 0
for _ in range(50000):
    prop = theta + random.gauss(0, 2.5)
    log_ratio = log_target(prop) - log_target(theta)
    if random.random() < math.exp(min(0.0, log_ratio)):
        theta, accepted = prop, accepted + 1
    draws.append(theta)

print('acceptance rate:', accepted / 50000)
print('estimated mean:', sum(draws) / len(draws))
print('estimated P(theta > 2):', sum(t > 2 for t in draws) / len(draws))

go deeper

for a junior

Be ready to state that the acceptance test compares the unnormalised posterior at two points, and that the unknown constant divides out because it appears identically above and below.

for a middle

Explain the full rule including the proposal-density ratio, show the algebra of the cancellation, and describe what happens on a rejection and why the test is done in logs.

for a senior

Demonstrate that you can spot the failure modes in someone else's implementation: a missing Hastings correction on an asymmetric proposal, underflow from raw densities, or a rejection path that quietly drops the draw.

for a principal

Own the wider call: when a hand-rolled Metropolis step is the right level of control for a model, and what you sacrifice by choosing a sampler whose output cannot give you the marginal likelihood you may later want for model comparison.

## The rule Metropolis-Hastings builds a chain over parameter space from two ingredients: a proposal distribution and an acceptance test. Starting at the current value `theta`, one iteration is: 1. Draw a candidate `theta'` from a proposal density `q(theta' | theta)`. 2. Compute the acceptance probability ``` alpha = min(1, [ p(theta') * q(theta | theta') ] / [ p(theta) * q(theta' | theta) ]) ``` 3. Draw `u` uniform on (0, 1). If `u < alpha`, the next state is `theta'`; otherwise the next state is `theta` again. Step 3 matters more than beginners expect. A rejection is not a skipped iteration — the current value is written into the output a second time. A chain that rejects heavily therefore contains long runs of repeated identical values. ## Where the constant goes The target `p` in that formula is the distribution you want to sample, here the posterior: ``` p(theta) = p_unnorm(theta) / Z, p_unnorm(theta) = p(y | theta) * p(theta), Z = integral of p_unnorm ``` Substitute this into the ratio: ``` p(theta') / p(theta) = [ p_unnorm(theta') / Z ] / [ p_unnorm(theta) / Z ] = p_unnorm(theta') / p_unnorm(theta) ``` The `Z` divides out exactly, because it is the same number in both places — it does not depend on `theta`. So the sampler only ever needs to evaluate likelihood times prior at two points and compare them. This is the reason Metropolis-Hastings is usable at all on real posteriors: the evidence integral that makes the posterior intractable in closed form never appears in the algorithm. The same cancellation happens with any target known up to a constant, which is why the method also shows up outside Bayesian inference wherever an unnormalised score function is available. ## The proposal correction term The `q` ratio, sometimes called the Hastings correction, is what keeps the chain honest when the proposal is not symmetric. Symmetric means `q(theta' | theta) = q(theta | theta')` — true for `theta' = theta + noise` with zero-mean symmetric noise, which is the common random-walk case. Then the correction is 1 and the rule reduces to the original Metropolis form `min(1, p(theta') / p(theta))`. With an asymmetric proposal — a lognormal step for a positive parameter, an independence proposal from some fitted approximation, a move that reflects off a boundary — the correction is not 1, and dropping it is a genuine bug rather than a rounding issue. The chain will still run and still look plausible, but it will converge to the wrong distribution, because the asymmetry biases which moves get offered and nothing compensates for it. The acceptance ratio is constructed so that the chain satisfies detailed balance with respect to the target; the correction term is exactly the piece that restores balance when proposals are lopsided. ## Working in logs On any real model, `p_unnorm` is a product of many densities and underflows to zero in double precision long before the parameters get extreme. Every implementation therefore computes a log unnormalised posterior — the log-likelihood plus the log-prior, both sums — and tests ``` log(u) < log p_unnorm(theta') - log p_unnorm(theta) ``` which is algebraically equivalent to `u < alpha` for the symmetric case and numerically stable. A parameter value outside the prior's support is handled by returning minus infinity from the log target, which makes the difference minus infinity and the proposal certain to be rejected. ## Reading the acceptance behaviour Two cases are worth naming explicitly. If the proposal has higher unnormalised posterior density than the current point, the ratio exceeds 1, `alpha` is 1, and the move is always accepted — the chain always goes uphill when offered. If the proposal is worse, it is accepted with probability equal to the ratio, so a slightly worse point is usually accepted and a much worse point is usually rejected. That downhill acceptance is not a flaw to be tuned away: it is what makes the chain explore a distribution rather than climb to the mode. A candidate who says the rule is accept-if-better is describing hill climbing, which finds a maximum and tells you nothing about posterior spread. ## What the trick does not buy you The cancellation makes the sampler runnable, not efficient. Nothing in it guarantees that the chain moves usefully — that depends on how proposals are generated, and a badly scaled proposal produces a valid chain that explores hopelessly slowly. The cancellation also gives you no access to `Z` itself; if you actually need the marginal likelihood, for example to compare models, plain Metropolis-Hastings output does not hand it to you, and you need a different technique.

  • What happens if the proposal is asymmetric and you drop the proposal-density ratio?
    The chain converges to the wrong distribution. The correction term compensates for a proposal that offers moves in one direction more readily than the reverse; without it the asymmetry leaks into the stationary behaviour. The run will still look healthy, which is what makes the bug dangerous — nothing crashes, the summaries are simply biased.
  • Why implement the acceptance test with logarithms?
    The unnormalised posterior is a product of many densities and underflows to zero in double precision. Working with the log-likelihood plus the log-prior turns the product into a sum, and the test becomes log(u) < log p(theta') - log p(theta). A point outside the prior support returns minus infinity, which rejects it automatically.
  • If the proposed point has higher unnormalised posterior density than the current one, what happens?
    The ratio exceeds one, so the acceptance probability is capped at one and the move is always accepted. Downhill moves are still accepted with probability equal to the density ratio — that is deliberate, and it is what separates a sampler from a hill-climbing optimiser that would only ever go up.

Comparing two prices quoted in an unknown currency: you cannot say what either costs, but you can say one is twice the other, because the unknown exchange rate cancels from the ratio.

saying these in an interview costs you the question

  • Thinks the evidence integral must be computed before sampling
  • Drops the proposal ratio for an asymmetric proposal
  • Says the rule accepts only when the proposal is better
  • Multiplies raw densities instead of adding logs
  • Claims a rejected proposal means the iteration produces no draw

context