skip to content

Why does a Gibbs sampler need no accept-reject step when it draws from full conditionals?

level: middleimportance: should knowfreq 48%

answer

  1. one block at a time
  2. condition on everything else
  3. the proposal is the exact slice
  4. a special case of Metropolis-Hastings
  5. acceptance ratio collapses to one

basics

~20 s

A Gibbs sampler draws each block from its exact full conditional, its distribution given the current values of all other parameters and the data. Because the proposal is already exact, the Metropolis acceptance probability equals one.

solid answer

~50 s

One Gibbs sweep updates parameters one block at a time. With parameters `theta_1, theta_2, theta_3`, you draw `theta_1` from `p(theta_1 | theta_2, theta_3, y)`, then `theta_2` from `p(theta_2 | theta_1, theta_3, y)` using the value of `theta_1` you just drew, then `theta_3` likewise. Each of these is a full conditional: the posterior for one block with every other parameter held at its current value. Gibbs is a special case of Metropolis-Hastings whose proposal is the full conditional itself; substitute that proposal into the acceptance ratio and the terms cancel to give an acceptance probability of exactly one, so nothing is ever rejected. The catch is that you must be able to sample each full conditional directly, which usually needs conditional conjugacy. If one block's conditional is not a standard distribution, you can put a Metropolis step inside the sweep for that block only.

go deeper

for a junior

Be ready to say what a full conditional is — the distribution of one parameter given all the others and the data — and that a Gibbs sweep cycles through the parameters drawing each in turn.

for a middle

Explain why the acceptance probability collapses to one when the proposal is the exact conditional, and state the requirement that every conditional be directly samplable for plain Gibbs to apply.

for a senior

Show that you can diagnose a Gibbs sampler that is accepting everything and still going nowhere on correlated parameters, and reach for blocking or reparameterisation rather than simply running longer.

for a principal

Own the modelling call: whether to shape a model for conditional conjugacy so Gibbs applies, or accept a harder-to-sample formulation that states the science better and pay for a gradient-based sampler instead.

## The sweep Gibbs sampling attacks a multi-parameter posterior by never dealing with all parameters at once. Split `theta` into blocks — often individual coordinates, sometimes groups that are natural to update together. One sweep visits the blocks in turn and replaces each with a fresh draw conditioned on everything else: ``` theta_1^(new) ~ p(theta_1 | theta_2^(old), theta_3^(old), y) theta_2^(new) ~ p(theta_2 | theta_1^(new), theta_3^(old), y) theta_3^(new) ~ p(theta_3 | theta_1^(new), theta_2^(new), y) ``` Note the bookkeeping in the second line: `theta_2` is conditioned on the value of `theta_1` drawn moments ago, not on last sweep's value. Using stale values is a classic implementation bug that silently changes the distribution being sampled. The state after a complete sweep is one draw; the collection of states across sweeps is the sample you summarise. ## Why nothing is rejected A full conditional is `p(theta_j | theta_-j, y)`, where `theta_-j` denotes every parameter except block `j`. It is proportional to the joint unnormalised posterior with all the other coordinates frozen, so it is the exact conditional slice of the target, not an approximation of it. Gibbs is Metropolis-Hastings with that conditional used as the proposal. Write the acceptance ratio for a move that changes only block `j`: ``` alpha = min(1, [ p(theta') q(theta | theta') ] / [ p(theta) q(theta' | theta) ]) ``` with `q(theta' | theta) = p(theta_j' | theta_-j, y)`. Factor the joint as conditional times marginal, `p(theta) = p(theta_j | theta_-j, y) * p(theta_-j | y)`, and note that `theta_-j` is unchanged by the move, so `p(theta_-j | y)` is identical on both sides. Everything cancels and the ratio equals one. In plain words: you are proposing from precisely the distribution the move is supposed to leave in place, so there is nothing left for a correction to fix. Every proposal is accepted. This is a real advantage — no tuning parameter, no acceptance rate to monitor, no wasted proposals — and a real constraint, because it only works when you can actually draw from each conditional. ## Where the conditionals come from The usual source is conditional conjugacy. In a normal linear model with a normal prior on the coefficients and an inverse-gamma prior on the error variance, the coefficients given the variance are normal, and the variance given the coefficients is inverse-gamma. Neither the joint posterior nor its normaliser has a friendly closed form, but each conditional slice does, and Gibbs only ever needs the slices. Hierarchical models are the same story: group-level parameters given the hyperparameters are standard, hyperparameters given the group-level parameters are standard, and the sweep alternates. When one block resists — a non-conjugate prior, a link function that breaks the algebra — the standard remedy is Metropolis-within-Gibbs: update the well-behaved blocks by direct draws and the awkward one with a Metropolis step, keeping the overall sweep valid. You then inherit a proposal scale to tune, but only for that block. The scan order is flexible. Sweeping the blocks in a fixed sequence is the common choice; picking a block at random each iteration is also valid. What is not valid is conditioning on stale values, or updating a block using a conditional derived from a different model than the one you are fitting. ## The cost: axis-aligned moves Gibbs pays for its lack of rejection with the shape of its moves. Each update changes one block while the others are pinned, so in a two-parameter picture the chain moves only horizontally and vertically. When parameters are close to independent, that is fine. When they are strongly dependent, it is crippling. Take a bivariate normal posterior with correlation 0.99. Conditional on the current value of the first coordinate, the second is confined to a very narrow band around the ridge, so the vertical move is tiny; the following horizontal move is tiny for the same reason. The chain zig-zags in small steps along a long diagonal ridge and needs an enormous number of sweeps to travel from one end to the other. Every proposal is accepted, so a naive acceptance-rate check looks perfect, while the chain is barely exploring. This is the standard cautionary example: acceptance rate is not a measure of progress for Gibbs. The usual mitigations are structural rather than tuning: reparameterise so the coordinates are closer to independent (centring or non-centring a hierarchical model, orthogonalising correlated predictors), or block the correlated parameters together and draw them jointly from their multivariate conditional, which turns the zig-zag into a single move along the ridge. ## What an interviewer is listening for Strong answers name the full conditional precisely — given all other parameters and the data — explain that acceptance is one because the proposal is the exact conditional, state the requirement that each conditional be directly samplable, and volunteer the correlation weakness without being prompted. Weak answers describe Gibbs as drawing all parameters jointly, or claim that never rejecting makes it strictly better than Metropolis-Hastings.

  • What do you do when one full conditional is not a standard distribution you can sample?
    Use Metropolis-within-Gibbs: keep direct draws for the blocks whose conditionals are standard and update the awkward block with a Metropolis step targeting its conditional. The sweep stays valid, and you only take on a proposal scale to tune for that one block rather than for the whole parameter vector.
  • Why does a Gibbs sampler crawl on a bivariate normal posterior with correlation 0.99?
    Updates are axis-aligned. Conditional on the current first coordinate, the second is confined to a narrow band around a long diagonal ridge, so each move is tiny relative to the ridge's length, and the chain zig-zags slowly along it. Acceptance is still 100 percent, which makes the problem invisible to acceptance-rate checks. Blocking the pair or reparameterising fixes it.
  • Must the blocks be updated in a fixed order?
    No. A deterministic scan through the blocks is the usual choice and is perfectly valid; choosing a block at random each iteration is also valid. What matters is that each update conditions on the current values of the other blocks, including any updated earlier in the same sweep.

saying these in an interview costs you the question

  • Says Gibbs draws all parameters jointly each iteration
  • Conditions on the previous sweep's values instead of the updated ones
  • Claims Gibbs works for any posterior regardless of the conditionals
  • Treats 100 percent acceptance as proof the chain mixes well
  • Thinks Gibbs is always preferable because nothing is rejected

context