What does the Gelman-Rubin R-hat statistic compare across MCMC chains?
answer
- you need more than one chain
- two kinds of variance are compared
- between chains versus inside each chain
- a ratio that falls toward one
- split each chain to catch slow drift
basics
~20 sR-hat compares the variation between several MCMC chains with the variation within each chain. If the chains have converged to the same distribution the two agree and R-hat sits near 1; values above roughly 1.01 mean the chains still disagree.
solid answer
~50 sR-hat, the Gelman-Rubin potential scale reduction factor, is computed per parameter from several chains started at dispersed initial values. It contrasts the between-chain variance `B` — how far the chains' means sit from each other — with the mean within-chain variance `W`. Concretely it is `sqrt(V_hat / W)`, where `V_hat` is a variance estimate that mixes the two and overestimates the posterior variance while the chains are still spread out. If every chain is exploring the same distribution, pooling them adds nothing beyond what one chain already shows, so the ratio falls to 1. Modern practice flags anything above about 1.01 (older texts said 1.1) and uses *split* R-hat, which cuts each chain in half and treats the halves as separate chains so that a single chain drifting slowly cannot pass. R-hat near 1 is necessary, not sufficient: chains that all started near the same mode can agree with each other and still miss the posterior.
go deeper
Know that R-hat needs several chains and that you want it close to 1. Be able to say it contrasts how much the chains differ from each other with how much each one varies internally.
Explain the construction: within-chain variance under-estimates and the pooled estimate over-estimates while chains disagree, so their ratio falls to 1 as mixing improves. Know the modern threshold and why chains are split in half.
Demonstrate the diagnosis: dispersed inits, worst-parameter reporting, and the judgement that a large R-hat calls for reparameterisation rather than more iterations. Say plainly why a good R-hat is necessary but not sufficient.
Own the acceptance policy — how many chains, which threshold, which parameters gate a model from shipping, and what happens when a production refit fails it. Weigh the compute cost of many chains against the risk of publishing a fit nobody checked.
## What R-hat is for Running one MCMC chain gives you no external reference: whatever the chain does looks, from the inside, like the posterior. The Gelman-Rubin diagnostic buys a reference by running **several chains from deliberately dispersed initial values** and asking a simple question — do these independent explorations agree with one another? If they have all converged to the same target, then the spread of one chain's draws should already tell you about the spread of the pooled draws. If they are still in different places, pooling them inflates the apparent variance, because you have added the differences *between* chains on top of the variation *within* each one. ## The construction For one scalar parameter, with `m` chains of `n` retained draws each: - `W` = the average of the within-chain variances. This *under*-estimates the posterior variance while the chains are still confined to their own neighbourhoods, because no single chain has yet toured the whole distribution. - `B` = the variance of the chain means, scaled by `n`. This captures how far apart the chains sit. - `V_hat = ((n - 1) / n) * W + (1 / n) * B`, a pooled variance estimate that *over*-estimates the posterior variance while the chains disagree. `R_hat = sqrt(V_hat / W)`. Read it as the factor by which your interval estimates might still shrink if you kept sampling — hence the original name, potential scale reduction factor. As the chains mix, `B` shrinks toward the within-chain scale, the numerator and denominator converge, and R-hat descends to 1 from above. ## Thresholds and the split refinement The historical rule of thumb was `R_hat < 1.1`. Contemporary guidance is much stricter — around **1.01** — because the older threshold lets through chains whose effective sample size is far too small for reliable tail summaries. Interviewers do not usually care which number you quote as long as you say the modern bar is tight and that the direction is "close to 1 from above". Plain R-hat has a hole: a single chain that drifts slowly in one direction has a large within-chain variance, which flatters the denominator. **Split R-hat** closes it by cutting each chain into halves and treating the halves as separate chains. A drifting chain then disagrees with itself, and the statistic rises. Rank-normalised variants further protect against heavy tails and infinite-variance parameters, where the raw variance ratio is not well behaved. ## The failure it is designed to catch The canonical picture: four chains launched from dispersed starting points, three of them settling into the same region and one parking somewhere else entirely and never leaving. The chain means are far apart, `B` is large relative to `W`, and R-hat comes back at something like 1.4. The correct read is not "run longer and it will average out" — it is that the sampler is not moving between regions of the posterior, so no summary computed from the pooled draws means anything. The remedies are structural: reparameterise, tighten the model, or accept that the posterior is genuinely multimodal and report it as such rather than pooling. ## Necessary, not sufficient R-hat near 1 says the chains agree. It does not say they are right. - If all chains are initialised in the same small region, they can agree with each other while jointly missing posterior mass elsewhere. Dispersed initialisation is part of the diagnostic, not an optional extra. - R-hat near 1 with a tiny effective sample size means the chains agree about a distribution they have each sampled only a handful of independent times. R-hat and effective sample size are checked together; neither substitutes for the other. - A sampler flagging pathologies during its trajectories can produce a comfortable R-hat while systematically under-visiting a hard region of parameter space. ## Practical protocol Run at least four chains from dispersed inits. Discard warm-up. Compute split R-hat per parameter — including derived quantities and the log posterior density, not just the parameters you care about, since a problem often shows up first in a scale parameter. Look at the *maximum* R-hat over parameters, not the average. Confirm with a trace plot: converged chains overlay each other as one fat band, while chains at R-hat 1.4 show visibly separated ribbons. If any parameter fails, treat the entire fit as unusable rather than reporting the parameters that happened to pass.
- Four chains from dispersed starts give R-hat 1.4 for one parameter, with one chain parked away from the others. What do you conclude?That the sampler is not moving between regions of the posterior, so the pooled draws summarise nothing coherent. Running longer rarely helps, because the chain that parked elsewhere has no mechanism to leave. The productive moves are structural: reparameterise the model, tighten the specification, or acknowledge genuine multimodality and report the modes separately instead of averaging across them.
- Why is split R-hat preferred to the plain version?Plain R-hat can be fooled by a single chain drifting steadily in one direction: the drift inflates that chain's own variance, which sits in the denominator and pushes the ratio down. Splitting each chain in half and treating the halves as separate chains forces a drifting chain to disagree with itself, so the statistic rises and the problem surfaces.
- Can R-hat be near 1 while the fit is still unusable?Yes. R-hat only certifies that the chains agree with each other. Chains all initialised in the same neighbourhood can agree while jointly missing posterior mass elsewhere, and chains can agree about a distribution each has sampled only a few independent times. That is why R-hat is always read alongside effective sample size and any sampler-reported pathologies.
- Which parameters do you check R-hat on?All of them, plus derived quantities and the log posterior density, and you look at the worst value rather than an average. Trouble frequently appears first in a scale or variance parameter that nobody planned to report. If any parameter fails, the whole fit is suspect, not just that coefficient.
saying these in an interview costs you the question
- Computes R-hat from a single chain
- Says any value below 1.1 is fine in modern practice
- Treats R-hat near 1 as proof the posterior is fully explored
- Reports the average R-hat instead of the worst
- Confuses R-hat with effective sample size