skip to content

Your random-walk Metropolis sampler accepts 2% of proposals — what is wrong and how do you fix it?

level: seniorimportance: should knowfreq 55%

answer

  1. the step size is the suspect
  2. too big overshoots, too small crawls
  3. there is a sweet spot, not a maximum
  4. roughly a quarter accepted
  5. adapt in a window, then freeze

basics

~20 s

The proposal steps are far too large, so almost every candidate lands in a low-density region and is rejected and the chain sits still for long stretches. Shrink the proposal scale until acceptance rises to roughly a quarter.

solid answer

~50 s

A random-walk proposal is `theta' = theta + scale * noise`, and the scale sets a hard tradeoff. Too large and proposals overshoot the posterior bulk into near-zero density, giving an acceptance rate near 2 percent and a chain that repeats the same value for dozens of iterations. Too small and almost everything is accepted — 98 percent is the mirror-image failure — but each move is a crawl. Both extremes waste compute for the same reason: very little distance covered per iteration. For multi-parameter random-walk Metropolis the classical asymptotic optimum is near 0.234, so tune for roughly 20 to 30 percent. Fix it with a finite adaptation phase that scales the step towards that target, and use a proposal covariance shaped like the posterior rather than one scalar step for parameters on wildly different scales.

go deeper

for a junior

Be ready to name the proposal step size as the knob behind a very low acceptance rate, and to say that both extremely low and extremely high acceptance are bad.

for a middle

Explain the mechanism on both sides: oversized steps land in near-zero density and get rejected, undersized steps are accepted but cover no ground, and the useful range is roughly 20 to 30 percent for multi-parameter models.

for a senior

Show how you would actually operate it: a pilot run, a finite adaptation window that is then frozen, a proposal covariance estimated from the pilot when parameters are correlated or differently scaled, and awareness that acceptance rate alone does not certify the run.

for a principal

Own the engineering tradeoff: how much tuning effort a hand-rolled random-walk sampler deserves before you switch to a method that adapts its own geometry, and what that choice costs in compute, gradient availability and team familiarity.

## Reading the symptom A 2 percent acceptance rate in random-walk Metropolis has one dominant cause: the proposal distribution is much wider than the posterior. The candidate `theta' = theta + scale * noise` is being thrown far outside the region where likelihood times prior is appreciable, the density ratio is tiny, and the test rejects. Because a rejection re-records the current state, the output contains the same value repeated fifty times on average before anything changes. The chain is technically valid — it would eventually explore the posterior — but the compute is being spent on proposals that are thrown away. ## The mirror-image failure The instinct to push acceptance as high as possible is exactly wrong, and interviewers probe for it. Shrink the scale far enough and nearly every proposal is a hair away from the current point, the density ratio is close to one, and acceptance climbs towards 98 percent. Now nothing is rejected — and the chain still explores terribly, because every accepted move is microscopic. It takes an enormous number of iterations to cross the posterior, and consecutive draws are so similar that they carry almost no new information. So the two failure modes are: - **Scale too large:** acceptance near 2 percent, chain sticks at one point for long stretches. - **Scale too small:** acceptance near 98 percent, chain moves constantly but in useless increments. Both are the same disease measured differently: little effective distance travelled per unit of computation. The useful setting sits between them. ## The target acceptance rate The well-known asymptotic analysis of random-walk Metropolis on high-dimensional targets identifies an optimal acceptance rate of about 0.234, which is why 20 to 30 percent is the standard operating range for multi-parameter models. In a single dimension the corresponding optimum is closer to 0.44. These are guidelines derived under idealised assumptions, not laws; they are a good target to tune towards and a poor thing to defend to three decimal places. The right way to use them is as a cheap, always-available signal: an acceptance rate of 2 percent or 98 percent tells you something is wrong without any further diagnostics, while a rate in the twenties tells you the scale is not the limiting problem. ## How to tune The crude fix is manual: run a short pilot, look at the acceptance rate, multiply or divide the scale by a factor of two or three, repeat. It works and it is honest. The standard automatic fix is an adaptation phase. During a defined initial window, adjust the scale after each batch of iterations — increase it when acceptance is above target, decrease it when below — then freeze it and run the sampler you actually keep. The frozen-then-run structure matters: a chain whose proposal keeps changing based on its own history is no longer a plain Markov chain, and the usual guarantees do not directly apply. Two disciplined ways out are (a) adapt during a finite warm-up window whose draws are discarded, then hold the scale fixed, or (b) use a diminishing-adaptation scheme where the size of adjustments shrinks towards zero as the run proceeds. What you must not do is tune the scale using the whole run and then report those same draws as a clean sample. ## When a single scalar is the wrong knob One scale for all parameters assumes they all have similar posterior spread. Real models violate this constantly: an intercept on the scale of hundreds beside a coefficient on the scale of hundredths. One scalar is then simultaneously far too large for the tight directions, which drives rejections, and far too small for the wide ones, which drives crawling — and no single value fixes both. Tuning it will land you at some mediocre compromise with a passable acceptance rate and bad exploration. Two remedies: - **Rescale the parameters.** Standardise predictors and put the coefficients on comparable scales so one step size is meaningful. - **Use a proposal covariance.** Propose from a multivariate normal whose covariance is proportional to an estimate of the posterior covariance, obtained from a pilot run. This handles both differing scales and correlation between parameters, since the proposal is then stretched along the same directions the posterior is. Correlated parameters are the case where a scalar step is most obviously hopeless: the posterior is a narrow diagonal ridge, an isotropic proposal has to be small enough to fit the ridge's width, and it therefore takes forever to travel the ridge's length. ## What acceptance rate does not tell you An acceptance rate in the target range is necessary, not sufficient. It says the step size is sensible relative to local geometry; it says nothing about whether the chain has found all the posterior mass, whether it is stuck in one mode of a multimodal posterior, or whether the model is even identified. Treat it as the first, cheapest thing to check when a sampler behaves badly, and expect a follow-up question about what else you would look at before trusting the output.

  • What does a 98 percent acceptance rate tell you about the same sampler?
    That the proposal scale is far too small. Nearly every candidate is a hair from the current point, so the density ratio is near one and everything is accepted, but each move covers almost no ground. It is the mirror image of the 2 percent case and just as wasteful — high acceptance is not a health signal.
  • Why is a single scalar step size a poor choice when parameters differ in scale by orders of magnitude?
    One value cannot suit both: it is too large for the tightly constrained directions, which causes rejections, and too small for the wide ones, which causes crawling. Standardise the parameters so their posterior spreads are comparable, or propose from a multivariate normal whose covariance is estimated from a pilot run.
  • Can you keep adapting the proposal scale for the whole run?
    Not freely. A proposal that keeps changing based on the chain's own history breaks the plain Markov structure the guarantees rest on. The safe patterns are adapting during a finite warm-up window whose draws you discard, or using a scheme where the adjustments shrink towards zero as the run proceeds.

Exploring a dark room: shuffling a centimetre at a time never bumps into anything but takes all night, while leaping across the room mostly slams into walls and leaves you where you started.

saying these in an interview costs you the question

  • Says a higher acceptance rate is always better
  • Fixes 2 percent acceptance by running more iterations
  • Tunes the scale on the whole run and keeps those draws
  • Uses one scalar step for parameters on wildly different scales
  • Treats a healthy acceptance rate as proof the sampler is fine

context