skip to content

Why are the warm-up (burn-in) draws of an MCMC chain discarded before summarising the posterior?

level: juniorimportance: should knowfreq 58%

answer

  1. the chain has to get somewhere first
  2. early states remember the starting value
  3. not yet in the typical set
  4. the sampler is still tuning itself
  5. removes bias, costs precision

basics

~20 s

Early MCMC draws still reflect where the chain was started rather than the posterior, and in adaptive samplers the tuning is still changing. Those draws are dropped so the summary uses only draws from the stationary distribution.

solid answer

~50 s

An MCMC sampler is a Markov chain whose *stationary* distribution is the posterior, but it only samples from that distribution once it has reached and settled into the region of high posterior mass — the typical set. The first iterations are dominated by the arbitrary starting value, so averaging them in pulls estimates toward wherever you happened to initialise. In adaptive samplers there is a second reason: during warm-up the step size and metric are still being tuned, so the transition kernel is changing from iteration to iteration and does not leave the posterior invariant at all — those draws are not valid posterior samples under any argument. Discarding warm-up removes initialisation bias; it costs some Monte Carlo precision, because you keep fewer draws, which is why warm-up length is a judgement rather than "more is always better". A common default is to discard the first half of each run.

go deeper

for a junior

Be ready to say in one breath that early draws still reflect the starting value, not the posterior, and that they are dropped for that reason. Knowing the term warm-up and the half-the-run default is enough here.

for a middle

Explain both mechanisms: the chain has not reached the typical set, and during adaptation the transition kernel is still changing so those iterations are not valid samples at all. Say how you would choose the length using dispersed chains.

for a senior

Show the judgement that a chain still unsettled after a long warm-up is reporting a geometry problem, not asking for more iterations. Be explicit that warm-up trades precision for the removal of initialisation bias.

for a principal

Own the convention your team runs on: what the default warm-up fraction is, whether diagnostics are computed on retained draws only, and how much compute you are willing to spend on warm-up versus retained draws across a whole fleet of fitted models.

## The problem warm-up solves Markov chain Monte Carlo does not draw independent samples from the posterior. It constructs a chain of dependent states whose *stationary* (long-run) distribution is the posterior. The theory guarantees that if you run the chain long enough, the distribution of its state converges to the posterior no matter where you started. It says nothing reassuring about iteration 3. You have to start the chain somewhere — a random point, a crude initial guess, all parameters at zero. That point is almost always far out in the tails, in a region the posterior gives very little mass to. The chain then drifts from there toward the **typical set**: the region that actually holds nearly all posterior probability. Draws collected during that drift are samples of "where a chain that started at my arbitrary initial value is after k steps", not samples of the posterior. Averaging them into your posterior mean drags the estimate toward the initial value. That is *bias*, and unlike Monte Carlo noise it does not shrink just because you collect more draws afterwards — it shrinks only because you throw those early draws away. ## The second, sharper reason: adaptation Modern samplers spend the warm-up period tuning themselves: choosing a step size to hit a target acceptance rate, estimating a mass matrix or covariance for the proposal geometry. During that period the transition rule itself is changing from iteration to iteration. A chain with a changing kernel is not a homogeneous Markov chain targeting the posterior, so the usual invariance argument does not apply to those iterations at all. Even if the trace *looks* flat and well-behaved during adaptation, those draws are not posterior samples and must be discarded. This is why the modern name is "warm-up" rather than "burn-in": the phase is doing tuning work, not merely forgetting the start. ## How much to discard There is no formula. Practical guidance: - A common default is to discard the first half of every run, then use the second half for inference. It is wasteful but safe. - Judge it with the diagnostics you would run anyway: start several chains from *dispersed* initial values and check that, after the discarded portion, they overlap on the trace plot and their between-chain and within-chain variation agree. - Compute convergence statistics on the *post*-warm-up draws only. Computing them on the full run mixes the transient into the diagnostic and can either hide a problem or invent one. ## What warm-up does not do Warm-up is a fix for one specific failure: dependence on the starting point. It does not: - **Reduce autocorrelation.** Successive draws remain correlated after warm-up; that is a separate property, measured by effective sample size, and no amount of discarding at the front changes the correlation structure of what remains. - **Rescue a chain stuck in the wrong place.** If the chain has entered a region it cannot leave, throwing away the first half just gives you a longer run of the same stuck behaviour. The trace will look deceptively calm. - **Fix a misspecified or badly conditioned model.** A chain that will not settle after a long warm-up is usually telling you about the geometry of the posterior, not asking for more iterations. ## The cost side Discarding is not free. Every discarded draw is a draw you paid for and cannot use, so the remaining sample is smaller and the Monte Carlo error of your posterior summaries is larger. The trade is bias against precision: too little warm-up leaves initialisation bias in the estimate; too much leaves an honest but noisier estimate. Because bias is the more dangerous failure — it does not announce itself in an interval — the conventional choice errs toward discarding generously, and then buys the precision back by running the sampler longer rather than by keeping the transient. ## What an interviewer is listening for That you can say *why* the early draws are not posterior draws (start dependence, plus a changing kernel during adaptation), that you check the decision with dispersed chains rather than a rule of thumb, and that you do not oversell warm-up as a cure for mixing or for a chain trapped away from the bulk of the posterior.

  • How do you decide how many warm-up iterations are enough?
    Not by a fixed count. Start several chains from dispersed initial values, discard a generous prefix (half the run is a common default), and check that the retained draws from different chains overlap on the trace and agree in variance. If the retained draws still disagree, the honest read is usually a hard posterior geometry rather than a demand for a longer warm-up.
  • Does discarding too much warm-up bias the results?
    No — it costs precision, not accuracy. Every extra discarded draw is a valid posterior draw you are not using, so posterior summaries get noisier while staying centred correctly. That asymmetry is why the safe error is to discard too much and then run longer, rather than to keep a transient you cannot detect in the final interval.
  • Why can adaptation-phase draws not be kept even when the trace already looks stationary?
    Because the transition kernel is still changing while step size and metric are being tuned. The invariance argument that makes MCMC valid assumes a fixed kernel targeting the posterior, and a chain whose rule changes each iteration satisfies no such guarantee. A flat-looking trace during adaptation is not evidence that those draws came from the posterior.

It is like timing a runner from the moment they leave the changing room: the first stretch measures how far away they parked, not how fast they run.

saying these in an interview costs you the question

  • Claims warm-up removes autocorrelation between draws
  • Says discarding warm-up fixes a chain stuck in one region
  • Treats a fixed number like 1000 as universally correct
  • Keeps adaptation-phase draws because the trace looked flat
  • Computes convergence diagnostics on the full run including warm-up

context