Does thinning an MCMC chain to every 10th draw improve the posterior estimates?
answer
- correlated is not the same as worthless
- you paid for the discarded draws
- the ratio improves, the count does not
- a resource decision, not a statistical one
- only warm-up removal fixes bias
basics
~20 sNo. For a fixed number of iterations, keeping every tenth draw throws away information, so estimates are no better and usually slightly noisier than using all draws. Thinning is justified by storage or downstream cost, not accuracy.
solid answer
~50 sThinning makes the retained draws look more independent, which is cosmetically appealing, but the estimator that uses all draws is at least as precise as the one built from every tenth. The discarded draws were correlated with their neighbours, not worthless — a correlated draw still carries some information about the posterior, and dropping it can only remove information. The legitimate reasons to thin are practical: the draws will not fit in memory or storage, or each retained draw feeds an expensive downstream step such as a simulation or a per-draw prediction, and you would rather pay for 1,000 of those than 10,000. What thinning never does is fix a problem. It does not raise effective sample size, it does not remove initialisation bias — that is what discarding warm-up is for — and it does not turn a chain that has not converged into one that has. Between the two, only discarding warm-up actually improves the estimate.
go deeper
Know that thinning means keeping only every kth draw and that it does not make results more accurate. Do not present it as a required step in a Bayesian workflow.
Explain why the all-draws estimator is at least as precise for a fixed run: correlated draws still carry information, so discarding them can only lose some. Distinguish thinning clearly from discarding warm-up.
Show the resource reasoning — storage limits, expensive per-draw downstream work — as the only real justification, and say what you do instead when the chain is sticky: reparameterise, change sampler, or add parallel chains.
Own the storage-versus-precision policy across many fitted models: what gets archived at full resolution, what gets thinned on the way to a store, and how you keep a thinned archive from being read as if it were the full sample.
## The intuition that motivates thinning, and why it misleads MCMC draws are autocorrelated. Keeping only every tenth one visibly reduces that correlation, and the retained sequence starts to resemble the independent sample that textbook Monte Carlo formulas assume. It feels like a cleanup step. The mistake is assuming that "looks more independent" means "is more informative". It does not. Consider what you are doing: you generated 10,000 draws, paid for all of them, and are now choosing to compute your posterior mean from 1,000 of them. The 9,000 you discarded were correlated with their neighbours, but correlated is not the same as redundant — each still carries some information about where posterior mass lies. Averaging over all 10,000 uses that information; averaging over 1,000 does not. The standard result is that, for a fixed number of iterations, the estimator using every draw has variance no larger than the thinned estimator, and usually smaller. Thinning is an information-discarding operation dressed as a variance-reduction one. ## What actually changes when you thin - **Stored draws**: down by the thinning factor, by construction. - **Effective sample size**: down as well, though by less than the thinning factor. The *ratio* of effective to stored draws improves, which is exactly why thinning looks good on a diagnostic dashboard while making the estimate slightly worse. - **Monte Carlo error of the posterior mean**: unchanged at best, larger in practice. - **Bias from a chain that has not converged**: unchanged. Every tenth draw from a stuck chain is still from a stuck chain. That third and fourth line are the ones interviewers probe. Thinning is a bookkeeping change, not a statistical fix. ## When thinning is genuinely the right call Thinning is a *resource* decision, and there are real cases: - **Memory and storage.** A model with tens of thousands of parameters times tens of thousands of iterations produces a draw matrix that will not fit anywhere convenient. Storing a tenth of it is a reasonable trade of a little precision for feasibility. - **Expensive per-draw downstream work.** If each retained draw drives a costly simulation, a per-draw forecast over a long horizon, or a human-reviewed artefact, the cost is per *retained* draw. Thinning buys the compute budget directly. - **Transport and reproducibility.** Shipping a draw archive to collaborators or storing it long-term has its own economics. In each case you are trading a small, quantifiable loss of precision for a resource you actually lack. That is a defensible engineering trade. "It makes the autocorrelation plot look nicer" is not. ## Thinning versus discarding warm-up These two are constantly confused because both discard draws, but they address opposite problems: - **Discarding warm-up** removes draws that came from the *wrong distribution* — states dominated by the starting value, or produced while the sampler was still tuning itself and therefore not targeting the posterior at all. Keeping them biases the estimate, and no amount of extra sampling later removes that bias. Discarding them genuinely improves the estimate. - **Thinning** removes draws that came from the *right distribution* and were merely correlated with their neighbours. Keeping them helps; removing them costs. If an interviewer asks which of the two actually improves your posterior estimate, the answer is: only the warm-up discard. Thinning at best breaks even. ## How to answer the question well Say no, explain that correlated draws are still informative and that the all-draws estimator dominates for a fixed run length, then name the resource cases where you would thin anyway. Add that thinning does not raise effective sample size and does not repair a convergence failure. Finish with what you would do instead when the chain is sticky: reparameterise the model, use a sampler that takes longer informed moves, or run more chains in parallel — all of which raise the information in the sample rather than rearranging it.
- If thinning does not help accuracy, why is it still widely used?Because of resources rather than statistics. Large parameter vectors times long runs produce draw archives that will not fit in memory or ship easily, and each retained draw may drive an expensive downstream simulation or forecast. Paying for a tenth of that work in exchange for a small precision loss is a legitimate engineering trade. Doing it to tidy an autocorrelation plot is not.
- Between thinning and discarding warm-up, which one actually improves the estimate?Discarding warm-up. Those draws came from the wrong distribution — dominated by the starting value, or produced while the sampler was still tuning and not targeting the posterior — so keeping them biases every summary. Thinned draws come from the right distribution and are merely correlated; discarding them removes real information and can only cost precision.
- Does thinning make the diagnostics look better?Some of them, deceptively. The ratio of effective to stored draws rises and the autocorrelation at lag 1 of the retained series drops, so a dashboard reads cleaner. The effective sample size itself falls, and any convergence failure survives untouched, so the improvement is presentational. Judge a thinned chain on the same absolute numbers you would demand from an unthinned one.
saying these in an interview costs you the question
- Says thinning reduces Monte Carlo error
- Claims thinning is required for valid inference
- Thinks thinning raises the effective sample size
- Uses thinning to rescue a chain that has not converged
- Confuses thinning with discarding warm-up draws