skip to content

Why does an MCMC run with 10,000 stored draws report an effective sample size of only 200?

level: middleimportance: must knowfreq 64%

answer

  1. the draws are not independent
  2. consecutive states are nearly identical
  3. count information, not iterations
  4. the sum runs over every lag
  5. precision scales with its square root

basics

~20 s

Because MCMC draws are dependent. When each draw is nearly a repeat of the previous one, 10,000 correlated draws carry about as much information as 200 independent ones, and effective sample size reports that count rather than the stored count.

solid answer

~50 s

MCMC produces a dependent sequence, so the stored draw count overstates how much information you have. Effective sample size answers "how many independent draws would give a posterior mean this precise?" and is approximately `ESS = N / (1 + 2 * sum of the autocorrelations at lags 1, 2, 3, ...)`. With a lag-1 autocorrelation around 0.98 the chain barely moves between iterations, the sum in the denominator is large, and 10,000 draws collapse to a few hundred effective ones. The consequence is precision: the Monte Carlo standard error of a posterior mean scales with the posterior standard deviation divided by `sqrt(ESS)`, so an ESS of 200 gives roughly a seventh of the precision the raw 10,000 suggests. Fixes attack the mixing, not the bookkeeping — reparameterise, use a sampler that takes longer informed moves, or run more chains. Tail effective sample size needs its own check.

code

python · 15 lines
python
import random, statistics

rho, n = 0.98, 10000
x, chain = 0.0, []
for _ in range(n):
    x = rho * x + random.gauss(0, (1 - rho ** 2) ** 0.5)
    chain.append(x)

m = statistics.fmean(chain)
num = sum((chain[i] - m) * (chain[i + 1] - m) for i in range(n - 1))
den = sum((v - m) ** 2 for v in chain)
r1 = num / den

print("lag-1 autocorrelation:", round(r1, 3))
print("effective draws:", round(n * (1 - r1) / (1 + r1)))

go deeper

for a junior

Know that MCMC draws are dependent and that effective sample size counts how many independent draws they are worth. Recognise that a big gap between stored and effective draws means the chain is mixing slowly.

for a middle

Explain the formula with the autocorrelation sum over all lags, why positive correlation shrinks the count, and that precision of a posterior mean scales with the square root of the effective count rather than the stored count.

for a senior

Show what you actually do about it: reparameterise or change sampler before adding iterations, check tail effective sample size when the deliverable is an interval, and refuse to quote digits the effective count cannot support.

for a principal

Own the reporting standard — the minimum effective sample size a model must clear before its numbers are published, whether tail figures gate interval claims, and how much parallel compute the organisation spends chasing that bar.

## Dependent draws are worth less Ordinary Monte Carlo reasoning assumes independent samples. MCMC gives you a *chain*: each state is generated from the previous one, so consecutive draws resemble each other. If the sampler takes small steps, draw 4,001 is nearly a copy of draw 4,000, and the pair together tells you barely more than either alone. The stored draw count is therefore a bookkeeping number, not an information count. **Effective sample size (ESS)** converts the stored count into the honest one. Its definition is operational: the number of *independent* draws that would give an estimate with the same variance as the one your correlated chain produced. ## The formula For a chain of `N` draws with autocorrelation `rho_k` at lag `k`: ``` ESS = N / (1 + 2 * (rho_1 + rho_2 + rho_3 + ...)) ``` Every lag contributes. Positive autocorrelation makes the denominator larger than 1 and pushes ESS below `N`. Two consequences worth stating in an interview: - The relevant quantity is the *sum over all lags*, not the lag-1 value alone. Two chains can share a lag-1 correlation of 0.98 and have quite different ESS if one's correlation decays quickly after lag 1 and the other's decays slowly. - Negative autocorrelation is possible — some samplers deliberately produce antithetic behaviour — and then ESS can exceed `N`. It is not a contradiction; it means the chain over-corrects and averages more efficiently than independent draws would. For the idealised case where correlations decay geometrically as `rho_k = rho^k`, the sum has a closed form and `ESS = N * (1 - rho) / (1 + rho)`. At `rho = 0.98` that is about one percent of `N` — 10,000 draws worth roughly a hundred. Real chains are messier, which is why a reported ESS of 200 from 10,000 draws is entirely plausible. ## Why it matters: precision The reason to care is the Monte Carlo error of your reported numbers. The standard error of a posterior mean estimated from a chain behaves like ``` MCSE = posterior_sd / sqrt(ESS) ``` with `ESS`, not `N`, in the denominator. So the ESS of 200 means your posterior mean is roughly `sqrt(10000 / 200)` — about seven times — less precise than the stored count would suggest. If you report a posterior mean to three significant figures on the strength of 10,000 draws while the effective count is 200, most of those figures are noise you have dressed up as inference. A useful sanity rule: you want ESS in the low hundreds at minimum for stable means and intervals, and more for anything in the tails. ## Bulk versus tail Modern reporting splits ESS in two. **Bulk-ESS** governs how well the centre of the distribution is estimated — means, medians. **Tail-ESS** governs the extreme quantiles used for interval endpoints. Tail-ESS is usually the smaller of the two, because the chain visits the tails rarely and those visits are strongly clustered. A fit with an acceptable bulk-ESS can still have a 95% interval whose endpoints wobble noticeably between reruns. If your deliverable is an interval, tail-ESS is the number that gates it. ## What actually fixes a low ESS - **Reparameterise.** Correlated or badly scaled parameters force the sampler to take tiny steps. Recentring, rescaling, or restating a hierarchy in a form with flatter geometry often buys an order of magnitude. - **Use a sampler that takes longer informed moves.** Gradient-informed trajectories traverse the distribution in one transition where a small random-walk step needs hundreds. - **Run longer.** ESS grows roughly linearly in iterations for a fixed chain, so this always works and is always the expensive answer. - **Run more chains.** ESS is additive across chains, and chains run in parallel, so wall-clock cost is often flat. ## What does not fix it **Thinning does not raise ESS.** Keeping every tenth draw removes nine tenths of the draws and roughly a corresponding share of the effective ones; the ratio ESS/stored-draws improves, but the ESS itself does not, and it usually falls slightly. Thinning changes the bookkeeping, not the information. Likewise, discarding more warm-up removes start-up bias but leaves the correlation structure of the retained draws untouched. ## Reading it next to convergence ESS and a between-chain convergence statistic answer different questions. A convergence statistic asks *are the chains describing the same distribution*; ESS asks *how much information do the draws carry*. A chain can be beautifully converged and nearly useless, or highly informative within a region it should never have been trapped in. Both are checked, per parameter, before any number is reported.

  • Can effective sample size ever exceed the number of stored draws?
    Yes, when the autocorrelations are negative. Some samplers move in an antithetic way, so a high draw tends to be followed by a low one and the errors partially cancel. The autocorrelation sum in the denominator then falls below 1 and the effective count rises above the stored count. It is a sign of an unusually efficient sampler, not a computation error.
  • Why report bulk and tail effective sample size separately?
    They govern different summaries. Bulk-ESS controls the precision of central quantities like the posterior mean, while tail-ESS controls the extreme quantiles that form interval endpoints. Tail-ESS is normally the smaller, because the chain visits the tails rarely and in clusters, so a fit with a healthy bulk figure can still produce interval endpoints that move between reruns.
  • What raises effective sample size for a fixed compute budget?
    Improving the geometry rather than adding iterations: reparameterising to decorrelate or rescale parameters, and using a sampler whose transitions travel further per iteration. Running more chains in parallel also helps, because effective sample size adds across chains while wall-clock time stays roughly flat. Simply running one chain longer works but buys precision at linear cost.

Polling the same person ten times gives ten answers but one opinion; a sticky chain is doing the same thing with parameter values.

saying these in an interview costs you the question

  • Says thinning increases the effective sample size
  • Reports precision using the stored draw count
  • Thinks lag-1 autocorrelation alone determines it
  • Assumes a good convergence statistic implies a high effective sample size
  • Ignores tail effective sample size when reporting intervals

context