A hierarchical model's sampler reports divergent transitions clustered where the group-level scale is near zero. What do you do?
answer
- the sampler is warning, not just complaining
- bias in one region, not extra noise
- curvature changes with the scale parameter
- a wide mouth and a thin neck
- sample standardised offsets instead
basics
~10 sTreat the fit as biased, not noisy. Divergences near a near-zero scale parameter signal the funnel geometry a centred hierarchical parameterisation creates; the fix is a non-centred parameterisation, then a smaller step size.
solid answer
~50 sDivergences are the sampler telling you its simulated trajectory blew up numerically because the posterior's curvature is extreme relative to the step size. Clustered near a near-zero group-level scale, that is Neal's funnel: in a centred parameterisation the group parameters are drawn around a common location with width `tau`, so the region they may occupy narrows to a needle as `tau` shrinks and stays wide when it is large. No single step size fits both. The consequence is bias, not noise — the neck is systematically under-visited, so `tau` is over-estimated. Do not ignore a handful of divergences and do not just run longer. Reparameterise non-centred: sample `z_j ~ Normal(0, 1)` and set `theta_j = mu + tau * z_j`, which straightens the geometry. If they persist, shrink the step size by raising the target acceptance rate, and check whether the scale prior leaves too much mass at zero.
go deeper
Recognise that a divergence warning means the sampler failed to explore part of the posterior and that the fit should not be reported as it stands. Know the term and that it is a hard warning, not a soft one.
Explain the mechanism: the numerical trajectory loses energy where curvature is extreme relative to the step size, and a hierarchy with a near-zero scale creates exactly that geometry. Be able to write the non-centred change of variables.
Diagnose and act: name the funnel, state that the consequence is bias with a predictable direction on the scale parameter, reparameterise first, then tune the step size, and compare posteriors before and after rather than just silencing the warning.
Own the standard: whether any divergence blocks a model from being reported, which parameterisation is the house default for hierarchies, and how you keep teams from tuning warnings away under deadline without recording what the posterior did.
## What a divergent transition is A gradient-based sampler proposes a move by numerically simulating a trajectory across the parameter space. The simulation uses discrete steps, and the discretisation is only accurate when the step size is small relative to the local curvature of the log density. Where the density's geometry turns sharply, a step of the usual size overshoots, the simulated trajectory gains energy that the exact dynamics would have conserved, and it flies off. The sampler detects this energy error, rejects the trajectory, and records a **divergent transition**. The important consequence is not the wasted iteration. It is that divergences are *not random*: they occur in the specific regions where the geometry is hard, so the sampler systematically fails to visit exactly those regions. That produces **bias** in the posterior summaries, and bias does not average away with more iterations. A run with divergences is not a noisy run; it is a run whose answer is wrong in a direction you can often predict. ## Why they cluster at small scale in a hierarchy Write a hierarchy in its natural, **centred** form: a common location `mu`, a group-level scale `tau`, and per-group parameters `theta_j` drawn around `mu` with spread `tau`. Now look at the joint posterior over `(tau, theta)`. - When `tau` is large, the `theta_j` are free to spread widely: the region of appreciable posterior density is broad. - When `tau` approaches zero, the `theta_j` are forced to be nearly identical to `mu`: the region collapses to a needle. Plot this and you get **Neal's funnel** — a wide mouth narrowing to a thin neck. The width of the typical set varies by orders of magnitude along the `tau` axis. A step size tuned for the mouth is catastrophically too large in the neck; a step size small enough for the neck would make exploring the mouth impossibly slow. The sampler picks one, and it diverges in the neck. The bias has a predictable sign. Because the neck is under-visited, small values of `tau` are under-represented, so the posterior for `tau` comes back too large and the groups look more different from each other than the data actually says. Any downstream statement about how much the groups resemble one another inherits that error. ## The fix: non-centred parameterisation Change variables so the difficult dependence disappears from the geometry. Instead of sampling `theta_j` directly with a width that depends on `tau`, sample standardised offsets and rebuild the quantity of interest: ``` z_j ~ Normal(0, 1) theta_j = mu + tau * z_j ``` The sampled parameters are now `mu`, `tau` and the `z_j`, and the `z_j` have a fixed unit scale regardless of what `tau` does. The posterior in these coordinates is far closer to isotropic, one step size works everywhere, and the neck stops being a wall. The model is unchanged — the implied distribution over `theta_j` is identical — only the coordinates the sampler works in have changed. This is the single highest-value move in the toolkit for hierarchical models with modest data per group. Note the direction is not universal: with *lots* of data per group the likelihood dominates and the centred form can be the better-conditioned one. Non-centred is the right default for weakly informed groups. ## The rest of the toolkit If divergences persist after reparameterising: 1. **Shrink the step size** by raising the sampler's target acceptance rate. This is a genuine remedy for mild curvature and a bad one for a real funnel: it costs compute proportionally and, if the geometry is pathological, only reduces the divergence count without removing the bias. If raising the target makes divergences vanish *and* the posterior for the scale parameter shifts materially, that shift is evidence the earlier run was biased. 2. **Constrain the scale parameter's prior.** A prior placing substantial mass near zero, with too little data to contradict it, keeps the neck alive. A weakly informative prior that discourages implausibly small scales removes the region the sampler cannot handle — a modelling decision to be stated openly, not slipped in. 3. **Reconsider the hierarchy.** With very few groups the scale is barely identified, and no amount of sampler tuning creates information that is not there. ## What not to do - **Do not ignore a small number of divergences.** Even a handful concentrated in one region indicates systematic under-exploration there. "Only 12 out of 4,000" is not a defence. - **Do not just run longer.** More iterations of a sampler that cannot enter the neck give a more precise estimate of the wrong answer. - **Do not raise the target acceptance rate until the warnings stop and then report the fit without comparing.** Suppressing the warning is not the same as fixing the geometry; compare the scale parameter's posterior before and after and say what changed. - **Do not report only the parameters that look fine.** The bias propagates to anything computed from the affected parameters. ## How to present the diagnosis The senior answer names the mechanism (numerical integration error where curvature is extreme), names the geometry (a funnel created by the centred parameterisation), states the consequence in the right terms (bias, systematically over-estimating the scale, not extra variance), and gives the fix in order: reparameterise first, tune second, revisit the prior and the model third.
- Why are divergences a bias problem rather than a precision problem?Because they are not scattered at random. They occur exactly where the geometry is hard, so the sampler fails to visit that specific region and every summary is computed from a truncated exploration. Extra iterations give a tighter estimate of the same wrong answer. In a hierarchy the direction is usually predictable: the narrow region is under-visited, so the group-level scale comes back too large.
- When is the centred parameterisation the better choice?When each group has plenty of data. Then the likelihood pins the group parameters firmly and the posterior is well conditioned in centred coordinates, while the non-centred form can introduce its own awkward dependence. Non-centred is the default for weakly informed groups with few observations each; the two are alternative coordinate systems for the same model, and which conditions better depends on how much the data constrains each group.
- Is raising the target acceptance rate until the divergence warnings disappear an acceptable fix?Only if you check what changed. A smaller step size genuinely helps mild curvature, but against a real funnel it can silence the warning while leaving the region under-explored. Compare the posterior for the scale parameter before and after: a material shift is evidence the earlier run was biased and that the geometry, not the step size, is the actual problem.
- How many divergences are acceptable in a reported fit?The working answer is none. Divergences cluster in one region by construction, so even a handful signals that the sampler could not enter part of the posterior. There is no threshold below which the bias is known to be negligible, which is why the standard practice is to fix the geometry and rerun rather than to argue that the count was small.
It is like driving one fixed gear through both a motorway and a hairpin: the gear that works on the straight throws you off the bend, so the bends never get driven.
saying these in an interview costs you the question
- Dismisses a small number of divergences as noise
- Says running more iterations will average the problem out
- Silences the warning by tuning without comparing posteriors
- Treats divergences as extra variance rather than bias
- Cannot state the non-centred reparameterisation