skip to content

Would you standardise on equal-tailed or highest-density credible intervals across your team's readouts?

level: principalimportance: nice to knowfreq 18%

answer

  1. a governance call, not a maths one
  2. which rule two analysts reproduce identically
  3. what happens when the scale changes
  4. some posteriors break the single-interval assumption
  5. default plus written exception triggers

basics

~20 s

Make equal-tailed the default, because it is reproducible, survives rescaling and is always one interval. Require the highest-density version, plus the posterior plot, in named exception cases: strongly skewed posteriors, posteriors piled against a boundary, and multimodal ones.

solid answer

~50 s

I would set equal-tailed as the house default and treat highest-density as a documented exception. The default has three organisational virtues: it is two quantiles, so any two analysts computing it from the same draws agree; it is invariant under monotone rescaling, so a rate interval and a log-odds interval tell the same story; and it always returns one contiguous range, which dashboards and decision memos assume. The costs are real: on a skewed posterior it is longer than needed, and against a boundary it excludes zero by construction. So I would mandate the highest-density interval in three named cases: heavy skew, a posterior whose mass piles at a boundary, and multimodality, where the honest density region is two disjoint chunks that a single interval would bridge with implausible values. Whichever is used, the readout must name the rule, name the prior, and carry the posterior plot.

go deeper

for a junior

Know that a report should say which interval rule produced the numbers, since equal-tailed and highest-density endpoints differ on the same posterior.

for a middle

Be able to explain why one rule reproduces exactly from quantiles while the other is a search over draws, and why that difference matters for two analysts comparing results.

for a senior

Show that you would pick the rule from the posterior's shape - skew, boundary mass, extra modes - and attach the plot rather than trusting two numbers.

for a principal

Own the tradeoff between per-case optimality and organisation-wide comparability, and be ready to defend a written default with mechanical exception triggers rather than case-by-case taste.

## The decision Equal-tailed and highest posterior density (HPD) intervals both hold the stated share of the posterior; they differ in which slice they take. The equal-tailed one runs between the 2.5th and 97.5th posterior percentiles at the 95% level. The HPD one is the shortest set holding 95%, equivalently every value whose posterior density exceeds a threshold. As a personal analysis choice this is a small matter. As a **reporting standard across a team** it is not, because summaries get compared across experiments, scales and people, and comparability is the whole point of a standard. ## The case for equal-tailed as default **Reproducibility.** It is two sample quantiles. Two analysts with the same posterior draws get the same endpoints. The HPD estimated from draws is a narrowest-window search over the sorted sample, which is noisier in the tails and needs more draws before it settles - so two people can report visibly different endpoints from the same model. **Invariance under rescaling.** Percentiles map through any increasing transformation, so the equal-tailed interval for a rate transforms into the equal-tailed interval for the odds or the log-odds. HPD does not: a change of variable rescales the density by a Jacobian factor and reorders which values count as highest density. In an organisation where some teams report rates and others report log-odds or lifts, HPD makes two correct reports of the same posterior disagree at the endpoints for no substantive reason. **It is always one interval.** Dashboards, decision memos and downstream tooling assume a low and a high number. HPD does not guarantee that. ## The case for the exceptions **Heavy skew.** On a strongly asymmetric posterior the equal-tailed interval keeps a long thin tail while excluding denser values on the other side, so it is longer than necessary and its centre is misleading. When the report is about which values are most plausible, HPD is the honest summary. **Boundary mass.** When the posterior's density is highest at a boundary of the parameter space - say a rate after observing no events at all - the equal-tailed rule must discard 2.5% of the mass below its lower endpoint, so it reports a strictly positive lower bound and excludes zero purely as an artefact of the rule. The HPD region runs to the boundary and includes it. Reading the equal-tailed lower bound as evidence against zero is a real misinterpretation that shows up in safety and defect-rate reporting. **Multimodality.** With two separated peaks and a low-density valley between them, the 95% HPD region can be two disjoint chunks, one around each peak. The equal-tailed interval spans the valley and so reports a contiguous range that includes values the posterior considers relatively implausible. Here the standard should force the *posterior plot*, because no two-number summary is adequate; the disjoint region is a signal that the model or the data are telling you something structural. ## What the standard should mandate 1. **A named default** - equal-tailed - so the unmarked case is unambiguous. 2. **The rule printed with the numbers.** Any interval in a readout says which rule produced it; an unlabelled interval is not comparable to anything. 3. **The prior printed with the numbers,** since the probability statement is conditional on it. 4. **Exception triggers written down,** not left to taste: skew beyond some agreed threshold, mass within some distance of a boundary, or any detected second mode. 5. **The posterior plot attached** whenever an exception fires, and ideally always. The plot is the artefact that makes a bad summary self-correcting. 6. **One scale per metric family,** so the invariance question stops arising for the metrics people compare most often. ## Tradeoffs to acknowledge aloud Standardising costs something. The default is sometimes the worse summary, and an analyst who knows their posterior is skewed will occasionally be forced into an exception process for a case they could have judged alone. The alternative - each analyst choosing per report - produces intervals that cannot be compared and readers who cannot tell whether a difference between two readouts is substantive or a difference of convention. For a reporting layer read by non-specialists, comparability usually outweighs per-case optimality, and the exception list is what buys back the lost precision. An interviewer at this level is looking for you to name that tradeoff rather than to declare one rule universally correct. ## The answer that does not land "HPD, because it is shorter" treats the choice as an optimisation problem when it is a communication and governance problem. Shortest is a property of one report; reproducible, comparable and always-an-interval are properties of a hundred of them.

  • What exception triggers would you actually write into the standard?
    Three, all checkable from the posterior draws: skew beyond an agreed threshold, a substantial share of mass within a small distance of a parameter boundary, and any detected second mode. Each triggers the highest-density summary plus a mandatory posterior plot, so the exception is mechanical rather than a matter of taste.
  • How would you handle a readout where the density region comes out as two disjoint chunks?
    Do not collapse it into one interval. Report both chunks with their masses and lead with the posterior plot, because the disjointness is itself the finding - it usually means a mixture, an unmodelled subgroup or a badly specified model. A contiguous summary here would assert plausibility for values the posterior actively disfavours.
  • Why does the reporting scale matter for this choice?
    Equal-tailed endpoints transform correctly under any increasing rescaling, so a rate interval and a log-odds interval agree. Highest-density endpoints do not, because the density picks up a Jacobian factor under the change of variable. If teams report on different scales, the density rule makes two correct summaries of one posterior disagree.

saying these in an interview costs you the question

  • Declares the shortest interval universally correct
  • Ignores that a density region can be disjoint
  • Leaves the interval rule unlabelled in reports
  • Treats reparameterisation invariance as academic trivia
  • Sets a default with no written exception process

context