skip to content

Your posterior for a parameter is bimodal — what goes wrong if you report only the posterior mean?

level: seniorimportance: should knowfreq 48%

answer

  1. look at the shape before summarising
  2. the mean is a balance point
  3. between two peaks lies a valley
  4. the interval spans the trough too
  5. report mass per mode, not a location

basics

~10 s

With two separated modes the posterior mean lands in the trough between them, a value the posterior itself calls unlikely. The single number also hides that two competing explanations are in play.

solid answer

~50 s

A bimodal posterior says the data plus prior support two distinct explanations with a low-density valley between them. The mean is a balance point, so it falls in that valley: you report a value the posterior actively disfavours, and an equal-tailed interval around it spans the trough as if it were solid ground. Report the shape instead — the density itself, the posterior mass in each mode, and the decision-relevant probability such as the chance the parameter exceeds your action threshold. If one number is genuinely required, choose it from the loss and say so: the mean is optimal under squared-error loss even when it is an implausible parameter value. Then diagnose the bimodality — weak identifiability, two regimes in the data, or prior-data conflict — because it is usually a finding, not a nuisance.

go deeper

for a junior

Remember that a posterior is a whole distribution and that the mean is only one summary of it. Be able to say that with two peaks the mean falls in between, where the posterior is low.

for a middle

Explain why the mean is a mass-weighted balance point and why the median can also land in the gap. Know that a highest-density region may consist of two disjoint intervals while an equal-tailed interval spans the trough.

for a senior

Show the working habit: inspect the shape, report mass per mode and the decision-relevant probability, and pick any point estimate from the loss function while naming it. Then diagnose the cause, from weak identifiability to a genuine subgroup.

for a principal

Own how uncertainty is allowed to travel through the organisation: which decisions may consume a point estimate at all, what must accompany one, and how to stop dashboards from silently flattening multi-modal or skewed posteriors into a single misleading number.

## What bimodality is telling you A posterior with two separated peaks is not noise; it is the analysis saying that two different parameter regions each account for the data well, and that the region between them accounts for it poorly. That is substantive information. Collapsing it to one number throws away exactly the part that was informative. ## Why the mean lands in the wrong place The posterior mean is a mass-weighted balance point. With mass piled at two separated locations, the balance point sits between them — in the valley. If the two modes are near-equally weighted, the mean sits close to the middle of the trough, which may be a region of very low posterior density. Nothing in the number itself signals this. A reader who receives only the mean will reasonably assume the posterior is a single hump centred there, and will act on a value the analysis considers among the less plausible. The median has a related problem: it is the value with half the mass on either side, which for two separated humps can again fall in the gap. Both summaries answer a question about location, and location is not what a two-humped posterior is about. ## The interval makes it worse, not better The standard reflex is to attach an interval. An equal-tailed interval — trim a fixed share of mass from each end — will typically stretch from inside the left mode to inside the right mode, covering the trough as though it were solid support. It reports the right total mass while implying the parameter is plausibly anywhere in between, which is precisely the opposite of what the posterior says. A highest-density region is honest here, because it can come out as a union of two disjoint intervals with a hole in the middle; the hole is the point. ## What to report instead Order of preference: 1. **The posterior itself.** A plot of the density is the fastest way to convey two competing explanations and costs one figure. 2. **The mass in each mode.** Something like: about two thirds of posterior mass sits near the lower region, one third near the upper. Two numbers, and the reader knows the whole story. 3. **The decision-relevant probability.** Most consumers of the analysis do not need the parameter, they need an action. The probability that the parameter exceeds a threshold, or that one option beats another, is a single number that stays meaningful no matter how strange the shape. 4. **A point estimate chosen from the loss, and labelled as such.** If a downstream system requires one number, derive it: the mean minimises expected squared error, the median minimises expected absolute error, and different losses can select different values. Reporting a mean and saying it is the squared-error-optimal action is defensible. Reporting a mean as if it were the plausible value of the parameter is not. ## Diagnose before you summarise Separated modes usually have a cause worth naming: - **Weak identifiability.** Two different parameter regions imply nearly the same observable predictions, so the data cannot separate them. More of the same data will not help; a different measurement or design might. - **Two regimes in the data.** A subgroup, a period, or a mixture whose components pull toward different parameter values. Modelling the structure explicitly usually replaces one confusing posterior with two clear ones. - **Symmetry in the model.** Some parameterisations admit an exchange that leaves the fit unchanged, producing mirror-image modes that are two labels for the same explanation. Here the fix is a constraint or a relabelling, not a summary. - **Prior-data conflict.** A prior concentrated somewhere the likelihood is not can leave mass in two separated places. That is a modelling conversation, not a reporting one. ## The senior habit The generalisable lesson goes beyond bimodality: always look at the posterior before choosing how to summarise it. Skew, heavy tails, hard boundaries at zero and modes sitting on a constraint all break the implicit assumption that a mean plus an interval is a faithful summary. The mean is a good default only when the posterior is roughly single-humped and symmetric — and you only know that by looking. Reporting summaries computed by habit and never inspecting the shape is a failure interviewers probe for deliberately, because it is the one that survives into production dashboards where nobody sees the density.

  • Is the posterior mean ever the right answer when the posterior is bimodal?
    Yes, when the decision genuinely has squared-error loss and demands a single number — then the mean is the optimal action even though it is an implausible parameter value. The requirement is to say so explicitly: report it as the loss-optimal action, alongside the posterior shape, never as the value the parameter probably takes.
  • Why is a highest-density region more honest than an equal-tailed interval here?
    A highest-density region collects the most probable values, so with two separated modes it can come out as two disjoint intervals with a gap. That gap communicates the trough. An equal-tailed interval trims fixed mass from each end and therefore spans the trough, implying support the posterior does not give.
  • What are the common causes of a bimodal posterior worth checking?
    Weak identifiability, where two parameter regions imply nearly the same predictions; genuine subgroups or regimes that a single-parameter model is straining to cover; symmetries in the parameterisation that create mirror modes; and prior-data conflict, where prior and likelihood concentrate in different places. Each implies a different fix.

Asking for the average of two crowded railway platforms on either side of the tracks gives you a point on the tracks — arithmetically correct, and the one place nobody is standing.

saying these in an interview costs you the question

  • Computes summaries without ever plotting the posterior
  • Assumes any posterior is approximately a single symmetric hump
  • Treats the posterior mean as the most probable parameter value
  • Quotes an equal-tailed interval spanning a low-density trough as solid support
  • Calls bimodality a nuisance to smooth away rather than a finding

context