When is the posterior mode a poor summary of a skewed posterior distribution?
answer
- peak location versus where the mass is
- skew separates three summaries
- each summary minimises a different loss
- mode matches an all-or-nothing loss
- asymmetric costs point to a quantile
basics
~10 sA skewed posterior has three different point summaries. Under right skew the mode is the smallest of the three, and it optimises an all-or-nothing loss that matches almost no real decision.
solid answer
~50 sThe mode is only the location of the peak; it says nothing about where the probability mass sits. For a right-skewed posterior on a latency parameter with a Gamma shape of 3 and rate 1, the mode is 2, the median about 2.7 and the mean 3. Reporting 2 quietly picks the smallest of the three and hides the long right tail that drives operational risk. The clean way to choose is decision-theoretic: the posterior mean minimises expected squared-error loss, the median minimises expected absolute-error loss, and the mode is the limiting answer under an all-or-nothing loss that rewards only being exactly right. Almost no real decision has that loss. If under-provisioning costs far more than over-provisioning, the right report is an upper quantile, not any measure of centre. In high dimensions the mode can also be atypical, with almost no posterior mass around it.
go deeper
Recall that mode, median and mean coincide only when the distribution is symmetric, and that for a right-skewed distribution the mode is the smallest of the three and the mean the largest.
Be ready to compute the three summaries for a named skewed distribution and explain why the long tail pulls the mean away from the peak while the median lands between them.
Demonstrate the decision-theoretic reasoning out loud: name the loss, derive the summary, and say plainly when asymmetric costs mean the right report is a quantile rather than any measure of centre.
Own the reporting standard your organisation uses. Decide when a single number may be handed to a downstream consumer at all, and how the loss that number will be used against gets documented alongside it.
## What a point summary is being asked to do A posterior is a whole distribution. Compressing it to one number is a decision, and different compressions answer different questions. The three standard summaries are: - the **mode**, the value where the posterior density is highest — this is the MAP estimate; - the **median**, the value with half the posterior mass on each side; - the **mean**, the probability-weighted average of the parameter. For a symmetric unimodal posterior all three coincide, and the choice is free. Skew is what separates them, and it is exactly where the choice starts to matter. ## A concrete skewed case Suppose the posterior for a latency parameter has a Gamma shape with shape parameter 3 and rate 1. Then: - mode = `(shape - 1) / rate` = 2 - median is approximately 2.67 - mean = `shape / rate` = 3 The ordering `mode < median < mean` is the signature of right skew, and it is worth being able to state on sight: the long right tail drags the mean up, the mode stays back at the peak, and the median sits between them. Three summaries, three answers, one posterior. Anyone who reports a single number without saying which summary it is has left the reader unable to reconstruct what was claimed. ## The decision-theoretic account The principled way to pick is to name the loss function. If `theta` is the truth and `a` is the number you report, Bayesian decision theory says to choose the `a` that minimises the posterior expected loss. - Squared-error loss, `(theta - a)^2`, is minimised by the **posterior mean**. - Absolute-error loss, `|theta - a|`, is minimised by the **posterior median**. - An all-or-nothing loss that charges a fixed penalty unless `a` is essentially exactly `theta` is minimised, in the limiting sense, by the **posterior mode**. That last loss is the one almost nobody has. Reporting a latency parameter, a conversion rate or a coefficient is not a game where being close counts for nothing. So the mode's popularity comes from computational convenience — it is the answer optimisation gives you, without any integration — rather than from a decision anyone was actually making. ## Asymmetric costs Many real problems have costs that are not symmetric at all. If under-provisioning a service costs far more than over-provisioning it, then no measure of centre is the right report. The loss-minimising action sits deliberately off-centre — for a strongly asymmetric loss, at an upper quantile of the posterior. Answering *which point summary should I use* with *neither, because your loss is asymmetric and the answer is a quantile* is the response that separates a senior candidate from a competent one. ## Why the mode gets worse as dimensions grow In a one-dimensional problem the mode at least sits in a region of appreciable probability. In many dimensions this stops being true. Density and mass come apart: the highest-density point occupies a vanishingly small volume, while the bulk of the posterior mass lives in a shell some distance away from it. A parameter vector drawn from the posterior will typically look nothing like the mode — its norm will be larger, its components less extreme in aggregate. So the MAP vector can be an untypical draw from the very distribution it summarises, and predictions made by plugging it in can be systematically different from predictions averaged over the posterior. ## Other failure modes to name - **Sharpness is invisible.** Two posteriors with the same mode can have wildly different spreads. The mode alone cannot distinguish a decisive result from a shrug. - **Boundaries.** When the posterior peaks at the edge of the parameter space — a variance at zero, a probability at one — the mode sits on the boundary and reports a degenerate answer that the mass of the posterior contradicts. - **Density units.** The mode is the argmax of a density, and a density is not a probability; it depends on how the parameter was scaled. The mean and the median depend on the distribution's mass rather than its peak height, which is why they behave more stably. ## What to do instead Report the summary that matches the decision, and say which one it is. When the posterior is skewed or the audience will act on the number, hand over more than a point: the shape, or at least a second summary that reveals the skew. If a single number is genuinely required, derive it from the loss the number will be used against rather than from whichever quantity was easiest to compute.
- Which point summary minimises expected squared-error loss, and which minimises absolute-error loss?The posterior mean minimises expected squared-error loss and the posterior median minimises expected absolute-error loss. The mode corresponds only to an all-or-nothing loss that charges the same penalty for every miss, however small. Naming the loss first is what turns the choice of summary from taste into derivation.
- Under-provisioning a service costs ten times what over-provisioning costs. Which posterior summary do you report?None of the three centres. With strongly asymmetric costs the loss-minimising action is an upper quantile of the posterior, chosen so the residual probability of under-provisioning is worth its cost. Reporting the mode, median or mean here optimises a symmetric loss nobody is paying.
- Why can the posterior mode be atypical in a high-dimensional parameter space?Density and mass separate as dimension grows. The highest-density point occupies negligible volume, while most posterior mass sits in a shell away from it. So a draw from the posterior typically looks unlike the mode, and plugging the mode in can give predictions that differ systematically from posterior-averaged ones.
The mode is the tallest point of a mountain range; the mean is where the rock actually is. On a lopsided range those are not the same place, and only one of them tells you how much rock there is.
saying these in an interview costs you the question
- Treats mode, median and mean as interchangeable
- Says the mode is always smaller than the mean
- Picks a summary without naming a loss function
- Assumes the mode always sits in a high-mass region
- Reports a point estimate for a strongly asymmetric cost
- Says skew does not matter if the sample is large