Why is a MAP estimate not invariant under reparameterisation while the MLE is?
answer
- a density must still integrate to one
- the likelihood is not a density
- change of variables brings a Jacobian
- a varying factor tilts the curve
- the median survives, the mode does not
basics
~20 sThe posterior is a probability density, so changing variables multiplies it by a Jacobian that reshapes the curve and can move its peak. The likelihood is not a density over the parameter, gets no Jacobian, so its maximiser transforms along.
solid answer
~50 sIf you transform a parameter one-to-one, densities do not just relabel — they pick up a Jacobian factor to keep total probability equal to one. That factor varies across the parameter range, so it tilts the posterior and the mode moves to a place that is not the transform of the old mode. The likelihood is a function of the parameter but not a density over it, so no Jacobian appears; if `theta_hat` maximises the likelihood then `g(theta_hat)` maximises it in the `g` parameterisation, for any one-to-one `g`. A concrete case: a `Beta(5, 2)` posterior for a probability has mode `4/5 = 0.8`, but writing the same posterior in terms of log-odds and taking its mode gives a value that maps back to `5/7`, about 0.71. Same beliefs, two different MAP answers. The posterior median is invariant under monotone transforms; the mean and the mode are not.
go deeper
Recall the headline fact: maximum likelihood transforms cleanly under a one-to-one reparameterisation, and the posterior mode does not. Knowing the direction of that asymmetry is enough at this level.
Be ready to state the change-of-variables rule with its Jacobian factor and explain why a varying multiplier can move a peak while leaving a non-density function's maximiser exactly where it was.
Show that you have hit this in practice: parameters optimised on an unconstrained scale and reported on the natural one. Be able to say which summaries survive the transform and choose one deliberately.
Own the convention. Decide which scale reported estimates are defined on across a team's models, and be able to defend reporting an invariant summary or the full posterior instead of a mode that depends on a symbol choice.
## Invariance, stated precisely An estimator is **invariant under reparameterisation** if reparameterising the model and re-estimating gives the transformed version of the original answer. Formally, for a one-to-one transform `g`, the estimator `T` is invariant when the estimate computed in the `phi = g(theta)` parameterisation equals `g(T)`, where `T` is the estimate computed in the `theta` parameterisation. This matters because the parameterisation is a modelling convenience. A variance or its logarithm, a probability or its log-odds, a rate or a mean waiting time — these describe the same beliefs. An estimate that changes when you rewrite the same model in different symbols has an arbitrary component. ## Why the MLE is invariant The likelihood `L(theta) = p(D | theta)` is a function *of* the parameter but not a probability distribution *over* the parameter. It has no normalisation requirement in `theta`, and it does not integrate to one over the parameter space. So under `phi = g(theta)` with `g` one-to-one, the likelihood in the new parameterisation is simply the old function composed with the inverse: `L_new(phi) = L(g_inverse(phi))`. Composing with a one-to-one map relabels the horizontal axis without changing any value, so the peak keeps its height and its identity. Whatever value of `theta` maximised the old function, its image `g(theta_hat)` maximises the new one. This is the well-known functional invariance of maximum likelihood: the MLE of a function of a parameter is that function of the MLE. ## Why the posterior mode is not The posterior *is* a density in the parameter — it integrates to one over the parameter space, and that constraint is what breaks the argument. Under `phi = g(theta)`, the change-of-variables rule for densities is ``` p_phi(phi) = p_theta(theta) * |d theta / d phi| ``` The extra factor, the absolute Jacobian, is what keeps total probability equal to one after the axis is stretched unevenly. But it is generally a function of the parameter, not a constant. Multiplying a curve by a varying positive function tilts it: regions where the transform stretches the axis are pushed up, regions where it compresses are pulled down. The peak of the tilted curve need not sit above the image of the old peak. So the mode is not invariant, and the MAP estimate inherits that. Notice this is a statement about the mode as a summary; the posterior itself is unchanged — it is the same beliefs, correctly re-expressed. Only the summary is parameterisation-dependent. ## A worked example Take a posterior on a probability `p` of the form `Beta(5, 2)`, whose density is proportional to `p^4 * (1 - p)^1`. Its mode is `(a - 1) / (a + b - 2) = 4 / 5 = 0.8`. Now reparameterise to log-odds, `phi = log(p / (1 - p))`. The derivative `d p / d phi` equals `p * (1 - p)`, so the density in `phi` is proportional to `p^4 * (1 - p) * p * (1 - p) = p^5 * (1 - p)^2`, expressed through `p`. Maximising `p^5 * (1 - p)^2` over `p` gives the stationary condition `5 / p = 2 / (1 - p)`, hence `p = 5 / 7`, about 0.714. Two MAP estimates from one posterior: 0.8 if you optimise on the probability scale, roughly 0.71 if you optimise on the log-odds scale and map back. Neither is wrong; the question *what is the most probable value* simply has no parameterisation-free answer for a continuous parameter. Under the same transform, an MLE would have produced the same answer on both scales. ## What survives the transform - The **posterior median** is invariant under any strictly increasing transform, because the transform preserves the ordering and therefore preserves which value has half the mass below it. - Posterior **probabilities of intervals** are invariant, since a region carries the same mass whatever coordinates describe it. - The **posterior mean** is not invariant, because averaging does not commute with a nonlinear function. - The **mode**, and hence MAP, is not invariant, for the Jacobian reason above. ## Practical consequences This is not a purely theoretical curiosity. Parameters that live on constrained scales are routinely optimised on an unconstrained one — a variance through its logarithm, a probability through its log-odds — because optimisation is easier there. If a mode is taken on the transformed scale and reported on the original, the number reported is a different estimator from the one the original-scale objective would have produced, and the discrepancy grows with the skew of the posterior. The defensible responses are: choose the parameterisation deliberately, on the scale the decision is actually made on; prefer a summary that is invariant, such as the median, when the scale is arbitrary; or report the posterior rather than a point and let the consumer take whatever summary their decision requires.
- Which posterior summaries do survive a monotone reparameterisation?The median does, because a strictly increasing transform preserves order and therefore preserves which value has half the mass below it. Posterior probabilities of intervals also survive, since a region carries the same mass in any coordinates. The mean does not, because averaging does not commute with a nonlinear function.
- Does this failure of invariance ever bite in practice?Yes. Constrained parameters are often optimised on an unconstrained scale, such as a variance through its logarithm or a probability through its log-odds. A mode taken there and reported back on the original scale is a different estimator from the original-scale mode, and the gap widens as the posterior gets more skewed.
- What would you do about it when the natural scale is genuinely ambiguous?Pick the scale the decision is made on, since that is the only non-arbitrary choice available. Failing that, report an invariant summary such as the median, or hand over the posterior itself so the consumer applies the summary their own loss demands rather than inheriting yours.
Stretch a rubber sheet unevenly and the highest point of a drawn hill can shift to a different part of the picture, even though no rock was moved. The peak is a property of the drawing, not of the terrain.
saying these in an interview costs you the question
- Says maximum likelihood is not invariant either
- Claims the posterior itself changes under reparameterisation
- Forgets the Jacobian in the change-of-variables rule
- Thinks the posterior mean is invariant under monotone transforms
- Says invariance failure only matters for improper posteriors