skip to content

MAP Versus MLE

Maximum likelihood picks the parameter that best explains the data while MAP maximises the posterior, so a Gaussian prior becomes ridge and a Laplace prior becomes lasso. Interviewers love this link.

on this pageshow

questions

4

How does a MAP estimate differ from a maximum likelihood estimate of the same parameter?

level: juniorimportance: must knowfreq 68%

answer

  1. one of them adds a prior
  2. likelihood times prior, then maximise
  3. argmax of the posterior density
  4. a flat prior makes them identical
  5. both return one number, not a distribution

basics

~20 s

Maximum likelihood picks the parameter value that makes the observed data most probable. MAP picks the value that maximises the posterior, which is the likelihood multiplied by a prior. When the prior is flat, the two estimates coincide.

solid answer

~50 s

Both return a single number found by maximising a function of the parameter, but they maximise different functions. The MLE maximises the likelihood `p(data | theta)`. The MAP estimate maximises the posterior `p(theta | data)`, which by Bayes' rule is proportional to `p(data | theta) * p(theta)`. Taking logs makes the relationship obvious: MAP maximises `log-likelihood + log-prior`, so MAP is MLE plus a term that pulls the answer toward parameter values the prior favours. Concretely, three flips giving three heads: the MLE of the heads probability is 1, because `p^3` is largest at `p = 1`. Put a `Beta(2, 2)` prior on it and the posterior is `Beta(5, 2)`, whose mode is 0.8 — the prior drags the estimate off the boundary. Use a flat `Beta(1, 1)` prior instead and the log-prior is constant, so the posterior mode goes back to 1, exactly the MLE.

go deeper

for a junior

Be ready to write both objectives on a whiteboard: likelihood alone versus likelihood times prior. Recall that logs turn the product into a sum and that a flat prior makes the two estimates identical.

for a middle

Expect to derive the collapse rather than assert it, and to work a small conjugate example end to end, such as a Beta prior with binomial data, showing exactly how the prior moves the estimate off a boundary value.

for a senior

Show judgment about when the prior term is actually doing work: small samples, near-boundary parameters, or a likelihood that is flat in some direction. Be explicit that a point estimate discards the uncertainty the analysis produced.

for a principal

Own the framing question of whether a single number should be reported at all. Be able to argue when a point summary is the right interface for a downstream consumer and when handing over the mode alone misleads the people acting on it.

## Two estimators, two objective functions Both maximum likelihood estimation (MLE) and maximum a posteriori estimation (MAP) answer the same shaped question — *give me one number for the unknown parameter* — and both answer it by maximising a function. They differ only in which function. Write the unknown parameter as `theta` and the observed data as `D`. - The **likelihood** is `L(theta) = p(D | theta)`: how probable the data you actually saw would be, if `theta` were the truth. The MLE is `argmax_theta L(theta)`. Note that the likelihood is a function of `theta` but is *not* a probability distribution over `theta` — it does not integrate to one over the parameter. - The **posterior** is `p(theta | D)`, obtained from Bayes' rule as `p(theta | D) = p(D | theta) * p(theta) / p(D)`. Here `p(theta)` is the prior, the distribution describing beliefs about `theta` before seeing the data, and `p(D)` is a normalising constant that does not depend on `theta`. The MAP estimate is `argmax_theta p(theta | D)`. Because `p(D)` does not involve `theta`, maximising the posterior is the same as maximising `p(D | theta) * p(theta)`. Taking logarithms — which does not move the maximiser, since the logarithm is increasing — gives the cleanest statement of the difference: ``` MLE: argmax log p(D | theta) MAP: argmax log p(D | theta) + log p(theta) ``` MAP is maximum likelihood with one extra additive term. That term is the only difference. ## The flat-prior collapse If the prior is constant over the region of the parameter space where the likelihood lives, `log p(theta)` is a constant, and adding a constant to a function does not change where it peaks. The MAP estimate is then identical to the MLE. This is the sense in which maximum likelihood is a special case of MAP estimation: it is MAP under a prior that expresses no preference between parameter values. One caveat worth carrying into an interview: flatness is a property of the parameterisation you happened to write down. A prior that is uniform on a probability is not uniform on that probability's log-odds, so the collapse is a statement about one specific parameterisation, not a universal one. ## A worked example: three heads in three flips A coin is flipped three times and lands heads three times. Let `p` be the probability of heads. The likelihood is `p^3`, which increases on the whole interval from 0 to 1, so the MLE is `p = 1`. Interviewers like this case precisely because the answer is absurd: three flips is not evidence that the coin can never land tails, yet the MLE sits on the boundary. Now add a `Beta(2, 2)` prior, a mild symmetric bump peaking at 0.5. The Beta family is conjugate to the binomial, so the posterior is `Beta(2 + 3, 2 + 0) = Beta(5, 2)`. The mode of a `Beta(a, b)` with both parameters above 1 is `(a - 1) / (a + b - 2)`, giving `4 / 5 = 0.8`. The estimate has been pulled inward, away from the impossible-looking boundary value. Replace that prior with the flat `Beta(1, 1)` — the uniform distribution on the interval — and the posterior is `Beta(4, 1)`, whose density is proportional to `p^3`. That is exactly the likelihood again, so its mode is 1: MAP has collapsed onto the MLE, as the general argument predicted. ## What MAP is and is not MAP is a **point summary of a posterior**. It is the location of the highest posterior density, and nothing more. It reports no spread, no interval, no shape. Two posteriors — one razor-sharp, one nearly flat — can share the same mode and imply completely different levels of confidence, and the MAP number cannot tell them apart. Doing a full Bayesian analysis and then reporting only the mode throws away most of what the analysis computed. MAP is also not the posterior mean and not the posterior median. For a symmetric unimodal posterior all three coincide, which is why they are so often conflated; for a skewed posterior they are three different numbers. ## How the two behave as data accumulates The log-likelihood grows with the sample size — it is a sum over observations — while the log-prior stays fixed. So the prior term is progressively outvoted, and under standard regularity conditions (a fixed prior that assigns positive density near the true value, a well-behaved model) the MAP estimate and the MLE converge to each other and to the same answer. The prior matters most exactly where you would want it to: when data are scarce, when a parameter is near a boundary, or when the likelihood is nearly flat in some direction and needs something to break the tie.

  • Does a uniform prior always guarantee that MAP equals the MLE?
    It guarantees it in the parameterisation where the prior is actually constant, and only if that constant covers the region containing the likelihood's peak. A prior uniform on a probability is not uniform on its log-odds, so the same prior can collapse MAP onto the MLE in one parameterisation and not in another.
  • Does a MAP estimate tell you anything about uncertainty?
    No. It is the location of the posterior's peak and carries no information about how sharp that peak is. A near-flat posterior and a very concentrated one can share a mode. Any statement about uncertainty has to come from the posterior's spread or shape, not from the mode.
  • What happens to the gap between MAP and MLE as the sample grows?
    It shrinks. The log-likelihood accumulates one term per observation while the log-prior stays fixed, so the prior contribution is progressively outweighed. Under standard regularity conditions, with a fixed prior that puts positive density near the true value, the two estimates converge.

Maximum likelihood listens only to today's evidence. MAP listens to the evidence and to what you already believed, then reports the single most plausible answer after hearing both.

saying these in an interview costs you the question

  • Says MAP is the mean of the posterior
  • Claims MLE makes no assumptions while MAP does
  • Thinks MAP returns a distribution rather than a number
  • Says the prior dominates regardless of sample size
  • Confuses maximising the posterior with maximising the prior
  • Cannot state that a flat prior makes MAP equal the MLE

context

open as a page

Why does a zero-mean Gaussian prior on regression coefficients turn MAP into ridge regression?

level: middleimportance: should knowfreq 55%

basics

~20 s

Taking logs turns the posterior into log-likelihood plus log-prior. A zero-mean Gaussian prior contributes minus the sum of squared coefficients divided by twice the prior variance, so maximising it is least squares with an L2 penalty attached.

open as a page

When is the posterior mode a poor summary of a skewed posterior distribution?

level: seniorimportance: should knowfreq 45%

basics

~10 s

A skewed posterior has three different point summaries. Under right skew the mode is the smallest of the three, and it optimises an all-or-nothing loss that matches almost no real decision.

open as a page

Why is a MAP estimate not invariant under reparameterisation while the MLE is?

level: seniorimportance: nice to knowfreq 22%

basics

~20 s

The posterior is a probability density, so changing variables multiplies it by a Jacobian that reshapes the curve and can move its peak. The likelihood is not a density over the parameter, gets no Jacobian, so its maximiser transforms along.

open as a page