skip to content

Why does a zero-mean Gaussian prior on regression coefficients turn MAP into ridge regression?

level: middleimportance: should knowfreq 55%

answer

  1. log posterior splits into two additive terms
  2. log of a Gaussian is a squared term
  3. log of a Laplace is an absolute value
  4. penalty weight is a variance ratio
  5. lambda equals noise variance over prior variance

basics

~20 s

Taking logs turns the posterior into log-likelihood plus log-prior. A zero-mean Gaussian prior contributes minus the sum of squared coefficients divided by twice the prior variance, so maximising it is least squares with an L2 penalty attached.

solid answer

~50 s

With Gaussian noise of variance `sigma^2`, the log-likelihood of a linear model is `-RSS / (2 * sigma^2)` plus a constant, where `RSS` is the residual sum of squares. Put an independent zero-mean Gaussian prior with variance `tau^2` on each coefficient and the log-prior is `-sum(w_j^2) / (2 * tau^2)` plus a constant. Maximising the sum of those two is the same as minimising `RSS + (sigma^2 / tau^2) * sum(w_j^2)`, which is exactly the ridge objective with `lambda = sigma^2 / tau^2`. So the penalty strength is not a free knob in this view — it is the ratio of noise variance to prior variance. A tight prior (small `tau^2`) means a large `lambda` and heavy shrinkage; letting `tau^2` grow without bound sends `lambda` to zero and recovers ordinary least squares. Swapping the Gaussian prior for a Laplace prior replaces the squared term with an absolute-value term, giving the lasso objective instead.

go deeper

for a junior

Recall the direction of the mapping: a Gaussian prior on coefficients corresponds to a squared penalty and a Laplace prior to an absolute-value penalty. Knowing which prior goes with which penalty is the bar here.

for a middle

Be ready to do the algebra out loud: write the log-likelihood, add the log-prior, drop constants and read off the penalty weight as the ratio of noise variance to prior variance.

for a senior

Show that you know the correspondence holds for the posterior mode only, and explain why the posterior mean under a Laplace prior is not sparse. Mention leaving the intercept unpenalised and what that means as a prior.

for a principal

Own the framing that a penalty is an assumption about coefficient magnitudes made explicit. Be able to argue when stating that assumption as a prior improves how a modelling choice is communicated and reviewed, and when it is just relabelling.

## The correspondence in one line MAP estimation maximises `log-likelihood + log-prior`. Penalised estimation minimises `loss + penalty`. Because minimising a negated quantity is the same as maximising it, every prior corresponds to a penalty and every penalty corresponds to a prior: ``` penalty(w) = -(constant) * log p(w) ``` The two most familiar penalties in linear regression fall straight out of the two most familiar symmetric priors centred at zero. ## Deriving the ridge case Take a linear model where each response is the dot product of the predictors with a coefficient vector `w`, plus independent Gaussian noise of variance `sigma^2`. The log-likelihood of the data is ``` log p(data | w) = -RSS(w) / (2 * sigma^2) + constant ``` where `RSS(w)` is the residual sum of squares, `sum over i of (y_i - prediction_i)^2`. Maximising this alone is ordinary least squares — the Gaussian-noise MLE. Now place an independent zero-mean Gaussian prior with variance `tau^2` on every coefficient. A Gaussian density is proportional to `exp(-w_j^2 / (2 * tau^2))`, so summing over coefficients gives ``` log p(w) = -sum(w_j^2) / (2 * tau^2) + constant ``` Add the two and drop the constants: ``` maximise -RSS(w) / (2 * sigma^2) - sum(w_j^2) / (2 * tau^2) ``` Multiplying by `-2 * sigma^2` flips it into a minimisation without moving the optimum: ``` minimise RSS(w) + (sigma^2 / tau^2) * sum(w_j^2) ``` That is the ridge objective, and the correspondence is `lambda = sigma^2 / tau^2`. The penalty weight is the noise-to-prior variance ratio, which is where the usual shorthand *prior variance behaves like one over lambda* comes from. The two limits are worth being able to state instantly. As `tau^2` grows without bound the prior flattens, `lambda` goes to zero, and the MAP estimate becomes ordinary least squares — the flat-prior collapse of MAP onto the MLE, seen through a regression lens. As `tau^2` shrinks toward zero the prior insists the coefficients are near zero, `lambda` explodes, and the fitted coefficients are crushed toward the origin. One practical detail that interviewers like to hear: the intercept is normally left unpenalised, which in the Bayesian reading means giving it a flat prior rather than a zero-mean Gaussian one. Shrinking an intercept toward zero would be a statement that the response is centred near zero, which is rarely something you believe. ## Deriving the lasso case Keep the same Gaussian likelihood and swap the prior for an independent zero-mean Laplace (double-exponential) prior with scale `b`, whose density is proportional to `exp(-|w_j| / b)`. Then ``` log p(w) = -sum(|w_j|) / b + constant ``` and the same rearrangement gives ``` minimise RSS(w) + (2 * sigma^2 / b) * sum(|w_j|) ``` which is the lasso objective with an L1 penalty. The Laplace density has a sharp kink at zero — its derivative jumps rather than passing smoothly through — while the Gaussian density is flat-topped there, with zero derivative. That kink is exactly why the L1 posterior mode can sit at precisely zero for a coefficient whose data support is weak: the penalty keeps a constant pull toward zero no matter how close you get, whereas the quadratic pull of the Gaussian prior fades away as the coefficient approaches the origin. The Gaussian prior shrinks; the Laplace prior shrinks and can zero out. ## Where the analogy stops The correspondence is between penalised estimation and the **mode** of the posterior, not the posterior itself, and this is the point candidates most often miss. With a Gaussian likelihood and a Gaussian prior the posterior over the coefficients is itself Gaussian, hence symmetric, so the mode and the mean agree and the ridge solution really is the whole posterior's centre. Under a Laplace prior the posterior is not symmetric in the relevant way and, crucially, its **mean is not sparse**: the posterior assigns zero probability to a coefficient being exactly zero, so averaging over the posterior gives small non-zero values. Only the mode lands on exact zeros. A candidate who says *the Bayesian version of the lasso gives you sparsity for free* has confused the mode with the posterior. Similarly, reading `lambda` as `sigma^2 / tau^2` explains what the penalty *means* but does not by itself hand you a number, because you rarely know `tau^2`. What the Bayesian reading buys you is interpretation: a penalty is an assumption about coefficient magnitudes, stated as a distribution, and the strength of that assumption is measured relative to how noisy the data are.

  • Which prior gives the lasso instead, and what makes it produce exact zeros?
    A zero-mean Laplace, or double-exponential, prior. Its log density is proportional to minus the absolute value of the coefficient, giving an L1 penalty. The density has a kink at zero rather than a smooth flat top, so the pull toward zero does not fade as a coefficient approaches it, and the posterior mode can land exactly on zero.
  • Under a Laplace prior, is the posterior mean sparse as well as the mode?
    No. The posterior is a continuous distribution, so the event that a coefficient is exactly zero has probability zero and the mean is a small non-zero number. Only the mode sits at exact zeros. Sparsity here is a property of the point summary you chose, not of the posterior.
  • What does the ridge correspondence say about how the penalty should scale with noise?
    Since lambda equals noise variance over prior variance, holding beliefs about coefficient sizes fixed means noisier data should be penalised harder. The prior is a fixed statement about plausible coefficient magnitudes; the noisier the likelihood, the more weight that statement deserves relative to the fit.

saying these in an interview costs you the question

  • Says the prior changes the objective's shape but not its optimum
  • Claims a Gaussian prior yields the L1 penalty
  • Thinks the Bayesian posterior mean under a Laplace prior is sparse
  • Cannot connect prior variance to penalty strength at all
  • Says larger prior variance means stronger shrinkage
  • Treats ridge and MAP as unrelated coincidences

context