skip to content

Which prior makes the ridge estimate the MAP solution of a Bayesian linear regression?

level: seniorimportance: nice to knowfreq 26%

answer

  1. think likelihood times prior
  2. negative log turns products into sums
  3. a Gaussian exponent is a squared term
  4. zero-mean Gaussian on every weight
  5. lambda = noise variance over prior variance

basics

~20 s

An independent zero-mean Gaussian prior on the weights. With Gaussian noise of variance sigma squared and prior variance tau squared, the posterior mode is exactly the ridge estimate, with lambda equal to sigma squared divided by tau squared.

solid answer

~50 s

Put a Gaussian likelihood on the data - `y = Xw + noise` with noise `N(0, sigma^2)` - and an independent prior `w_j ~ N(0, tau^2)` on each weight. Take the negative log of likelihood times prior and the products become sums: the likelihood contributes the residual sum of squares divided by `2*sigma^2`, and the prior contributes `sum(w_j^2)` divided by `2*tau^2`. Maximising the posterior is therefore minimising `RSS + (sigma^2 / tau^2) * sum(w_j^2)`, which is ridge with `lambda = sigma^2 / tau^2`. The reading is intuitive: lambda is how noisy the data is relative to how large you believe the weights can be. A tight prior means strong shrinkage; a very vague prior recovers least squares. Note that this delivers only the posterior mode - the full posterior would also give you uncertainty on each weight.

go deeper

for a junior

Hold the headline only: a squared penalty is what you get from believing, before seeing any data, that coefficients are centred on zero and unlikely to be large.

for a middle

Be able to sketch the derivation - take the negative log of likelihood times prior, and the Gaussian prior's exponent turns into the sum of squared weights alongside the residual sum of squares.

for a senior

Explain what lambda equals in that view, noise variance over prior variance, and use it to reason about how much shrinkage a stated domain belief or a given data volume actually justifies.

for a principal

Decide whether the team needs the full posterior - uncertainty on weights and predictive intervals - or whether a tuned point estimate is sufficient for the decisions the model feeds, and budget the modelling effort accordingly.

## The setup Bayesian linear regression makes two modelling statements. First a likelihood: the target is a linear function of the features plus Gaussian noise, `y_i = x_i . w + e_i` with `e_i ~ N(0, sigma^2)`, independent across observations. Second a prior on the unknown weights, expressing what you believe before seeing data. Choose an independent zero-mean Gaussian, `w_j ~ N(0, tau^2)` for every j: you believe coefficients are centred on zero and unlikely to be enormous, with `tau` setting how unlikely. MAP - maximum a posteriori - estimation returns the single weight vector that maximises the posterior density, which by Bayes' rule is proportional to likelihood times prior. ## The derivation Maximising a product is awkward, so take the negative logarithm and minimise instead; the logarithm is monotone, so the location of the optimum is unchanged. The Gaussian likelihood of the data contributes, up to constants that do not involve `w`, `sum_i (y_i - x_i . w)^2 / (2 * sigma^2)` because the exponent of a Gaussian density is a negative squared deviation over twice its variance. The Gaussian prior contributes, again up to constants, `sum_j w_j^2 / (2 * tau^2)` for exactly the same reason - it is a Gaussian in `w_j` centred on zero. Adding them and multiplying through by `2 * sigma^2`, which does not move the minimum, gives `RSS + (sigma^2 / tau^2) * sum_j w_j^2` That is the ridge objective, with `lambda = sigma^2 / tau^2`. ## Reading lambda The ratio is the whole payoff of the derivation. Lambda is noise variance over prior variance: how much the data is expected to lie, divided by how much freedom you grant the weights. - Noisy data, `sigma^2` large: lambda large, shrink hard, because individual observations should not be trusted to move a coefficient far. - A confident prior that effects are small, `tau^2` small: lambda large again, for a different reason. - A vague prior, `tau^2` very large: lambda tends to zero and MAP tends to plain least squares. A flat prior adds no penalty at all, which is why maximum likelihood and unpenalised least squares coincide here. This also explains a practical observation. The prior term is a fixed cost while the likelihood term accumulates over observations, so as the sample grows the data outvotes the prior and the fitted weights drift toward the least-squares answer. The tuned lambda typically falls as data accumulates. ## The prior assumes a common scale Using one `tau` for every weight is a statement that all coefficients are expected to be about the same size. That is only defensible once the predictors are on comparable scales - a single shared prior variance across a column in metres and a column in millions of currency units is claiming something nobody intended. This is the Bayesian restatement of the practical rule that predictors are standardised before a penalised fit. The intercept, by the same logic, is given a flat prior rather than the shared Gaussian, which is why it is left unpenalised. ## What MAP throws away Under a Gaussian likelihood and a Gaussian prior the posterior over the weights is itself Gaussian, with mean `(X'X + lambda I)^-1 X'y` and covariance proportional to `(X'X + lambda I)^-1`. Two things follow. First, because a Gaussian's mode and mean coincide, the ridge estimate in this conjugate case is both the posterior mode and the posterior mean. Second, the point estimate discards the covariance - and that covariance is the interesting part of a Bayesian treatment, since it says how uncertain each weight is and supports predictive intervals rather than bare point predictions. If uncertainty is what the decision needs, MAP is not enough; if a good point predictor is what the decision needs, the extra machinery may not be worth it. ## Why the view is worth having Three reasons. It makes lambda interpretable instead of a number pulled off a validation grid - you can reason about a plausible prior scale for coefficients in your domain and get an order of magnitude for lambda before tuning. It explains why shrinkage is principled rather than a hack: you are not distorting the fit, you are combining data with a prior belief. And it lets you place other penalties in the same frame - a different prior shape gives a different penalty, which is how the absolute-value penalty arises from a heavier-tailed prior. ## Where the analogy stops MAP is not fully Bayesian. It depends on the parameterisation - reparameterise the weights and the mode moves while the posterior does not - and it gives no uncertainty. And in practice you rarely know `sigma^2` or `tau^2`, so lambda is still chosen by held-out error or by estimating it from the data through the marginal likelihood, rather than read off from stated beliefs.

  • If ridge is only the MAP estimate, what does the point estimate leave on the table?
    The posterior itself. With a Gaussian likelihood and Gaussian prior the posterior over weights is Gaussian with covariance proportional to `(X'X + lambda I)^-1`, which quantifies how uncertain each weight is and supports predictive intervals. Ridge reports only the mode of that posterior - which in this conjugate case is also its mean - and discards the spread entirely.
  • What does the Gaussian-prior view say about how lambda should change as the dataset grows?
    The prior contributes a fixed cost while the likelihood term accumulates across observations, so the prior's influence per unit of evidence fades and the fit drifts toward least squares. In practice the tuned lambda falls as data accumulates: with more evidence you need less help from a prior belief, and heavy shrinkage that helped at a thousand rows will simply add bias at a million.
  • Does an equal prior variance on every weight assume anything about the features?
    Yes - that the coefficients live on comparable scales. A single tau across all weights says every feature's effect is expected to be about the same size, which only makes sense after the predictors have been standardised. If one column is in metres and another in millions of currency units, the shared prior is making a claim nobody chose to make.

saying these in an interview costs you the question

  • Names a uniform or flat prior as the one giving ridge
  • Says the prior is Laplace, confusing it with the absolute-value penalty
  • Treats lambda as the prior variance rather than a ratio of variances
  • Claims MAP delivers full posterior uncertainty
  • Puts the shared Gaussian prior on the intercept too

context