skip to content

Explicit Penalties

Ridge shrinks every coefficient toward zero, lasso drives some to exactly zero, and elastic net mixes the two. Interviewers ask 'L1 vs L2?' to see whether you can explain the geometry, not recite it.

on this pageshow

explore

questions

10

What penalty does elastic net add to a linear model, and why mix L1 with L2?

level: juniorimportance: must knowfreq 62%

answer

  1. two penalty terms, not one
  2. sparsity from one, stability from the other
  3. a dial that slides between lasso and ridge
  4. absolute values zero, squares group

basics

~20 s

Elastic net penalises the sum of absolute coefficients (L1) and the sum of squared coefficients (L2) together. The L1 part drives weak predictors to exactly zero; the L2 part keeps correlated predictors together and makes the fit stable.

solid answer

~50 s

Elastic net minimises `loss + lambda * [alpha * sum |w_j| + (1 - alpha) * sum w_j^2 / 2]`. The mixing ratio `alpha` is the L1 share: 1 is a pure lasso, 0 is a pure ridge, anything between is a real mixture, and `lambda` sets the overall strength. You mix because each penalty alone has a defect on wide, correlated data. A pure L1 fit gives you selection but picks roughly one predictor out of each correlated group, close to arbitrarily, and its solution is not unique when columns are collinear. A pure L2 fit is stable and unique but never zeroes anything. Adding the squared term makes the objective strictly convex, so the solution is unique and correlated predictors share weight, while the absolute-value term keeps the kink at zero that produces genuine sparsity.

go deeper

for a junior

Be ready to name both terms and say what each does: absolute values create exact zeros, squares shrink smoothly and never zero. Knowing which model each extreme of the mixing ratio recovers is the screening bar.

for a middle

An interviewer expects the mechanics: why the kink at zero produces sparsity while a squared term cannot, and why adding the squared term makes the objective strictly convex and the solution unique.

for a senior

Show when you would reach for the mix in practice — wide, correlated designs where a pure L1 fit selects a different set on every refit — and be candid about what it costs in parsimony and tuning effort.

for a principal

Own the framing that every penalty buys variance reduction with bias, and that the mixing choice is really a policy about what your team ships: a short unstable feature list or a longer reproducible one.

## What an elastic net actually minimises An unpenalised linear fit picks coefficients `w` to minimise the training loss alone — for squared-error regression, `loss = sum over rows of (y - w.x)^2`. When predictors are many or strongly related to each other, that objective has too much freedom: coefficients grow large, absorb noise, and swing dramatically if you refit on a slightly different sample. A penalty puts a price on large coefficients: ``` objective = loss + lambda * penalty(w) ``` Two classic penalties sit at either end: - **L2 (ridge)**: `penalty = sum w_j^2`. Every coefficient is pulled toward zero roughly in proportion to its size. Nothing ever lands exactly on zero, so no predictor is removed. - **L1 (lasso)**: `penalty = sum |w_j|`. Coefficients are shrunk and, once the penalty outweighs a predictor's contribution to the loss, pinned at exactly zero — so the fit performs selection. Elastic net applies **both at once**, inside a single convex objective: ``` penalty = alpha * sum |w_j| + (1 - alpha) * sum w_j^2 / 2 objective = loss + lambda * penalty ``` `alpha` in [0, 1] is the **mixing ratio** — the share of the penalty that is L1. `alpha = 1` is a pure lasso, `alpha = 0` is a pure ridge, and anything between is a genuine mixture. `lambda` is the **overall strength**: how hard the whole penalty presses. A warning worth carrying into interviews: symbol conventions differ between implementations — some use `alpha` for the strength and a separate name for the ratio — so always confirm which dial you are turning. ## Why mixing beats either penalty alone A pure L1 fit has three well-known weaknesses on wide, correlated data: 1. **Arbitrary choice among correlated predictors.** Given a group of predictors that carry essentially the same information, an L1 fit tends to keep about one of them and zero the rest, and which one it keeps is close to a coin flip decided by sampling noise. 2. **A hard cap when there are more predictors than rows.** With `p > n`, an L1 solution cannot place more than `n` non-zero coefficients; the path saturates there. 3. **Non-uniqueness.** With collinear columns the L1 objective is convex but not *strictly* convex, so several coefficient vectors can achieve the same optimum. Refits on resampled data then disagree about which predictors were "selected". A pure L2 fit has none of those problems — it is strictly convex, unique, and stable — but it never zeroes anything, so on 8,000 predictors you ship 8,000 coefficients and no selection story. The mixture keeps the useful half of each. The absolute-value term still creates a kink at zero, so exact zeros and genuine selection survive. The squared term makes the objective **strictly convex** even when `p > n`, which guarantees a unique solution, lifts the `n` cap, and pulls the coefficients of correlated predictors toward each other rather than letting one win — the **grouping effect**. In geometric terms, the feasible region keeps the L1 corners that sit on the axes (that is where zeros come from) while the rest of its boundary is rounded by the squared term (that is where the grouping and the stability come from). ## Three things the mix does not mean - **A mixing ratio of 0.5 does not mean half the coefficients become zero.** How many zeros you get depends on the overall strength and on how much unique signal each predictor carries. - **Adding an L2 share does not weaken regularisation.** It adds a second penalty term; the fit is more constrained, not less. What changes is the *character* of the constraint. - **The two penalties are not applied in sequence.** There is no "run lasso, then ridge the survivors". One objective, both terms, solved together. ## Choosing the regime - Genuinely sparse signal with near-independent predictors: a pure L1 fit is fine and is more parsimonious. - Dense signal where nearly every predictor contributes a little: a pure L2 fit is the honest model; forcing zeros only discards signal. - Wide and correlated — the case elastic net was designed for: mix, and let cross-validation tell you how much L1 the data will support. If your tuning lands on `alpha = 1` or `alpha = 0`, that is information, not a failure: the data is telling you it is in one of the pure regimes. ## What the penalty costs you Two hyperparameters now need tuning instead of one, and they interact. Every penalty also buys its variance reduction with bias: the surviving coefficients are shrunk toward zero, so they are *not* unbiased effect-size estimates and should not be read as if they were. The intercept is normally left out of the penalty, so that shifting the target up or down does not change the fit's character.

  • What fit do you get at each extreme of the mixing ratio?
    An L1 share of 1 leaves only the absolute-value term, which is a pure lasso. An L1 share of 0 leaves only the squared term, which is a pure ridge. Neither extreme removes the penalty itself — that only happens when the overall strength goes to zero, recovering the unpenalised fit.
  • Does an elastic net still produce coefficients that are exactly zero?
    Yes, as long as the L1 share is non-zero. The absolute-value term is non-differentiable at zero, and that kink is what lets the optimum sit exactly on zero for a weak predictor. At the same overall strength you usually get somewhat fewer zeros than a pure L1 fit, because part of the penalty budget is now doing smooth shrinkage instead.
  • If tuning lands on a mixing ratio of zero, what has the data told you?
    That the signal is dense rather than sparse: enough predictors carry a little information that zeroing any of them costs more than it saves. A pure squared penalty is then the honest model, and forcing sparsity would be discarding signal to satisfy a preference for short feature lists.

Lasso is a sharp knife that cuts predictors away; ridge is a clamp that holds the whole fit steady but cuts nothing. Elastic net cuts with the clamp on, so it still removes predictors but not at random.

saying these in an interview costs you the question

  • Says elastic net is just a lasso with a different penalty strength
  • Claims a mixing ratio of 0.5 makes half the coefficients zero
  • Says an elastic net can never produce exactly zero coefficients
  • Thinks adding the squared term makes the model less regularised
  • Describes the two penalties as applied one after the other

context

open as a page

In ridge regression, what happens to the coefficients as lambda goes to 0 and as lambda grows very large?

level: juniorimportance: must knowfreq 68%

basics

~20 s

Ridge with lambda = 0 reproduces the ordinary least squares fit. As lambda grows, every slope coefficient shrinks smoothly toward zero and the model flattens toward a constant, trading variance for bias. No coefficient reaches exactly zero at finite lambda.

open as a page

Why does an L1 penalty on a linear model's coefficients drive some of them to exactly zero?

level: middleimportance: must knowfreq 78%

basics

~20 s

An L1 penalty adds lambda times the absolute value of each coefficient, and the slope of that term stays at lambda right up to zero. That constant pull can push a coefficient exactly to zero; an L2 penalty's pull fades away and never does.

open as a page

What does an L2 (ridge) penalty do to a linear regression's coefficients?

level: middleimportance: must knowfreq 78%

basics

~20 s

A ridge penalty adds lambda times the sum of squared coefficients to the squared-error objective. Every coefficient is pulled proportionally toward zero, none lands exactly on zero, and the fit trades a little bias for much steadier estimates.

open as a page

In a lasso, what happens to the fitted coefficients as the penalty strength lambda increases?

level: juniorimportance: should knowfreq 62%

basics

~20 s

Raising lambda shrinks every coefficient toward zero and pushes more of them to exactly zero, so the fit uses fewer predictors. Past a large enough lambda every slope is zero and only the intercept survives.

open as a page

Elastic net has two dials, the mixing ratio and the penalty strength — what does each change?

level: middleimportance: should knowfreq 44%

basics

~20 s

The mixing ratio sets the penalty's character — how much of it is L1 versus L2 — and the strength sets how much shrinkage is applied in total. They interact, so the best strength changes whenever the ratio changes.

open as a page

Why does elastic net beat lasso on an 8,000-gene panel of co-expressed modules with 200 patients?

level: seniorimportance: should knowfreq 36%

basics

~20 s

A pure L1 fit can place at most 200 non-zero coefficients when there are only 200 patients, and inside a co-expressed module it keeps roughly one gene chosen by sampling noise. Adding a squared-penalty share lifts that cap and keeps correlated genes together.

open as a page

Why does a lasso keep just one of two near-duplicate predictors, with the pick flipping across resamples?

level: seniorimportance: should knowfreq 48%

basics

~20 s

The L1 penalty charges the same total for splitting one effect across two nearly identical columns as for loading it all on one, so it is indifferent between them. A tiny noise-level difference decides the winner, and resampling can reverse it.

open as a page

Why does ridge stabilise a gearbox model whose three vibration channels are near-collinear?

level: seniorimportance: should knowfreq 55%

basics

~20 s

Three sensors on one housing carry nearly the same signal, so least squares gives huge cancelling coefficients. The squared penalty crushes exactly that badly determined direction and spreads the shared effect evenly across the channels.

open as a page

Which prior makes the ridge estimate the MAP solution of a Bayesian linear regression?

level: seniorimportance: nice to knowfreq 26%

basics

~20 s

An independent zero-mean Gaussian prior on the weights. With Gaussian noise of variance sigma squared and prior variance tau squared, the posterior mode is exactly the ridge estimate, with lambda equal to sigma squared divided by tau squared.

open as a page