skip to content

What penalty does elastic net add to a linear model, and why mix L1 with L2?

level: juniorimportance: must knowfreq 62%

answer

  1. two penalty terms, not one
  2. sparsity from one, stability from the other
  3. a dial that slides between lasso and ridge
  4. absolute values zero, squares group

basics

~20 s

Elastic net penalises the sum of absolute coefficients (L1) and the sum of squared coefficients (L2) together. The L1 part drives weak predictors to exactly zero; the L2 part keeps correlated predictors together and makes the fit stable.

solid answer

~50 s

Elastic net minimises `loss + lambda * [alpha * sum |w_j| + (1 - alpha) * sum w_j^2 / 2]`. The mixing ratio `alpha` is the L1 share: 1 is a pure lasso, 0 is a pure ridge, anything between is a real mixture, and `lambda` sets the overall strength. You mix because each penalty alone has a defect on wide, correlated data. A pure L1 fit gives you selection but picks roughly one predictor out of each correlated group, close to arbitrarily, and its solution is not unique when columns are collinear. A pure L2 fit is stable and unique but never zeroes anything. Adding the squared term makes the objective strictly convex, so the solution is unique and correlated predictors share weight, while the absolute-value term keeps the kink at zero that produces genuine sparsity.

go deeper

for a junior

Be ready to name both terms and say what each does: absolute values create exact zeros, squares shrink smoothly and never zero. Knowing which model each extreme of the mixing ratio recovers is the screening bar.

for a middle

An interviewer expects the mechanics: why the kink at zero produces sparsity while a squared term cannot, and why adding the squared term makes the objective strictly convex and the solution unique.

for a senior

Show when you would reach for the mix in practice — wide, correlated designs where a pure L1 fit selects a different set on every refit — and be candid about what it costs in parsimony and tuning effort.

for a principal

Own the framing that every penalty buys variance reduction with bias, and that the mixing choice is really a policy about what your team ships: a short unstable feature list or a longer reproducible one.

## What an elastic net actually minimises An unpenalised linear fit picks coefficients `w` to minimise the training loss alone — for squared-error regression, `loss = sum over rows of (y - w.x)^2`. When predictors are many or strongly related to each other, that objective has too much freedom: coefficients grow large, absorb noise, and swing dramatically if you refit on a slightly different sample. A penalty puts a price on large coefficients: ``` objective = loss + lambda * penalty(w) ``` Two classic penalties sit at either end: - **L2 (ridge)**: `penalty = sum w_j^2`. Every coefficient is pulled toward zero roughly in proportion to its size. Nothing ever lands exactly on zero, so no predictor is removed. - **L1 (lasso)**: `penalty = sum |w_j|`. Coefficients are shrunk and, once the penalty outweighs a predictor's contribution to the loss, pinned at exactly zero — so the fit performs selection. Elastic net applies **both at once**, inside a single convex objective: ``` penalty = alpha * sum |w_j| + (1 - alpha) * sum w_j^2 / 2 objective = loss + lambda * penalty ``` `alpha` in [0, 1] is the **mixing ratio** — the share of the penalty that is L1. `alpha = 1` is a pure lasso, `alpha = 0` is a pure ridge, and anything between is a genuine mixture. `lambda` is the **overall strength**: how hard the whole penalty presses. A warning worth carrying into interviews: symbol conventions differ between implementations — some use `alpha` for the strength and a separate name for the ratio — so always confirm which dial you are turning. ## Why mixing beats either penalty alone A pure L1 fit has three well-known weaknesses on wide, correlated data: 1. **Arbitrary choice among correlated predictors.** Given a group of predictors that carry essentially the same information, an L1 fit tends to keep about one of them and zero the rest, and which one it keeps is close to a coin flip decided by sampling noise. 2. **A hard cap when there are more predictors than rows.** With `p > n`, an L1 solution cannot place more than `n` non-zero coefficients; the path saturates there. 3. **Non-uniqueness.** With collinear columns the L1 objective is convex but not *strictly* convex, so several coefficient vectors can achieve the same optimum. Refits on resampled data then disagree about which predictors were "selected". A pure L2 fit has none of those problems — it is strictly convex, unique, and stable — but it never zeroes anything, so on 8,000 predictors you ship 8,000 coefficients and no selection story. The mixture keeps the useful half of each. The absolute-value term still creates a kink at zero, so exact zeros and genuine selection survive. The squared term makes the objective **strictly convex** even when `p > n`, which guarantees a unique solution, lifts the `n` cap, and pulls the coefficients of correlated predictors toward each other rather than letting one win — the **grouping effect**. In geometric terms, the feasible region keeps the L1 corners that sit on the axes (that is where zeros come from) while the rest of its boundary is rounded by the squared term (that is where the grouping and the stability come from). ## Three things the mix does not mean - **A mixing ratio of 0.5 does not mean half the coefficients become zero.** How many zeros you get depends on the overall strength and on how much unique signal each predictor carries. - **Adding an L2 share does not weaken regularisation.** It adds a second penalty term; the fit is more constrained, not less. What changes is the *character* of the constraint. - **The two penalties are not applied in sequence.** There is no "run lasso, then ridge the survivors". One objective, both terms, solved together. ## Choosing the regime - Genuinely sparse signal with near-independent predictors: a pure L1 fit is fine and is more parsimonious. - Dense signal where nearly every predictor contributes a little: a pure L2 fit is the honest model; forcing zeros only discards signal. - Wide and correlated — the case elastic net was designed for: mix, and let cross-validation tell you how much L1 the data will support. If your tuning lands on `alpha = 1` or `alpha = 0`, that is information, not a failure: the data is telling you it is in one of the pure regimes. ## What the penalty costs you Two hyperparameters now need tuning instead of one, and they interact. Every penalty also buys its variance reduction with bias: the surviving coefficients are shrunk toward zero, so they are *not* unbiased effect-size estimates and should not be read as if they were. The intercept is normally left out of the penalty, so that shifting the target up or down does not change the fit's character.

  • What fit do you get at each extreme of the mixing ratio?
    An L1 share of 1 leaves only the absolute-value term, which is a pure lasso. An L1 share of 0 leaves only the squared term, which is a pure ridge. Neither extreme removes the penalty itself — that only happens when the overall strength goes to zero, recovering the unpenalised fit.
  • Does an elastic net still produce coefficients that are exactly zero?
    Yes, as long as the L1 share is non-zero. The absolute-value term is non-differentiable at zero, and that kink is what lets the optimum sit exactly on zero for a weak predictor. At the same overall strength you usually get somewhat fewer zeros than a pure L1 fit, because part of the penalty budget is now doing smooth shrinkage instead.
  • If tuning lands on a mixing ratio of zero, what has the data told you?
    That the signal is dense rather than sparse: enough predictors carry a little information that zeroing any of them costs more than it saves. A pure squared penalty is then the honest model, and forcing sparsity would be discarding signal to satisfy a preference for short feature lists.

Lasso is a sharp knife that cuts predictors away; ridge is a clamp that holds the whole fit steady but cuts nothing. Elastic net cuts with the clamp on, so it still removes predictors but not at random.

saying these in an interview costs you the question

  • Says elastic net is just a lasso with a different penalty strength
  • Claims a mixing ratio of 0.5 makes half the coefficients zero
  • Says an elastic net can never produce exactly zero coefficients
  • Thinks adding the squared term makes the model less regularised
  • Describes the two penalties as applied one after the other

context