skip to content

Regularization

You will learn how penalty terms trade a little bias for a lot of variance, why L1 zeroes weights while L2 shrinks them, and how to tune the regularization strength. Interviewers use 'L1 vs L2?' as a fast probe of whether your bias-variance understanding is operational.

on this pageshow

explore

questions

26

In a lasso coefficient path drawn with lambda decreasing left to right, what do the two ends show?

level: juniorimportance: must knowfreq 55%

answer

  1. one line per predictor, not per row
  2. the horizontal axis is penalty strength
  3. one end is a constant-only model
  4. the other end pays no penalty at all

basics

~20 s

At the far left the penalty is largest and every coefficient is exactly zero, so the model predicts a constant. Moving right, predictors enter one by one, and the far right is the unpenalised least-squares fit.

solid answer

~50 s

A coefficient path draws one line per predictor: the fitted coefficient as a function of the penalty strength lambda. With lambda decreasing left to right, the left edge sits at `lambda_max`, the smallest penalty that drives every coefficient to zero, so the model there is just the intercept. As lambda falls, predictors switch on one at a time and their lines move away from zero; at the right edge, where lambda is effectively zero, you recover the ordinary least-squares fit. On a used-car listing-price model you would typically see mileage enter first, then model year, then trim level, with colour arriving only at the far right. That axis is a bias-variance dial: the all-zero left end is maximally biased with almost no variance, the unpenalised right end has the lowest bias and the highest variance, and a useful model normally sits somewhere in between.

go deeper

for a junior

Be ready to say what sits at each end: an all-zero, intercept-only model at the largest penalty and the plain least-squares fit at zero penalty. Knowing that lasso lines can hit exactly zero while ridge lines do not is enough here.

for a middle

Explain the mechanics: why a finite lambda_max zeroes everything, why predictors enter one at a time, and why the lasso path is straight lines between knots while a ridge path is smooth. Expect to contrast the two penalties on the same plot.

for a senior

Show how you use the plot in real work — sanity-checking that the predictors you expect enter early, spotting that a near-unpenalised end is unstable, and refusing to pick lambda by eye when an out-of-sample estimate is what decides it.

for a principal

Own the framing: the path is the cheapest artefact for arguing capacity tradeoffs with non-specialists, and the temptation to present it as a feature-importance ranking is the thing to head off before it reaches a slide deck.

## What a coefficient path is A penalised linear model does not produce one set of coefficients; it produces a *family* of them, one for every value of the penalty strength lambda. The lasso minimises ``` (1 / 2n) * sum_i (y_i - b0 - sum_j x_ij * b_j)^2 + lambda * sum_j |b_j| ``` and the solution `b(lambda)` moves continuously as you turn lambda. A **coefficient path plot** simply draws that family: the horizontal axis is the penalty (usually `log(lambda)`, or the total L1 norm of the solution), the vertical axis is coefficient value, and each predictor contributes one line. It is the single most informative picture of what a penalty is doing to your model. ## The left end: lambda_max and the null model There is a finite penalty above which nothing survives. Call it `lambda_max`: it is the smallest lambda for which the whole coefficient vector is zero, and it equals the largest absolute inner product between any single predictor column and the response, scaled by the sample size (with the `1 / 2n` least-squares convention above). Intuitively, a predictor only earns a nonzero coefficient when the squared-error it removes outweighs the `lambda * |b_j|` it must pay; when lambda exceeds the best predictor's payoff, no predictor can pay, and every coefficient is zero. So at the far-left edge the fitted model is a constant — the intercept, which is left unpenalised — and it predicts the mean response for every row. That model has no variance worth mentioning (resample the data and it barely moves) and enormous bias if the predictors really do carry signal. ## The right end: the unpenalised fit At the other extreme lambda goes to zero, the penalty term vanishes, and the objective reduces to plain least squares. Every predictor is in the model at its full unshrunk coefficient. This end has the lowest bias the linear family can offer and the highest variance: with many predictors, or with correlated ones, small changes in the training rows swing these coefficients a lot. When predictors outnumber rows the least-squares fit is not even unique, which is why paths are usually stopped short of lambda = 0. ## The middle: entry order and shrinkage Between the ends, two things happen at once. Predictors **enter** — their coefficients leave zero — and predictors already in the model keep **growing** away from zero as the penalty relaxes. For a used-car listing-price model, mileage might enter first, model year next, trim level later, and colour only near the unpenalised end; a reader can see at a glance which handful of predictors carries most of the signal at any given penalty. Two mechanical facts are worth knowing. First, the lasso path is *piecewise linear* in lambda: between the penalty values where the set of nonzero coefficients changes (the knots), every coefficient traces a straight line. Second, the ridge (L2) path looks quite different — coefficients shrink smoothly toward zero but generically never reach it, so all lines are nonzero everywhere and there is no entry order at all. If a path plot shows lines flattening onto zero at distinct points, you are looking at an L1 penalty. Paths are conventionally drawn after standardising the predictors, so that the vertical positions of different lines are comparable; that convention is a topic in its own right. ## Reading the axis as a dial The most useful habit is to read the horizontal axis as a capacity dial. Sliding left buys stability at the price of systematic error; sliding right buys fit at the price of instability. Neither end is the answer. Where exactly to stop is decided by out-of-sample error, not by looking at the coefficient lines — the path tells you *what the model becomes* at each penalty, not *which penalty generalises best*. ## Common misreadings The path is not a training history: nothing is iterating, and there is no time axis. Moving right is not "the model learning"; it is you choosing a weaker penalty. Also, a coefficient sitting at zero over most of the path does not prove its predictor is unrelated to the response — it may simply be redundant given the predictors that entered earlier.

  • How does the same plot look if you swap the lasso's L1 penalty for a ridge L2 penalty?
    Every line is nonzero across the whole plot. Ridge shrinks coefficients smoothly toward zero as the penalty grows but generically never sets one exactly to zero, so there is no entry order and no sparsity: all predictors are present at every penalty, just progressively smaller. The lines are also smooth curves rather than the lasso's straight segments joined at knots.
  • What exactly is lambda_max, and why is the solution there known without fitting anything?
    It is the smallest penalty that makes the whole coefficient vector zero, equal to the largest absolute predictor-response inner product scaled by the sample size. Above it, no predictor removes enough squared error to pay its own penalty, so the optimum is the all-zero vector. That makes the left edge a free, exact starting point rather than something you have to solve for.
  • Can you pick the deployed model by eye from the path plot alone?
    No. The path shows what the model becomes at each penalty, not how well each one generalises. Choosing a penalty needs an out-of-sample estimate of error over the same lambda grid; the path is then read at the winning lambda to see which predictors survived and how large their coefficients are.

It is a dimmer switch photographed at every setting: full dark is the intercept-only model, full brightness is the unpenalised fit, and each predictor's light comes on at its own point on the dial.

saying these in an interview costs you the question

  • Treats the unpenalised right end as the best model
  • Reads the horizontal axis as training iterations or epochs
  • Says ridge paths also drive coefficients to exactly zero
  • Claims a coefficient stuck at zero proves no relationship exists
  • Thinks every predictor enters the path at the same penalty

context

open as a page

What is early stopping in an iteratively fitted model, and why does it act as regularization?

level: juniorimportance: must knowfreq 72%

basics

~20 s

Early stopping halts an iterative fit once a held-out validation score stops improving, and keeps the best-scoring iterate. Each extra iteration lets the model absorb finer detail from the training rows, so stopping sooner limits effective capacity and cuts variance.

open as a page

What penalty does elastic net add to a linear model, and why mix L1 with L2?

level: juniorimportance: must knowfreq 62%

basics

~20 s

Elastic net penalises the sum of absolute coefficients (L1) and the sum of squared coefficients (L2) together. The L1 part drives weak predictors to exactly zero; the L2 part keeps correlated predictors together and makes the fit stable.

open as a page

In ridge regression, what happens to the coefficients as lambda goes to 0 and as lambda grows very large?

level: juniorimportance: must knowfreq 68%

basics

~20 s

Ridge with lambda = 0 reproduces the ordinary least squares fit. As lambda grows, every slope coefficient shrinks smoothly toward zero and the model flattens toward a constant, trading variance for bias. No coefficient reaches exactly zero at finite lambda.

open as a page

Why must predictors be standardised before fitting a ridge or lasso model?

level: juniorimportance: must knowfreq 72%

basics

~20 s

An L1 or L2 penalty adds up coefficient sizes, and a coefficient's size depends on the units of its predictor. Standardising puts every predictor on a common spread so one lambda penalises them all comparably.

open as a page

When you average 50 differently-seeded high-variance fits, what caps the variance reduction?

level: middleimportance: must knowfreq 58%

basics

~20 s

Correlation between their errors. Averaging M fits whose errors have pairwise correlation rho leaves rho times the single-fit variance no matter how large M grows; only the independent share, (1 - rho)/M of the variance, is averaged away.

open as a page

Why does an L1 penalty on a linear model's coefficients drive some of them to exactly zero?

level: middleimportance: must knowfreq 78%

basics

~20 s

An L1 penalty adds lambda times the absolute value of each coefficient, and the slope of that term stays at lambda right up to zero. That constant pull can push a coefficient exactly to zero; an L2 penalty's pull fades away and never does.

open as a page

What does an L2 (ridge) penalty do to a linear regression's coefficients?

level: middleimportance: must knowfreq 78%

basics

~20 s

A ridge penalty adds lambda times the sum of squared coefficients to the squared-error objective. Every coefficient is pulled proportionally toward zero, none lands exactly on zero, and the fit trades a little bias for much steadier estimates.

open as a page

In a lasso, what happens to the fitted coefficients as the penalty strength lambda increases?

level: juniorimportance: should knowfreq 62%

basics

~20 s

Raising lambda shrinks every coefficient toward zero and pushes more of them to exactly zero, so the fit uses fewer predictors. Past a large enough lambda every slope is zero and only the intercept survives.

open as a page

Why is adding Gaussian noise to a least-squares fit's inputs equivalent to a ridge penalty?

level: middleimportance: should knowfreq 32%

basics

~20 s

In expectation over the noise, the jittered squared error equals the clean squared error plus n*sigma^2 times the squared weight norm — the ridge objective. Large weights amplify input noise, so the fit keeps them small.

open as a page

Elastic net has two dials, the mixing ratio and the penalty strength — what does each change?

level: middleimportance: should knowfreq 44%

basics

~20 s

The mixing ratio sets the penalty's character — how much of it is L1 versus L2 — and the strength sets how much shrinkage is applied in total. They interact, so the best strength changes whenever the ratio changes.

open as a page

Under the one-standard-error rule, with 5-fold CV mean error 2.41 at lambda = 0.01 (standard error 0.09), 2.49 at 0.05 and 2.58 at 0.10, which lambda wins?

level: middleimportance: should knowfreq 45%

basics

~10 s

Lambda = 0.05. Take the lowest mean error, 2.41, add the standard error at that point, 0.09, to get a threshold of 2.50, then pick the largest lambda whose mean still sits under it.

open as a page

Why is the intercept left unpenalised in ridge and lasso regression?

level: middleimportance: should knowfreq 48%

basics

~20 s

The intercept only sets the model's overall level, and zero is not a neutral level. Shrinking it would drag every prediction toward zero and make the fit depend on the target's arbitrary origin, so it is fitted freely.

open as a page

Why is a lasso path's entry order a weak ranking of feature importance?

level: seniorimportance: should knowfreq 42%

basics

~20 s

Entry order records which predictor was most correlated with the leftover residual at that penalty, in this one sample. It shifts with scaling, with correlations between predictors, and across resamples, and says nothing about effect size.

open as a page

How do you set the patience for early stopping when the validation curve is jagged?

level: seniorimportance: should knowfreq 48%

basics

~20 s

Set early-stopping patience from the noise scale of the monitored curve: it must outlast the runs of worsening checks that random variation alone produces. Patience of one halts still-improving fits; over-long patience costs only compute.

open as a page

Why does elastic net beat lasso on an 8,000-gene panel of co-expressed modules with 200 patients?

level: seniorimportance: should knowfreq 36%

basics

~20 s

A pure L1 fit can place at most 200 non-zero coefficients when there are only 200 patients, and inside a co-expressed module it keeps roughly one gene chosen by sampling noise. Adding a squared-penalty share lifts that cap and keeps correlated genes together.

open as a page

Why does a lasso keep just one of two near-duplicate predictors, with the pick flipping across resamples?

level: seniorimportance: should knowfreq 48%

basics

~20 s

The L1 penalty charges the same total for splitting one effect across two nearly identical columns as for loading it all on one, so it is indifferent between them. A tiny noise-level difference decides the winner, and resampling can reverse it.

open as a page

A penalised model's CV-error curve is flat across a whole decade of lambda — why is the lambda with the lowest mean error a poor pick?

level: seniorimportance: should knowfreq 34%

basics

~10 s

On a flat curve the winner is decided by fold noise, not by any real difference: reshuffle the folds and the minimum jumps to a neighbouring lambda. That lowest mean is also optimistically low.

open as a page

Why does ridge stabilise a gearbox model whose three vibration channels are near-collinear?

level: seniorimportance: should knowfreq 55%

basics

~20 s

Three sensors on one housing carry nearly the same signal, so least squares gives huge cancelling coefficients. The squared penalty crushes exactly that badly determined direction and spreads the shared effect evenly across the channels.

open as a page

One pooled model with shared coefficients or five per-product-line fits — how do you decide?

level: principalimportance: should knowfreq 42%

basics

~20 s

Read parameter sharing as a regularizer: a shared coefficient block is estimated from all the data and cuts variance, at the cost of bias if the lines really differ. Decide on per-line volume, similarity, and cold start.

open as a page

Why is a lasso fit computed down a grid of lambdas from lambda_max rather than at one lambda?

level: middleimportance: nice to knowfreq 30%

basics

~20 s

You rarely know the right penalty in advance, so the whole path is needed anyway. The descending grid makes it cheap: at lambda_max the solution is known to be all zeros, and each fit resumes from the previous one.

open as a page

Why does early-stopped gradient descent shrink coefficients much like a ridge penalty?

level: middleimportance: nice to knowfreq 30%

basics

~20 s

Gradient descent started at zero moves fastest along directions the data determines well and slowest along weak ones, leaving the weak ones near zero. A ridge penalty shrinks exactly those most, and more iterations act like a smaller penalty.

open as a page

Why does smoothing training labels stop a log-loss classifier saturating at 0 and 1?

level: seniorimportance: nice to knowfreq 24%

basics

~20 s

Log loss against a hard 0/1 target keeps falling as the prediction approaches it, so the fit inflates coefficients without limit. Smoothing targets to 0.95 and 0.05 puts the minimum there, capping the log-odds and the weights.

open as a page

Which prior makes the ridge estimate the MAP solution of a Bayesian linear regression?

level: seniorimportance: nice to knowfreq 26%

basics

~20 s

An independent zero-mean Gaussian prior on the weights. With Gaussian noise of variance sigma squared and prior variance tau squared, the posterior mode is exactly the ridge estimate, with lambda equal to sigma squared divided by tau squared.

open as a page

Should one-hot dummies and rare binary flags be standardised before a penalised fit?

level: seniorimportance: nice to knowfreq 24%

basics

~10 s

There is no universal answer. Dividing a rare 0/1 flag by its tiny standard deviation stretches its range and lets it escape most of the shrinkage, though it rests on a handful of rows.

open as a page

When would you ship the CV-minimum lambda instead of the larger one-standard-error lambda?

level: principalimportance: nice to knowfreq 16%

basics

~10 s

Take the minimum when accuracy is the entire objective and the dip is deep relative to the error bars. The one-standard-error rule spends a little expected accuracy to buy a simpler, more reproducible model.

open as a page