Why does a lasso keep just one of two near-duplicate predictors, with the pick flipping across resamples?
answer
- what does the penalty cost for each split
- flat penalty, barely tilted loss
- the tie is broken at noise scale
- a zero means redundant, not irrelevant
- check selection frequency over resamples
basics
~20 sThe L1 penalty charges the same total for splitting one effect across two nearly identical columns as for loading it all on one, so it is indifferent between them. A tiny noise-level difference decides the winner, and resampling can reverse it.
solid answer
~50 sWhen two columns carry almost the same information, the penalty cannot tell the difference between them: for a fixed combined effect with matching signs, `|w1| + |w2|` is the same however you split it, so the penalty is flat along that direction. The residual-sum-of-squares term then breaks the tie, and with near-duplicates that tie is broken by noise-scale differences. Because the penalty is flat but the loss is slightly curved, the optimum sits at an end point: one column takes the whole effect and the other goes to zero. Resample the rows and the noise shifts, so the winner can swap. Practically this means a zero here says "redundant given the other one", never "irrelevant". I would report the survivor as a representative of a correlated cluster, measure selection frequency across bootstrap resamples before naming drivers, and consider a penalty that mixes in a squared term if I need the pair treated together.
go deeper
Remember the headline: with two nearly identical predictors a lasso tends to keep one and zero the other, and which one it keeps is not meaningful. Never present a zeroed feature as proven unimportant.
Explain the mechanism: the penalty costs the same however a shared effect is split, so it is flat along that direction and the residual term breaks the tie on tiny differences. Note that exactly identical columns make the solution non-unique.
Show the operating response: correlation-cluster the inputs, measure selection frequency across bootstrap resamples, pick the representative on engineering or availability grounds, and distinguish the negligible effect on accuracy from the serious effect on the explanation.
Own the call on what the model is for. If a sparse coefficient list is going to be read as a causal driver story, decide whether to ship it at all, what stability evidence must accompany it, and who is accountable when the list changes at the next refit.
## The situation A credit-risk model contains two bureau-derived utilisation ratios — revolving balance over limit, computed at two nearby cut dates. They correlate at something like 0.98 because they measure the same behaviour a few weeks apart. You fit a lasso, and one of them gets a healthy coefficient while the other is exactly zero. You refit on a bootstrap resample of the same rows and the two swap places. Nothing is broken; this is what the L1 penalty does with redundancy, and knowing why is the difference between reporting a finding and reporting an artefact. ## Why the penalty is indifferent Take two columns that are essentially the same, and suppose the model wants a combined effect of size `c` from them. Any split with matching signs — `w1 + w2 = c`, both non-negative — produces almost the same fitted values, and the penalty cost is `|w1| + |w2| = c` for *every* such split. The L1 term is completely flat along that direction. It has no preference between (c, 0), (0, c) and (c/2, c/2). If the two columns were *exactly* identical, that flatness would be the end of the story: the solution would not be unique, and any split would be an equally valid answer. In practice the columns differ slightly, so the squared-error term is very slightly curved along that direction and does have a preference — but the preference is decided by the small, largely accidental differences between the columns on the particular rows you happened to sample. And because the penalty contributes a flat surface while the loss contributes a shallow tilt, the minimum over the shared budget sits at an end of the range rather than in the middle. One column takes essentially the whole effect; the other is zeroed. That is also why the outcome is unstable. Draw a different sample, or a bootstrap resample of the same one, and the shallow tilt can point the other way. Predictive performance barely notices — both fits encode the same information — but the coefficient story changes completely, which matters enormously if the coefficient list is the deliverable. ## Contrast worth stating A squared penalty behaves differently on the same pair: because it charges `w1^2 + w2^2`, splitting an effect evenly is genuinely cheaper than concentrating it, so correlated columns get similar, smaller coefficients rather than one winner. That is the standard motivation for mixing a squared term into the L1 objective when the correlated pair should be treated as a unit — but note that you are then choosing stability over sparsity, since neither column will be dropped. ## What a zero means, precisely Write the interpretation down before anyone else does: a zero coefficient means *this column adds little once the columns already in the model are present, at this value of lambda, on this sample*. It does not mean the feature is unrelated to the outcome, it does not mean the survivor is the causally important one, and it is not a significance test. In the credit example, the dropped utilisation ratio may be every bit as predictive on its own; it simply has nothing left to contribute after its twin is included. ## How to handle it in practice - **Cluster before you interpret.** Group predictors by correlation and treat each cluster as one candidate driver. Report the survivor as "utilisation ratio (one of two near-identical cut-date variants)", not as a distinct finding. - **Measure selection stability.** Refit at the chosen lambda on many bootstrap resamples and record how often each column is selected. A feature chosen in 95% of resamples supports a much stronger claim than one chosen in 45%, and the frequencies are far more honest than a single fitted model's zeros and non-zeros. - **Decide the representative deliberately.** If one of the pair is cheaper to compute, available earlier at scoring time, or more robust upstream, choose it yourself and drop the other before fitting rather than letting noise choose. - **Separate the two jobs.** If the model is for prediction, the arbitrary pick is largely harmless. If it is for explanation, do not let a single sparse fit be the explanation; the sparsity you like is exactly what makes the story fragile. ## What interviewers are listening for The weak answer treats the zero as evidence and moves on. The strong answer names the flatness of the penalty along the redundant direction, notes that the tie is broken at noise scale, distinguishes the impact on prediction from the impact on interpretation, and offers a concrete stability check rather than a vague promise to "be careful with correlated features".
- How would you demonstrate this instability to a stakeholder who wants a driver list?Refit at the chosen lambda on a few hundred bootstrap resamples and report each predictor's selection frequency instead of a single yes/no list. Drivers picked in nearly every resample can be presented as findings; ones picked half the time are shown as a correlated cluster with a note that the model chooses among them arbitrarily. It converts a fragile binary into a claim with visible uncertainty.
- Does the arbitrary pick hurt the model's predictive accuracy?Usually very little. The two columns carry nearly the same information, so whichever survives the held-out error is about the same, and swapping the winner between refits barely moves the metric. What degrades is the interpretation: the coefficient list changes between runs, which undermines any narrative built on which features the model kept.
saying these in an interview costs you the question
- Reads a zero coefficient as proof the feature is irrelevant
- Says the survivor is the causally more important driver
- Expects the same selected set from every refit
- Claims correlated inputs make the predictions badly wrong
- Offers only 'drop correlated features' with no criterion