skip to content

Why is a lasso path's entry order a weak ranking of feature importance?

level: seniorimportance: should knowfreq 42%

answer

  1. it depends on what already entered
  2. conditional, not marginal, relationship
  3. refit on resamples and watch it reshuffle
  4. coefficient lines can flip sign
  5. late entry means redundant, not unrelated

basics

~20 s

Entry order records which predictor was most correlated with the leftover residual at that penalty, in this one sample. It shifts with scaling, with correlations between predictors, and across resamples, and says nothing about effect size.

solid answer

~50 s

The order predictors enter a lasso path is a greedy, sample-specific artefact, not a ranking. A predictor enters when its correlation with the current residual first justifies paying the penalty, so entry depends on what has already entered: on a used-car listing-price model, mileage enters first, model year next, and colour last — but colour may be strongly related to price on its own and merely redundant once trim level is in. Individual coefficient lines are not even monotone: a predictor can enter, shrink back to zero, and re-enter with the opposite sign once a correlated partner is admitted. Refit the path on bootstrap resamples and the order visibly reshuffles among near-tied predictors. If you need a ranking, report how often each predictor is selected across resamples, and at what penalty, rather than one path's ordering.

go deeper

for a junior

Recall that entering the path early means a predictor was worth its penalty soonest, and that this is not the same as being the most important. Do not turn a path plot into a ranked feature list.

for a middle

Explain the mechanism: entry is driven by correlation with the current residual, so it is conditional on what already entered, and correlated predictors make the order fragile and coefficients non-monotone.

for a senior

Show the operating habit — resample the path, report selection frequency and magnitudes at the deployed penalty, and investigate sign flips before letting any coefficient story reach a stakeholder.

for a principal

Own the boundary between predictive selection and causal claims when a business partner asks which factors drive the outcome, and set the standard your team uses to report variable selection so one path plot never becomes the official ranking.

## What entry order actually measures As lambda falls, a predictor's coefficient leaves zero at the moment its absolute correlation with the *current residual* — what the already-active predictors have not explained — crosses the penalty threshold. So entry order is a **greedy sequence conditioned on the model so far**, computed on one particular sample. Read literally it says: "given everything admitted at higher penalties, this predictor was the next one worth paying for here." That is a genuinely useful diagnostic, and it is not a ranking of importance. ## Four reasons the ordering is weak **1. It is conditional, not marginal.** A predictor that arrives late may be highly related to the response on its own and merely redundant given earlier arrivals. In a used-car listing-price model, mileage enters first, model year second, trim level later, and colour last. Colour genuinely does move price — certain colours cluster in premium trims — but by the time colour is considered, trim level has already absorbed that signal. Late entry means "adds little beyond what is in the model", not "unrelated to price". **2. It is not stable.** Two predictors whose residual correlations are close can swap order under a small perturbation of the data. Refit the path on bootstrap resamples of the same rows and you will see near-tied predictors trade places, and marginal ones flicker in and out of the model entirely. A single ordering reported to three decimal places of confidence is over-claiming. **3. It depends on the measurement scale.** The penalty is applied to the raw coefficient magnitudes, so which predictor "wins" the next entry slot moves when you change units. Paths are drawn on standardised predictors precisely so that this comparison is at least on common footing. **4. Entry order is not effect size.** A predictor can enter early and still carry a modest coefficient at the penalty you eventually deploy, while a late entrant grows quickly once admitted. Order tells you *when*, magnitude tells you *how much*, and the two are different columns of information. ## Non-monotone lines and sign flips A common surprise: individual coefficients need not move monotonically as lambda changes. A predictor can enter the path, shrink back to zero, and re-enter with the **opposite sign** once a correlated partner is admitted. This is not a solver bug. When two predictors carry overlapping information, the first one to enter acts as a stand-in for the pair; once the partner enters and takes over the shared component, the first coefficient can be re-purposed to correct the partner's overshoot, which may require the other sign. What *is* monotone is the total L1 norm `sum_j |b_j|` of the solution: it is non-increasing as lambda grows. No individual line carries that guarantee. Sign flips along the path are a warning worth acting on. They tell you the sign you would quote for that coefficient at a chosen penalty is not robust, so any narrative built on "this factor pushes price up" needs checking before it goes in front of anyone. ## What to do instead - **Resample the path.** Refit on many bootstrap or subsample draws and record, for each predictor, the fraction of draws in which it is selected at (or above) your chosen penalty. Selection frequency is a far more honest summary than one ordering, and it separates "always in" from "in half the time". - **Report magnitudes at the penalty you actually deploy**, on standardised predictors, alongside the frequency. That answers "how much" as well as "whether". - **Group correlated predictors before interpreting.** If trim level, body style and colour move together, treat them as a block; asking which of them the path admitted first is asking the data a question it cannot answer. - **Say what the path is for.** It is excellent for seeing how many predictors survive at a given penalty and for spotting instability. It is not evidence of causal importance, and no amount of penalty tuning turns a predictive ordering into a causal one. ## The interview answer in one line Entry order is a conditional, scale-dependent, sample-specific greedy sequence — informative as a diagnostic, unreliable as a ranking, and outright misleading when the predictors are correlated.

  • Is each coefficient guaranteed to move monotonically toward zero as lambda increases?
    No. Individual coefficients can grow, shrink, hit zero, return, and change sign as the active set changes around them, especially among correlated predictors. What is monotone is the total L1 norm of the solution, which is non-increasing as the penalty rises. Quoting a coefficient's sign from one point on the path without checking its neighbourhood is risky.
  • How would you check whether an entry ordering is stable enough to report?
    Refit the whole path on many bootstrap resamples of the rows and record, per predictor, how often it is selected at your chosen penalty and how early it enters. Report selection frequency and a spread rather than a single ranked list; predictors that appear in nearly every resample are worth naming, ones appearing in half are not.
  • A predictor never enters until the very bottom of the grid. Does that mean it does not affect the outcome?
    No. It may be redundant given predictors that entered earlier, weakly correlated with the residual in this particular sample, or genuinely small in effect — the path cannot distinguish these. Check its marginal relationship with the response and its correlation with the early entrants before concluding anything.
  • Why does a coefficient sometimes re-enter the path with the opposite sign?
    Because a correlated partner has just been admitted. The first predictor was standing in for shared signal; once the partner takes that over, the first coefficient can be re-purposed to correct the partner's overshoot, which may require the other sign. It is a property of the joint fit, not a numerical error.

It is like ranking a squad by the order a coach substitutes players in: the second choice depends entirely on who is already on the pitch, and a different opening lineup produces a different order.

saying these in an interview costs you the question

  • Presents entry order as a causal importance ranking
  • Says a late-entering predictor is unrelated to the outcome
  • Assumes the ordering is stable across resamples
  • Claims every coefficient shrinks monotonically as lambda grows
  • Treats a sign flip along the path as a solver bug
  • Compares entry order across predictors on different raw scales

context