skip to content

Why does elastic net beat lasso on an 8,000-gene panel of co-expressed modules with 200 patients?

level: seniorimportance: should knowfreq 36%

answer

  1. far more columns than rows
  2. the path saturates somewhere
  3. correlated genes travel as a module
  4. at most n non-zero coefficients
  5. strict convexity buys a unique solution

basics

~20 s

A pure L1 fit can place at most 200 non-zero coefficients when there are only 200 patients, and inside a co-expressed module it keeps roughly one gene chosen by sampling noise. Adding a squared-penalty share lifts that cap and keeps correlated genes together.

solid answer

~50 s

Two separate failures bite here. First, a counting limit: with more predictors than rows, an L1 solution cannot have more than `n` non-zero coefficients, so at most 200 of the 8,000 genes can ever be selected — if the truth is spread over 500 genes the model cannot represent it. Second, a selection limit: inside a module of genes correlated at 0.9-plus, keeping one gene costs the same penalty as splitting weight across several, so the fit keeps about one, chosen by sampling noise. Refit on a resample and a different gene comes back. Adding an L2 share fixes both because it makes the objective strictly convex: the solution becomes unique, the `n` cap no longer binds, and the grouping effect means that as two predictors' correlation approaches 1 their coefficients converge, so modules enter and leave together. The cost is a larger, less parsimonious selected set.

go deeper

for a junior

Recall the headline: with far more predictors than rows, an L1 fit is capped at as many non-zero coefficients as there are rows, and it keeps about one predictor from each correlated group.

for a middle

Explain why both limits exist — the active columns must be independent and there are only so many rows, and near-duplicate columns cost the penalty the same whether weight sits on one or is split.

for a senior

Diagnose it in front of the interviewer: show how you would detect the instability across resamples, and describe what changing the mixing ratio does to the selected set rather than just to the score.

for a principal

Frame the deliverable trade explicitly. A reproducible module-level list and a short cheap assay panel are different products, and the mixing ratio is where you choose between them; be prepared to defend that choice to a domain team.

## The setting An expression panel with 8,000 genes measured on 200 patients is the canonical `p >> n` problem: forty times more columns than rows. Genes do not vary independently — they come in co-expressed modules, groups of genes whose measurements rise and fall together because they sit in the same biological pathway. Within a module, pairwise correlations of 0.9 and above are ordinary. Fit a pure L1-penalised model here and you get something that looks impressive — a handful of genes with non-zero coefficients and decent cross-validated performance — but which frustrates the biologists on two separate counts. ## Problem one: the hard cap at n With more predictors than rows, an L1 solution can place **at most `n` non-zero coefficients** — here, at most 200 out of 8,000, no matter how the penalty strength is set. The reason is structural. At the optimum, the columns carrying non-zero coefficients must be linearly independent (otherwise weight could be shuffled between them at no cost, and the optimum would not be unique in the way the optimality conditions require). Those columns live in a space of dimension `n` — you only have 200 rows — so no more than 200 of them can be independent. The path saturates there. If the true biological signal is spread thinly across 400 or 600 genes, the L1 fit cannot represent it. It is not that the model chose 200; it is that 200 is the ceiling. ## Problem two: arbitrary choice inside a module Within a co-expressed module of, say, thirty genes carrying nearly the same information, the L1 penalty is close to indifferent about which one it keeps. Keeping one gene at weight `b` costs the same penalty as splitting `b` across two genes, and the loss barely distinguishes them — so the optimiser keeps roughly one, chosen by whichever gene happened to correlate marginally better with the outcome in this particular sample of 200 patients. Refit on a bootstrap resample and a different member of the module comes back. To a biologist who wants to know *which pathway matters*, a list that changes every time you rerun it is not a finding. ## What the L2 share changes Adding a squared-penalty share — the elastic net — changes both problems at once, and for the same underlying reason: the squared term makes the objective **strictly convex**. - **The cap disappears.** With a strictly convex objective the solution is unique and the `n` ceiling no longer binds; an elastic net can select more predictors than there are rows. One clean way to see it: an elastic net fit is equivalent to an L1 fit on an augmented design with `n + p` rows, and that augmented design has plenty of rows to spare. - **Correlated predictors are held together — the grouping effect.** For standardised predictors, the difference between two elastic-net coefficients is bounded by a quantity that shrinks as the correlation between the two predictors grows, and the bound is tighter the larger the L2 share. Stated at the limit: as the correlation between two predictors approaches 1, their elastic-net coefficients converge to the same value. Whole modules therefore enter or leave the model together, and the selected set stops flipping between refits. That is the answer the biologists actually wanted: not "gene 4,317 matters" but "this module matters", with a reproducible list. ## What it costs Nothing here is free. - **Parsimony.** Keeping modules intact means more surviving genes. A 30-gene module that the L1 fit compressed into one coefficient may now contribute thirty. If the deliverable is an assay panel whose per-gene cost is real, that matters, and you may deliberately push the mixing ratio back toward pure L1 and accept the instability. - **Interpretation.** Coefficients spread across a correlated module are individually small and individually meaningless; read them as a group, and never as unbiased effect sizes — every penalised coefficient is shrunk toward zero by construction. - **Tuning.** You now have two dials rather than one, and the mixing ratio has to be chosen on evidence rather than habit. ## How to talk about it in an interview The strong answer separates the two failure modes rather than blurring them into "lasso is unstable". One is a *counting* limit — at most `n` non-zeros — and it bites whenever the truth is denser than the sample is tall. The other is a *selection* limit — arbitrary choice among near-duplicates — and it bites whenever predictors travel in groups. The L2 share fixes both because strict convexity is what both were missing, and the price is a larger, less parsimonious model.

  • Why can an L1 fit never select more than n predictors when p exceeds n?
    Because the columns carrying non-zero coefficients have to be linearly independent at the optimum, and those columns live in a space whose dimension is the number of rows. With 200 rows, no more than 200 columns can be independent, so the active set saturates at 200 however the strength is set. It is a structural ceiling, not a tuning choice.
  • State the grouping effect precisely.
    For standardised predictors, the gap between two elastic-net coefficients is bounded by a quantity that shrinks as the correlation between those two predictors grows, and the bound tightens as the L2 share grows. At the limit, as the correlation approaches 1 the two coefficients converge to the same value. A pure L1 penalty has no such bound, which is why it can hand one predictor everything and the other nothing.
  • What do you give up by moving from a pure L1 fit to an elastic net here?
    Parsimony, mostly. Keeping whole modules means far more surviving genes, which costs interpretability and, if the deliverable is an assay panel, real money per gene measured. You also gain a second hyperparameter to tune. If a short panel matters more than reproducibility, pushing the ratio back toward pure L1 and accepting the instability can be the right call — but say so explicitly.
  • The coefficients inside a selected module are all small. What does that mean biologically?
    Usually that the module shares one signal and the fit has split the weight across its members, not that each gene is individually weak. Read penalised coefficients as a group, and never as unbiased effect sizes: every one of them is shrunk toward zero by construction, so their magnitudes understate the underlying association.

saying these in an interview costs you the question

  • Says a lasso can select any number of genes if you lower the penalty enough
  • Treats an unstable selected set as randomness rather than a property of the penalty
  • Claims the squared term makes the objective non-convex
  • Reads a shrunken coefficient as an unbiased effect size
  • Says elastic net loses the ability to zero coefficients entirely

context