skip to content

Which model family suits 20,000 gene-expression columns and only 180 patients?

level: middleimportance: should knowfreq 46%

answer

  1. columns outnumber rows a hundred to one
  2. many weight vectors fit perfectly
  3. impose structure, do not learn it
  4. shrinkage or sparsity on a linear model
  5. greedy splits chosen from 180 rows

basics

~20 s

With far more columns than rows, use a strongly regularized linear family - ridge, lasso or elastic net, or a linear SVM - possibly after dimension reduction. Flexible families such as deep trees have far too little data per decision.

solid answer

~50 s

When the column count is two orders of magnitude above the row count, almost any flexible model separates the training data perfectly, so the choice is about how much structure I impose rather than how much I can fit. I start with a strongly regularized linear model: ridge if I believe many genes contribute small effects, lasso or elastic net if I want a short signature, or a linear SVM, which is comfortable in high dimensions. Elastic net is usually safer than plain lasso here, because correlated gene groups make lasso's pick arbitrary and lasso can select at most n features when p exceeds n. Deep trees and boosted ensembles fit badly: each split is the best of 20,000 candidates judged on 180 rows, so spurious splits are easy to find. And I quote the score as an interval, not a number.

go deeper

for a junior

Know the shorthand: when columns vastly outnumber rows, constrained and simple beats flexible. Naming a regularized linear model as the first reach, and saying why complex models overfit here, is enough at this level.

for a middle

Explain why the fit is under-determined - 20,000 unknowns against 180 observations admit endless perfect solutions - and how a penalty on the weights picks one. Say which penalty you would choose and what it assumes.

for a senior

Show judgment about what the result means: feature selection that changes between resamples, optimistic internal scores from tiny folds, screening that must sit inside the resampling loop, and external validation before anyone believes the signature.

for a principal

Frame the real decision as whether the study can answer the question at all. Sometimes the right call is to fund more samples or narrow the endpoint rather than model 20,000 columns on 180 patients.

### What p much greater than n means With 20,000 gene-expression columns and 180 patients you have roughly a hundred times more unknowns than observations. A linear model on that data has 20,000 free weights constrained by 180 equations, so there are infinitely many weight vectors that reproduce the training labels exactly. Fitting is not the difficulty; choosing among the perfect fits is. Every useful method in this regime works by adding structure the data itself cannot supply. That reframes the question. You are not asking which family learns the most flexible function. You are asking which family imposes the most sensible constraint, because whatever the model does beyond the training rows is coming from the constraint, not from the 180 patients. ### The families that survive Strongly regularized linear models are the default here. - Ridge adds a penalty proportional to the sum of squared weights. It keeps every column and shrinks all of them toward zero, which is the right bet when the truth is many genes each contributing a small effect. It is usually the stronger pure predictor in that situation. - Lasso adds a penalty proportional to the sum of absolute weights and drives many weights exactly to zero, producing a short signature of genes. Two cautions: when p exceeds n it can select at most n features, and when genes are strongly correlated it picks one of the group essentially arbitrarily, so the chosen list moves around between resamples. - Elastic net mixes the two penalties and is the common compromise: sparse enough to hand a biologist a shortlist, stable enough that correlated genes tend to enter or leave together. - A linear support vector machine is also comfortable in high dimensions, since the fit depends on the margin rather than on estimating 20,000 well-determined coefficients. Reducing the dimension before fitting is the other lever: projecting onto a few dozen components, or screening columns down to a manageable set, then fitting a simple model on what remains. If you screen, the screening has to be redone inside every resampling fold, or the score you report has already seen the held-out patients. ### Why flexible families struggle A deep tree or a boosted ensemble is a poor match, and the reason is the split search rather than trees being inherently weak. Each split is chosen as the best of 20,000 candidate columns evaluated on a handful of rows. With that many candidates and that little data, some column separates this particular sample by luck, and the greedy search is designed to find exactly that column. The result is an optimistic training fit built on splits that will not reproduce. Adding more trees does not recover information the data does not contain; it mainly supplies more ways to fit the noise. Distance-based methods suffer differently: with thousands of mostly irrelevant columns, distances between patients become dominated by noise and neighbourhoods stop being informative. ### The part candidates forget In this regime the evaluation is harder than the modelling. A held-out fold contains a few dozen patients, so a handful of cases flipping moves the reported score by several points; any number you quote should carry a wide interval. Feature selection is unstable in a way that matters scientifically: rerun the fit on a different 90% of the patients and a lasso signature can change substantially while predictive performance barely moves, because correlated genes are interchangeable. That is a property of the data, not a bug in the method, but it does mean the selected gene list should be described as one of several equally good lists rather than as the answer. External validation on an independent cohort is worth more than any internal estimate. ### The regime, not the column count None of this is a rule about gene-expression data or about 20,000 columns. It is a rule about the ratio. Give the same 20,000 columns two million rows and the advice inverts: there is now ample data behind every split, deep interactions are estimable, and the ordinary tabular default returns. The question to ask on any new dataset is how much data stands behind each decision the model has to make, and to choose the family whose number of decisions the data can actually support.

  • Why is a boosted tree ensemble a risky choice when 20,000 candidate columns are scanned on 180 rows?
    Each split is the winner of 20,000 comparisons evaluated on a handful of rows, so the winner is often the column that got lucky rather than the one that matters. The greedy search is built to find exactly that column, and the ensemble compounds the optimism. Boosting cannot recover information the data does not contain; with this ratio it mostly supplies more ways to fit the noise.
  • Would you prefer lasso or ridge for a gene-expression classifier?
    It depends on what the model is for. Lasso and elastic net produce a short gene signature, which is what a biologist can follow up, but with strongly correlated gene groups the choice of representative is unstable across resamples. Ridge keeps every gene with a small weight and is usually the better pure predictor when the true signal is diffuse. Elastic net is the standard compromise between the two.
  • Does the same reasoning apply if you get 20,000 columns and two million rows?
    No - that is a different regime. Once the row count vastly exceeds the column count, the data can support a flexible model, and the ordinary tabular default returns: a boosted ensemble can estimate deep interactions because every split still has plenty of rows behind it. The advice is about the ratio and how much data stands behind each decision, not about the raw column count.

It is like fitting a curve through three points with a twenty-term polynomial: the data does not pick the curve, your choice of penalty does.

saying these in an interview costs you the question

  • Says more features always means more information
  • Reaches for a boosted ensemble because it wins on tabular data
  • Thinks a perfect training fit means the signal is real
  • Screens columns on all rows, then cross-validates the survivors
  • Reports a single accuracy number from 180 patients as fact
  • Presents an unstable feature list as the definitive gene set

context