skip to content

Why can the sqrt(p) features-per-split default fail on a 400-probe panel with few informative probes?

level: seniorimportance: should knowfreq 46%

answer

  1. count how many candidates 400 features actually gives
  2. how often does a split see a real probe
  3. the trees are diverse but empty
  4. the default assumes signal is spread out
  5. tune it upward, and shrink the noise columns

basics

~20 s

With 400 features, sqrt(p) offers only about 20 candidates per split. If about ten probes carry the signal, most splits see none and are made on noise — the trees end up diverse but uninformative. Raise the features per split.

solid answer

~50 s

The sqrt(p) rule is a default tuned for tables where many features carry some signal, and it breaks when signal is sparse. On a 400-probe gene-expression panel for tumour subtype, sqrt(400) = 20 candidates per node. If about ten probes are genuinely informative, the chance a given split even sees one is roughly 1 - (1 - 10/400)^20, about 40% — so around three splits in five are made on pure noise. Decorrelation is not the problem here; the trees are plenty different, they are just individually weak, and averaging weak-and-wrong is not the same as averaging weak-and-diverse. The fixes are to raise the number of features per split substantially — treating it as the main tuned parameter, not a fixed rule — and to reduce the noise dimension by screening or aggregating probes, with any screening fitted inside the cross-validation folds so it does not leak.

go deeper

for a junior

Know that the number of features offered at each split is a real parameter with a common starting value of about the square root of the feature count, and that it is a starting point rather than a fixed rule.

for a middle

Be able to state the direction of the trade: more candidates per split means stronger but more similar trees, fewer means weaker but more varied ones. Show you can compute how often a split even sees an informative feature.

for a senior

Diagnose the plateau. Show the refit-and-compare experiment that proves the splits were starved, and add the guardrail that any feature screening lives inside the cross-validation folds rather than in front of them.

for a principal

Own the framing that defaults encode an assumption about how signal is spread across columns. Decide when a wide, sparse problem should change model family or data representation instead of consuming the team's tuning budget.

## Where the default comes from Offering `sqrt(p)` features at each split for classification (and about `p/3` for regression) is a rule of thumb, not a theorem. It was calibrated on the kind of tabular problem where most columns carry at least some signal and many are partly redundant with each other. In that regime a small candidate set is cheap insurance: whichever features you draw, several of them are usable, so you lose little tree strength and gain a lot of decorrelation between trees. Sparse-signal, wide-data problems break that assumption. A 400-probe gene-expression panel used to classify tumour subtype may have only a handful of probes that actually separate the subtypes; the rest are biological and technical noise. Now the candidate draw is a lottery you usually lose. ## The arithmetic Draw 20 of 400 features at a node. If 10 are informative, the probability that **none** of the 20 is informative is approximately `(1 - 10/400)^20 = 0.975^20`, which is about 0.60. So roughly 40% of splits have at least one informative probe available and roughly **60% of splits are made on noise alone**. A split made on noise is not harmless. It partitions the data on an irrelevant axis, consumes rows, and pushes the informative probes further down the tree where fewer samples remain to detect them. The tree still terminates with pure-looking leaves, because with 400 columns and few rows there is always some column that separates a small group by chance. That is the mechanism by which wide, sparse data quietly turns tree-growing into noise-fitting. ## Why the symptom is easy to misread The forest will not look broken. It will produce plausible-looking predictions and an accuracy well above chance, because the 40% of splits that do see signal do real work. What you see is a forest that plateaus well below what a simpler model — a regularised linear classifier, say — achieves on the same panel, and whose accuracy is oddly insensitive to depth or tree count. The instinct is to add trees or tune depth. Neither helps, because the limitation is that most splits never had a useful feature to choose from. The diagnostic that separates this from ordinary underfitting is to refit with the number of features per split raised sharply — to 100, to 200, to all 400 — and watch whether validation performance climbs. If it does, you were starving the splits. In sparse-signal regimes the optimum often sits far above `sqrt(p)`, sometimes at `p` itself, which is to say the ideal model here is closer to bagged trees than to a heavily randomised forest. ## Understanding the trade in both directions It helps to hold the direction of the trade firmly: - **Raise the features per split**: each tree gets stronger, because the best available split is more often on offer; the trees also get **more correlated**, because they keep finding the same good features. Variance reduction from averaging drops. - **Lower the features per split**: trees get weaker individually and **less correlated**. Averaging removes more variance, but there is less signal in each tree to average. The optimum depends on how the signal is distributed across the columns. Dense, redundant signal favours a small candidate set. Sparse signal buried in many noise columns favours a large one. This is exactly why the features-per-split parameter is the one worth spending your tuning budget on, and why quoting `sqrt(p)` as if it were a law is a weak answer. ## The other half of the fix: shrink the noise dimension Tuning alone is often not enough when `p` dwarfs `n`. Complementary moves: - **Univariate screening.** Rank probes by a simple association with the outcome and keep the top few dozen. This raises the informative fraction directly. The trap is fatal and common: screening on the full dataset and then cross-validating the forest leaks the outcome into feature selection and produces an optimistic estimate. The screen must be refitted inside each training fold. - **Aggregation.** Collapse correlated probes into pathway or module scores. Fewer, denser features suit the default much better and are usually easier to interpret with a domain expert. - **Reconsider the model family.** With very wide, very sparse data and few rows, a strongly regularised linear model is a serious competitor and should be in the comparison rather than assumed away. ## What to say in the interview State the mechanism first — sqrt(p) gives 20 candidates out of 400, so most splits see no informative probe — then the direction of the fix, then the guardrail on screening. Being able to produce the rough 40% figure from `1 - (1 - 10/400)^20` shows you understand the default as an assumption about how signal is spread, rather than as a setting you inherited.

  • How does raising the number of features per split change the trees?
    Each tree gets stronger, because the genuinely best split is more often among the candidates, and the trees get more correlated, because they repeatedly discover the same strong features. That means less variance reduction from averaging. The right setting balances the two, and with sparse signal it usually sits far above sqrt(p) — sometimes at every feature, which makes the model effectively a bagged forest.
  • You screen the top 30 probes by association with the outcome and then cross-validate the forest. What is wrong?
    The screen has already seen every row, including the ones used as validation in each fold, so the selected probes are chosen partly for how well they happen to separate the validation data. The cross-validated score is then optimistically biased, sometimes dramatically with 400 probes and few samples. Refit the screen inside each training fold so selection is part of what is being validated.
  • Would more trees rescue a forest starved of informative candidate features?
    No. More trees reduce the noise in the average but cannot create signal that the individual trees never had access to. The error curve flattens at a plateau set by how good the trees are, and starved trees set a low plateau. The lever is the number of features offered per split, plus reducing the count of noise columns.

saying these in an interview costs you the question

  • Treats sqrt(p) as a rule rather than a tuned parameter
  • Says more trees will fix an underperforming wide-data forest
  • Thinks raising features per split lowers tree correlation
  • Screens features on the full dataset before cross-validating
  • Assumes a forest always beats a regularised linear model on wide data

context