What does a random forest add on top of bagged decision trees, and why does it help?
answer
- extra randomness beyond resampling the rows
- look at what each split is allowed to see
- one dominant feature at every root
- averaging helps least when trees agree
- the correlation term is the floor
basics
~20 sA random forest restricts every split to a random subset of the features. That stops one strong predictor dominating every tree, so the trees are less correlated — and correlation between trees is what limits how much variance averaging can remove.
solid answer
~50 sBagged trees are grown on resampled copies of the same data, so they look alike: a strongly predictive feature sits at or near the root of nearly every tree. Averaging helps least when the things you average agree. The variance of an average of `B` identically distributed predictors with variance `s2` and pairwise correlation `rho` is `rho*s2 + (1-rho)*s2/B` — more trees drive the second term to zero, but `rho*s2` is a floor you cannot average away. A random forest attacks that floor directly: at each node, before searching for the best split, it draws a random subset of `m` of the `p` features and allows the split to use only those. The dominant feature is unavailable at roughly a fraction of nodes, so other structure gets discovered, `rho` drops, and the averaged forest lands below bagging's floor. The cost is that each individual tree is slightly worse.
code
python · 13 linesdef ensemble_variance(rho, var_single, n_trees):
"""Variance of the average of n identically distributed tree
predictions with pairwise correlation rho."""
return rho * var_single + (1 - rho) * var_single / n_trees
sizes = (1, 10, 100, 1000)
for rho in (0.9, 0.6, 0.3):
row = [round(ensemble_variance(rho, 1.0, b), 3) for b in sizes]
print("rho =", rho, "-> variance at B = 1, 10, 100, 1000:", row)
# rho = 0.9 -> [1.0, 0.91, 0.901, 0.9]
# rho = 0.6 -> [1.0, 0.64, 0.604, 0.6]
# rho = 0.3 -> [1.0, 0.37, 0.307, 0.301]go deeper
Recall the one-line difference: a forest randomises the features available at each split, not just the training rows. Be able to say that this makes the trees less alike and that averaging unlike trees is what helps.
Explain the mechanics: a fresh random subset of features per node, the usual sqrt(p) starting point, and the fact that averaging correlated predictors leaves a floor equal to their correlation times a single tree's variance. Say out loud that each tree gets weaker.
Show you can reason about when the trade fails — very few informative features among many noise columns, or a problem whose error is bias rather than variance. Be ready to say what you would change and how you would tell the difference from a learning curve.
Own the framing that randomisation is a variance-budget decision, not a default. Be able to argue when a decorrelated averaging ensemble is the right family for a team at all, given the interpretability, serving cost and tuning effort it commits you to.
## The problem a random forest is solving A fully grown decision tree is a low-bias, high-variance model: it can carve the feature space finely enough to fit almost anything, but shift a handful of training rows and the split at the root can change, and everything below it changes with it. The classic remedy is to grow many trees on perturbed copies of the training data and average their predictions. That is bagging, and it is the base a random forest builds on. Averaging works because independent errors cancel. If you average `B` **independent** estimates each with variance `s2`, the average has variance `s2/B`, which goes to zero as `B` grows. Trees grown on resampled copies of one dataset are not independent, though — they see mostly the same rows and all of the same columns. The honest formula for identically distributed predictors with variance `s2` and **pairwise correlation** `rho` is: ``` Var(average of B trees) = rho*s2 + (1 - rho)*s2/B ``` As `B` grows, the second term vanishes and the whole expression converges to `rho*s2`. That is the ceiling on what averaging can buy you. If your trees are 90% correlated, you can add trees forever and still keep 90% of a single tree's variance. To do better you must make the trees disagree more. ## The mechanism: per-split feature subsampling The random forest change is small and precise. At **every node** of **every tree**, before evaluating candidate splits, draw a random subset of `m` out of the `p` features. The node may split only on those `m`; the other `p - m` are invisible for that decision. The next node draws a fresh subset. Consider a predictive-maintenance table where device age is by far the strongest single predictor of failure. Under plain bagging, device age wins the root split in nearly every tree, so every tree begins with the same partition and the ensemble is close to a single tree with sampling noise on top. Under per-split subsampling, device age is simply absent from a good share of nodes. Those nodes are forced to split on vibration amplitude, duty cycle, ambient temperature — secondary structure that the dominant feature was masking. The trees become genuinely different objects, `rho` falls, and the averaged prediction improves even though no single tree improved. Note the placement of the randomness: the subset is drawn **per split**, not once per tree. Drawing one subset per tree would be much cruder — a tree that never sees the informative features would be pure noise, and a tree that sees them would be an ordinary tree. Per-split drawing gives every tree access to every feature somewhere in its structure while still varying which features compete at each individual decision, which is what decorrelates the deep structure and not just the root. ## The knob and the trade The number of features considered per split, usually written `m`, interpolates between two extremes. Set `m = p` and you are back to bagged trees: maximum tree strength, maximum correlation. Set `m = 1` and each split is made on a single randomly chosen feature: minimum correlation, very weak individual trees. The conventional starting points are `m = sqrt(p)` for classification and `m = p/3` for regression, and `m` is the parameter most worth tuning. The trade is explicit. Lowering `m` **raises the bias of each tree**, because the best available feature is often not in the candidate set, and **lowers the correlation between trees**. The ensemble improves whenever the variance reduction from lower `rho` beats the bias increase. That is usually true on tabular data with many moderately informative, partly redundant features, and it is not automatic — with very few informative features among many noise features, aggressive subsampling starves most splits of signal. ## Why trees are the right base learner Variance-reduction schemes need base learners that are high-variance and low-bias, and that respond strongly to perturbation. Deep unpruned trees are the textbook case: they are unstable by construction, so there is a lot of variance to remove, and their greedy split search is easy to randomise. Averaging many high-bias, stable models (a set of near-identical linear fits, say) buys almost nothing, because their errors are the same error. ## Practical consequences Trees in a forest are grown deep and typically not pruned — pruning would attack variance a second time and buy bias you do not need, since averaging is already handling variance. Regression forests average the leaf values; classification forests normally average the per-class probabilities from each tree rather than taking a hard vote, which gives smoother scores. Because each tree is built without reference to any other tree, training is embarrassingly parallel: 500 trees across 16 cores is the same ensemble as 500 trees built one after another, just faster, and the only shared work is the data itself. That independence is a structural property of averaging ensembles, and it is what makes large forests cheap in wall-clock terms. Finally, be clear about what this does **not** fix. Feature subsampling is a variance device. If your model is underfitting because the signal genuinely is not in your features, or because the true relationship is smooth and extrapolates beyond the training range, more randomness will not rescue it — trees predict a constant outside the range they were trained on, and averaging constants gives a constant.
- Why draw the feature subset at each split rather than once per tree?One subset per tree is far cruder: a tree handed only noise features is useless, and a tree handed the informative ones is just an ordinary tree, so you get a mixture of good and wasted models. Drawing fresh at each node lets every tree reach every feature somewhere while still varying which features compete at each decision, which decorrelates the deep structure and not only the root split.
- Does the decorrelation come for free?No. Each tree is individually weaker, because the best split is often unavailable in the candidate subset, so per-tree bias rises. The ensemble wins only when the variance reduction from lower tree correlation outweighs that bias. On tables with many partly redundant informative features it usually does; when a handful of features carry all the signal, aggressive subsampling can leave most splits with nothing useful to split on.
- Why does this trick work so much better on trees than on linear models?Averaging removes variance, so it pays most when the base learner has a lot of variance and little bias. Deep unpruned trees are unstable by construction and their greedy split search is easy to randomise. Averaging many stable, higher-bias fits mostly reproduces the same systematic error many times over, so there is little for the average to cancel.
Ask twenty forecasters who all read the same single newspaper and you get twenty near-identical forecasts; averaging them barely helps. Force each one to consult a different random handful of sources and the average finally starts to beat any individual.
saying these in an interview costs you the question
- Says a random forest is just bagging applied to trees
- Claims feature subsampling reduces bias rather than variance
- Thinks the feature subset is chosen once per tree, not per split
- Believes enough trees drive ensemble variance to zero regardless of correlation
- Says each tree in a forest is individually stronger than a plain tree