skip to content

How do extremely randomized trees differ from a random forest, and when do they win?

level: middleimportance: nice to knowfreq 31%

answer

  1. the split search itself is what changes
  2. stop optimising the threshold
  3. usually the whole sample, not resampled copies
  4. worse trees, less alike, much faster
  5. helps most when thresholds are noise

basics

~20 s

Extremely randomized trees draw each candidate feature's cut-point at random instead of searching for the best threshold, and usually train on the full sample. That adds bias per tree, removes more variance, and makes training much faster.

solid answer

~50 s

A random forest still performs an exhaustive threshold search on each candidate feature: for a numeric column it evaluates candidate cut-points and keeps the one with the best split score. Extremely randomized trees skip that search. For each candidate feature they draw one cut-point uniformly at random from that feature's observed range, score only those random splits, and keep the best of them. The original method also grows trees on the whole training set rather than resampled copies. The consequences follow directly: individual trees are worse, so bias per tree rises, but the trees are far less alike, so averaging removes more variance. Training is also markedly cheaper, since scanning every threshold of every candidate feature is the dominant cost of growing a tree. On a 50-feature audio-loudness table with noisy targets, this often matches or beats a comparable forest in a fraction of the time.

go deeper

for a junior

Recall the headline difference: the split threshold is drawn at random rather than searched for, which makes training much faster. You are not expected to argue the bias and variance consequences in detail.

for a middle

Explain the mechanism precisely — one random cut-point per candidate feature, best of those kept, and normally the whole training set rather than resampled copies — and state that per-tree bias rises while correlation between trees falls.

for a senior

Show judgment about when to reach for it: noisy targets, continuous features, tight retraining budgets. Mention that it usually wants more trees to plateau, and that you would settle the choice by comparing both under the same validation scheme.

for a principal

Position it on the randomisation spectrum and reason about the cost curve for the organisation: how much accuracy a faster-to-retrain family is worth when models are rebuilt often and engineering time, not accuracy, is the binding constraint.

## The one change that defines the method Growing a decision tree node normally means a search. For each candidate feature you sort or bin the values, walk the possible thresholds, score each resulting two-way partition with an impurity or variance-reduction criterion, and keep the best threshold found. Extremely randomized trees replace that search with a draw. For each candidate feature, sample a single cut-point uniformly at random between that feature's minimum and maximum in the node, score the resulting split, and take the best among those few random splits — one per candidate feature — rather than the best among all possible thresholds. A second difference matters and is often forgotten: in the original formulation the trees are grown on the **entire** training set, not on resampled copies of it. The reasoning is that once split thresholds are random, you already have plenty of diversity between trees, and using all the data lets each tree see every row. ## What the change buys and what it costs **Bias goes up per tree.** A randomly chosen threshold is rarely the best one, so each tree fits the training signal less precisely. Its decision boundaries are cruder, and it needs more depth to express the same structure. **Correlation between trees goes down.** This is the payoff. Two trees in a random forest, offered the same candidate features at a node, will often pick the same feature and a very similar threshold, because they are both optimising. Two extremely randomized trees will pick different thresholds by construction. Since the variance of an averaged ensemble is floored by the pairwise correlation between its members, lowering that correlation lowers the floor. **Training gets much faster.** Threshold search is the dominant cost of growing a tree; evaluating one random cut-point per candidate feature instead of every candidate threshold removes most of that work, and it removes the need to sort values within each node. **The boundary geometry changes.** Because thresholds are drawn from a continuous range rather than snapped to observed data values, the ensemble's decision surface tends to be smoother, which can help on genuinely continuous relationships and can hurt when the true structure has a sharp threshold at a specific value. ## When to prefer them The honest summary is that the two are close in accuracy on most tabular problems and the choice is usually decided by noise and cost: - **Noisy targets or noisy features.** When the best threshold found by an exhaustive search is largely an artefact of the sample, optimising it hard is a way of fitting noise. Random cut-points refuse to chase it. A 50-feature audio-loudness table with subjective, noisy target values is the shape of problem where this shows up. - **Tight training budgets or frequent retraining.** If you retrain hourly, or you are sweeping other parameters, the speed difference is real and compounding. - **Many continuous features.** Random thresholds are natural on continuous columns; on columns with a few discrete levels, drawing a random cut-point within the range has less room to be different from the optimal one, so the advantage shrinks. When the signal is clean and sharply thresholded — a rule that genuinely triggers at a specific value — the exhaustive search of a plain forest finds that value and the randomised variant only approximates it, so accuracy can favour the forest. If you are unsure, both are cheap to fit; compare them under the same validation scheme rather than arguing from first principles. ## Tuning differences worth knowing Because each tree is weaker, extremely randomized trees typically want **more trees** to reach their plateau than a comparable forest, and they tolerate **deeper** growth, since randomised thresholds already restrain how precisely a tree can fit. The number of features considered per split remains the same kind of parameter with the same direction of effect: more candidates means stronger, more correlated trees. ## Framing it as one family It is cleanest to think of a spectrum of randomisation applied to independently grown trees: perturb the rows, perturb which features may compete at each split, perturb the threshold itself. Each addition trades individual tree quality for lower correlation between trees, and averaging then converts the lower correlation into lower variance. Extremely randomized trees are simply the far end of that spectrum, and knowing where a method sits on it tells you what its failure mode will be — every one of them fails the same way, by being too biased to represent the signal, not by overfitting.

  • Why does randomising the cut-point help when the target is noisy?
    An exhaustive threshold search returns the value that best separates this particular sample, and with a noisy target much of that separation is sampling accident. Optimising it hard bakes the accident into the tree. Drawing the cut-point at random cannot chase that accident, so each tree carries less sample-specific error, and the ensemble average is more stable even though every tree is individually cruder.
  • Do extremely randomized trees usually need more trees than a comparable forest?
    Yes, typically. Each tree is weaker and noisier, so the average needs more members before it settles, and the error curve flattens later. That is affordable because each tree is much cheaper to grow — no threshold scan and no per-node sorting — so the total training cost often still comes out lower than the forest's.
  • When would you expect a plain random forest to win?
    When the true structure has sharp, specific thresholds and the data is clean enough for the split search to find them. An exhaustive search lands on the real cut-point; a random draw only approximates it, and the ensemble spends bias approximating something that was findable. Clean, strongly structured data with relatively few features is the regime where the optimised search earns its cost.

A forest asks each surveyor to find the best place to draw a boundary line; extremely randomized trees ask them to throw the line down at random and keep the least-bad throw. Individually sloppier, collectively far more independent.

saying these in an interview costs you the question

  • Says the features are randomised, not the cut-points
  • Claims extremely randomized trees lower both bias and variance
  • Thinks they are slower because of the extra randomness
  • Assumes they always use resampled training sets
  • Presents them as strictly better than a random forest

context