skip to content

Trees, Forests, and Boosting

You will learn why tree ensembles dominate tabular ML: bagging cuts variance, boosting cuts bias, and a regularized objective made GBDT the default. 'Random forest vs boosting' is a staple screen.

on this pageshow

explore

questions

page 2 of 2

When is a bagged ensemble's out-of-bag error a trustworthy substitute for a hold-out set?

level: seniorimportance: should knowfreq 41%

basics

~20 s

Out-of-bag error is trustworthy when rows are independent and the ensemble is large, because each row is scored only by base models that never saw it. It misleads when rows are grouped, duplicated or time-ordered.

open as a page

Under leaf-wise growth, why is a cap of 63 leaves not equivalent to a depth cap of 6?

level: seniorimportance: should knowfreq 45%

basics

~10 s

Both permit roughly 64 leaves, but they constrain different things. A depth cap of 6 limits every prediction path to six splits. A cap of 63 leaves permits a chain 62 splits deep.

open as a page

Why does gradient boosting still need ordered boosting when its categorical encodings are already leak-free?

level: seniorimportance: should knowfreq 38%

basics

~20 s

Leak-free encodings fix the features, not the gradients. Standard boosting computes a row's residual from an ensemble already fitted on that row, so residuals are optimistic. Ordered boosting scores each row with a model trained only on the rows preceding it.

open as a page

How do you set a decision tree's minimum leaf size on a 900-row cohort with only 36 positive cases?

level: seniorimportance: should knowfreq 36%

basics

~20 s

Size leaves by expected positive cases, not total rows: at a 4% rate, 50-row leaves hold about two positives, so their probabilities are noise. Divide the events you need behind a prediction by the prevalence.

open as a page

Why can the sqrt(p) features-per-split default fail on a 400-probe panel with few informative probes?

level: seniorimportance: should knowfreq 46%

basics

~20 s

With 400 features, sqrt(p) offers only about 20 candidates per split. If about ten probes carry the signal, most splits see none and are made on noise — the trees end up diverse but uninformative. Raise the features per split.

open as a page

Your gradient-boosted trees stop growing far short of the depth cap - which parts of the regularized objective are refusing the splits?

level: seniorimportance: should knowfreq 46%

basics

~20 s

A split happens only if its gain - the drop it produces in the regularized objective - clears the per-leaf cost and the minimum-gain threshold. A large L2 penalty on leaf scores, a high threshold, or a child-curvature floor each veto splits early.

open as a page

Your booster's held-out fold stops training at round 740 of 3,000 — how do you report its accuracy and ship that round count?

level: seniorimportance: should knowfreq 57%

basics

~20 s

Treat them separately. The stopping fold chose a hyperparameter, so its score is optimistic and the reported accuracy must come from data that played no part in stopping. Ship 740 as a fixed round count, valid only for the learning rate and subsampling it was found under.

open as a page

Why does a voting ensemble of three separately tuned gradient-boosted tree models barely beat the best one?

level: seniorimportance: should knowfreq 44%

basics

~20 s

Because their errors are almost the same errors. Members with the same inductive bias, features and objective get the same rows wrong, and combining only cancels mistakes that members make independently. Diversity is the lever, not the number of models.

open as a page

How do you choose among XGBoost, LightGBM and CatBoost for a given table?

level: seniorimportance: should knowfreq 45%

basics

~20 s

Choose on constraints, not on accuracy claims. LightGBM for training throughput on very large, wide tables; CatBoost for high-cardinality categorical columns and fast uniform-depth inference; XGBoost as a strongly regularized general default. Tuned properly, their accuracies usually land within noise.

open as a page

How do you decide whether permutation-based ordered boosting is worth its training cost on a given dataset?

level: principalimportance: should knowfreq 34%

basics

~20 s

Weigh the bias it removes against wall-clock and memory. The prediction shift shrinks as rows accumulate, so on tens of millions of rows it buys little; on small, categorical-heavy data with many rare levels it can move the holdout number.

open as a page

In a bagged classifier, when does averaging predicted probabilities beat majority voting?

level: middleimportance: nice to knowfreq 36%

basics

~20 s

Averaging predicted probabilities beats majority voting when the base models differ in confidence: a few strongly negative models can outweigh a bare majority of barely-positive ones. Voting collapses each model to one ballot and discards that information.

open as a page

Why do some gradient-boosting algorithms grow oblivious (symmetric) trees with one split per level?

level: middleimportance: nice to knowfreq 26%

basics

~20 s

An oblivious tree uses the same feature and threshold at every node of a level, so a depth-d tree holds d tests and 2^d leaves. Scoring becomes d comparisons plus one array lookup, and the constrained shape acts as regularisation.

open as a page

Why can a decision tree's pre-pruning rules reject a split that post-pruning would keep?

level: middleimportance: nice to knowfreq 30%

basics

~10 s

Pre-pruning judges each split on its immediate gain with no lookahead, so a near-worthless split that unlocks a strong one a level below is refused. Post-pruning grows first and judges the pair together.

open as a page

How do extremely randomized trees differ from a random forest, and when do they win?

level: middleimportance: nice to knowfreq 31%

basics

~20 s

Extremely randomized trees draw each candidate feature's cut-point at random instead of searching for the best threshold, and usually train on the full sample. That adds bias per tree, removes more variance, and makes training much faster.

open as a page

How does a decision tree route a row whose split feature is missing, without imputing it?

level: seniorimportance: nice to knowfreq 30%

basics

~20 s

Two tree-native mechanisms exist. CART learns surrogate splits, backup features whose cut best reproduces the primary split, and uses the best available one. Boosted trees instead learn a default direction per node by testing which side gives more gain.

open as a page

Why can binning a price feature into 255 histogram buckets hide a real cut point in its tail?

level: seniorimportance: nice to knowfreq 28%

basics

~10 s

A histogram-based tree can only cut at bin edges. With 255 equal-count bins the top bin already holds about 0.4 percent of rows, so a real threshold at the 99.9th percentile falls inside it.

open as a page

Your booster's training loss reaches zero on 5%-mislabelled data while validation loss rises after round 300 — why?

level: seniorimportance: nice to knowfreq 36%

basics

~20 s

Boosting fits whatever the ensemble still gets wrong, and a mislabelled row is permanently wrong, so later trees carve tiny regions around those rows. Training loss collapses because the noise is being memorised; validation loss turns up because those rounds add nothing real.

open as a page

Why can a decision tree's greedy split search produce a globally suboptimal tree?

level: seniorimportance: nice to knowfreq 38%

basics

~20 s

A tree's splits are each chosen to minimise impurity at that node alone, with no lookahead, so the locally best cut can foreclose a better pair of cuts beneath it. Finding the globally optimal tree is NP-hard.

open as a page

Your model card lists impurity importance as the key drivers - what claims does that ranking support?

level: principalimportance: nice to knowfreq 30%

basics

~20 s

Only a narrow one: for this fitted model, on its training rows, splits on that column removed the largest share of impurity, given the other columns present. It is not a causal effect, not signed, not stable, and not a property of the data.

open as a page

Which features would you place monotone constraints on in a boosted lending model, and what does that cost?

level: principalimportance: nice to knowfreq 34%

basics

~20 s

Constrain only where the direction is a domain law or a written underwriting policy, such as risk never falling as debt-to-income rises. Each constraint buys behaviour you can promise and defend, and usually costs a little held-out accuracy.

open as a page

How do you decide whether a five-model stack is worth shipping under a 40 ms response budget?

level: principalimportance: nice to knowfreq 31%

basics

~20 s

Price both sides. Measure the stack's p99 latency, including fan-out and feature lookups, against the share of the 40 ms the model owns, and measure its gain in the business metric on untouched data. Ship the smallest ensemble clearing both.

open as a page

showing 31–52 of 52