Trees, Forests, and Boosting
You will learn why tree ensembles dominate tabular ML: bagging cuts variance, boosting cuts bias, and a regularized objective made GBDT the default. 'Random forest vs boosting' is a staple screen.
on this pageshowhide
explore
- Decision Trees12 questions
- Split Criteria4 questions
- Pruning and Stopping4 questions
- Axis-Aligned Boundaries4 questions
- Averaging Ensembles12 questions
- Bootstrap Aggregation4 questions
- Random Forests4 questions
- Stacking and Blending4 questions
- Sequential Weak Learners10 questions
- AdaBoost3 questions
- Gradient Boosting3 questions
- Shrinkage and Early Stopping4 questions
- Modern GBDT Algorithms15 questions
- Regularized Boosting Objective3 questions
- Leaf-Wise Histogram Growth4 questions
- Ordered Target Statistics4 questions
- Impurity-Based Importance4 questions
- Tabular Model Choice3 questions
questions
page 2 of 2When is a bagged ensemble's out-of-bag error a trustworthy substitute for a hold-out set?
basics
~20 sOut-of-bag error is trustworthy when rows are independent and the ensemble is large, because each row is scored only by base models that never saw it. It misleads when rows are grouped, duplicated or time-ordered.
Under leaf-wise growth, why is a cap of 63 leaves not equivalent to a depth cap of 6?
basics
~10 sBoth permit roughly 64 leaves, but they constrain different things. A depth cap of 6 limits every prediction path to six splits. A cap of 63 leaves permits a chain 62 splits deep.
Why does gradient boosting still need ordered boosting when its categorical encodings are already leak-free?
basics
~20 sLeak-free encodings fix the features, not the gradients. Standard boosting computes a row's residual from an ensemble already fitted on that row, so residuals are optimistic. Ordered boosting scores each row with a model trained only on the rows preceding it.
How do you set a decision tree's minimum leaf size on a 900-row cohort with only 36 positive cases?
basics
~20 sSize leaves by expected positive cases, not total rows: at a 4% rate, 50-row leaves hold about two positives, so their probabilities are noise. Divide the events you need behind a prediction by the prevalence.
Why can the sqrt(p) features-per-split default fail on a 400-probe panel with few informative probes?
basics
~20 sWith 400 features, sqrt(p) offers only about 20 candidates per split. If about ten probes carry the signal, most splits see none and are made on noise — the trees end up diverse but uninformative. Raise the features per split.
Your gradient-boosted trees stop growing far short of the depth cap - which parts of the regularized objective are refusing the splits?
basics
~20 sA split happens only if its gain - the drop it produces in the regularized objective - clears the per-leaf cost and the minimum-gain threshold. A large L2 penalty on leaf scores, a high threshold, or a child-curvature floor each veto splits early.
Your booster's held-out fold stops training at round 740 of 3,000 — how do you report its accuracy and ship that round count?
basics
~20 sTreat them separately. The stopping fold chose a hyperparameter, so its score is optimistic and the reported accuracy must come from data that played no part in stopping. Ship 740 as a fixed round count, valid only for the learning rate and subsampling it was found under.
Why does a voting ensemble of three separately tuned gradient-boosted tree models barely beat the best one?
basics
~20 sBecause their errors are almost the same errors. Members with the same inductive bias, features and objective get the same rows wrong, and combining only cancels mistakes that members make independently. Diversity is the lever, not the number of models.
How do you choose among XGBoost, LightGBM and CatBoost for a given table?
basics
~20 sChoose on constraints, not on accuracy claims. LightGBM for training throughput on very large, wide tables; CatBoost for high-cardinality categorical columns and fast uniform-depth inference; XGBoost as a strongly regularized general default. Tuned properly, their accuracies usually land within noise.
How do you decide whether permutation-based ordered boosting is worth its training cost on a given dataset?
basics
~20 sWeigh the bias it removes against wall-clock and memory. The prediction shift shrinks as rows accumulate, so on tens of millions of rows it buys little; on small, categorical-heavy data with many rare levels it can move the holdout number.
In a bagged classifier, when does averaging predicted probabilities beat majority voting?
basics
~20 sAveraging predicted probabilities beats majority voting when the base models differ in confidence: a few strongly negative models can outweigh a bare majority of barely-positive ones. Voting collapses each model to one ballot and discards that information.
Why do some gradient-boosting algorithms grow oblivious (symmetric) trees with one split per level?
basics
~20 sAn oblivious tree uses the same feature and threshold at every node of a level, so a depth-d tree holds d tests and 2^d leaves. Scoring becomes d comparisons plus one array lookup, and the constrained shape acts as regularisation.
Why can a decision tree's pre-pruning rules reject a split that post-pruning would keep?
basics
~10 sPre-pruning judges each split on its immediate gain with no lookahead, so a near-worthless split that unlocks a strong one a level below is refused. Post-pruning grows first and judges the pair together.
How do extremely randomized trees differ from a random forest, and when do they win?
basics
~20 sExtremely randomized trees draw each candidate feature's cut-point at random instead of searching for the best threshold, and usually train on the full sample. That adds bias per tree, removes more variance, and makes training much faster.
How does a decision tree route a row whose split feature is missing, without imputing it?
basics
~20 sTwo tree-native mechanisms exist. CART learns surrogate splits, backup features whose cut best reproduces the primary split, and uses the best available one. Boosted trees instead learn a default direction per node by testing which side gives more gain.
Why is a gradient boosting leaf's value found by line search rather than averaging?
basics
~20 sThe tree only chooses which rows move together; the leaf value is the constant that actually minimises the chosen loss for those rows given the current model. Under squared error that is their mean, under absolute error their median.
Why can binning a price feature into 255 histogram buckets hide a real cut point in its tail?
basics
~10 sA histogram-based tree can only cut at bin edges. With 255 equal-count bins the top bin already holds about 0.4 percent of rows, so a real threshold at the 99.9th percentile falls inside it.
Your booster's training loss reaches zero on 5%-mislabelled data while validation loss rises after round 300 — why?
basics
~20 sBoosting fits whatever the ensemble still gets wrong, and a mislabelled row is permanently wrong, so later trees carve tiny regions around those rows. Training loss collapses because the noise is being memorised; validation loss turns up because those rounds add nothing real.
Why can a decision tree's greedy split search produce a globally suboptimal tree?
basics
~20 sA tree's splits are each chosen to minimise impurity at that node alone, with no lookahead, so the locally best cut can foreclose a better pair of cuts beneath it. Finding the globally optimal tree is NP-hard.
Your model card lists impurity importance as the key drivers - what claims does that ranking support?
basics
~20 sOnly a narrow one: for this fitted model, on its training rows, splits on that column removed the largest share of impurity, given the other columns present. It is not a causal effect, not signed, not stable, and not a property of the data.
Which features would you place monotone constraints on in a boosted lending model, and what does that cost?
basics
~20 sConstrain only where the direction is a domain law or a written underwriting policy, such as risk never falling as debt-to-income rises. Each constraint buys behaviour you can promise and defend, and usually costs a little held-out accuracy.
How do you decide whether a five-model stack is worth shipping under a 40 ms response budget?
basics
~20 sPrice both sides. Measure the stack's p99 latency, including fan-out and feature lookups, against the share of the 40 ms the model owns, and measure its gain in the business metric on untouched data. Ship the smallest ensemble clearing both.
showing 31–52 of 52