skip to content

Why does a voting ensemble of three separately tuned gradient-boosted tree models barely beat the best one?

level: seniorimportance: should knowfreq 44%

answer

  1. Headcount is not the lever
  2. They get the same rows wrong
  3. Averaging cancels only independent error
  4. Correlate the out-of-fold residuals
  5. Check the best-member-per-row ceiling

basics

~20 s

Because their errors are almost the same errors. Members with the same inductive bias, features and objective get the same rows wrong, and combining only cancels mistakes that members make independently. Diversity is the lever, not the number of models.

solid answer

~40 s

Combining helps only where members disagree. Three boosted tree ensembles on the same features with the same objective differ mainly in hyperparameters, so they misclassify the same ambiguous rows, and averaging three copies of one mistake returns the mistake. For squared error this is exact: the ambiguity decomposition shows the averaged ensemble's error equals the average member's error minus the average disagreement, so zero disagreement means zero gain. I would check it directly before adding anything — pairwise correlation of the members' out-of-fold residuals, the rate at which they predict different labels, and the oracle ceiling, the score reached if something always picked the member that got each row right. If that ceiling sits barely above the best single model, no combiner can help and I ship the single model.

go deeper

for a junior

Remember that ensembles pay off when members make different mistakes, and that three near-identical models mostly repeat one model's answer. Know the word diversity and what it refers to here.

for a middle

Explain that averaging cancels only independent error, and name concrete sources of diversity: different families, feature views, encodings and objectives. Be able to say why reseeding one family buys little.

for a senior

Show the measurement, not just the intuition: residual correlations, disagreement rates and the oracle ceiling computed from out-of-fold predictions, and a willingness to ship the single model when the ceiling says the pool has nothing extra.

for a principal

Own the standard that ensembles must clear a measured lift on untouched data before they enter the codebase, and weigh a two-point gain against tripling the training, monitoring and on-call surface.

## The gain comes from disagreement, not from headcount Combining models helps only to the extent that they are wrong about **different rows**. Three gradient-boosted tree ensembles trained on the same features with the same objective, differing only in hyperparameters or seed, tend to be wrong about the *same* rows — the ambiguous ones, the mislabelled ones, the ones from a under-represented segment. Averaging three copies of the same mistake reproduces the mistake. For squared error this is exact rather than a metaphor. The **ambiguity decomposition** says that the squared error of a uniformly averaged ensemble equals the average squared error of its members *minus* the average disagreement among members (each member's average squared deviation from the ensemble's own prediction). Disagreement is a subtracted term: it is the entire benefit. Zero disagreement means the ensemble scores exactly what the average member scores. The same intuition carries to classification without the tidy identity. So the lever is not the number of models. It is how differently they fail. ## Measuring it before you build anything Generate out-of-fold predictions for each candidate model — one per training row, from a model that did not see the row — and then look at: - **Pairwise correlation of the out-of-fold predictions**, or better, of the residuals. Two models correlated at 0.99 on residuals will not help each other. - **Disagreement rate** for classifiers: the fraction of rows where two models predict different labels. - **The oracle ceiling**: the accuracy you would get if, for every row, an oracle picked the member that got it right. If the oracle is barely above the best single model, no combiner can find much, because the correct answer simply is not present in the pool. If the oracle is far above, there is real signal to be routed and a meta-learner may capture some of it. That last check is the one people skip, and it settles the question quickly and cheaply. ## Where diversity actually comes from Ordered roughly by how much they buy you: 1. **Different inductive biases.** A regularised linear model, a nearest-neighbour model and a boosted tree ensemble carve the space in genuinely different ways: a smooth global surface, a local one, an axis-aligned piecewise-constant one. This is the biggest lever, and it is why the classic strong stack mixes families rather than tuning one family harder. 2. **Different views of the data.** Different feature subsets, different encodings of the same categorical column, different aggregation windows, different missing-value handling. 3. **Different objectives.** Optimising absolute error versus squared error, or different class weights, produces models that fail differently on the tails. 4. **Different resamples and seeds.** The cheapest and the weakest source when members are already high-capacity models fitted to convergence — enough to smooth out a little instability, not enough to change what the family cannot see. ## What a meta-learner can and cannot fix A level-1 model can reweight members, recalibrate them, and learn that one member is better in a particular region of the input space. It cannot manufacture signal that no member has. Put three near-identical members under it and the honest outcome is a small gain — sometimes a loss once the fitted weights carry fold noise. The typical symptom is a stack that lands within noise of its best member while costing three times the training and serving effort. ## The judgment this leads to When an ensemble barely beats its best member, do not add a fourth copy. Either add something structurally different and re-measure the oracle ceiling, or accept that the pool has no complementary signal and ship the single model, whose training pipeline, monitoring and latency profile are all a third the size. "Our stack is five models" is not a result. "Our stack beats the best single model by an amount that survives a hold-out the stack has never seen" is.

  • How would you measure ensemble diversity before committing to building the stack?
    Generate out-of-fold predictions for each candidate, then compute pairwise correlations of their residuals and, for classifiers, the fraction of rows where they predict different labels. Then compute the oracle ceiling: the score reached if the right member were chosen per row. That ceiling bounds anything a combiner can achieve, and it costs one pass over predictions you already have.
  • Can a meta-learner rescue an ensemble of highly correlated base models?
    No. It can reweight members, recalibrate their scores and learn that one is stronger in a particular region, but it cannot invent signal that none of them carries. Over three near-identical members the honest result is a small gain, and sometimes a loss once fitted weights absorb fold noise. The remedy is a structurally different member, not a smarter combiner.
  • You only have one strong model family available. What is the cheapest way to add diversity?
    Vary the view rather than the family: different feature subsets, different encodings of the same categorical columns, different aggregation windows, different handling of missing values, different objectives such as absolute versus squared error. Reseeding alone is the weakest option for high-capacity models fitted to convergence. Add one deliberately different simple model as a contrast and measure whether the oracle ceiling moves.

Asking three people who read the same newspaper for their forecast is not three opinions, it is one opinion said three times. The fourth reader adds nothing; a reader of a different newspaper might.

saying these in an interview costs you the question

  • Adds more members of the same family expecting more gain
  • Claims an ensemble always beats its best member
  • Never checks whether members err on the same rows
  • Treats the meta-learner as able to invent missing signal
  • Counts models as the achievement rather than the measured lift

context