What is bagging, and why does averaging bootstrap-trained models cut variance but not bias?
answer
- same learner, many resampled training sets
- errors that differ cancel out
- shared systematic error does not
- variance term shrinks, bias term stays
- unstable deep trees gain the most
basics
~20 sBagging trains many copies of one model on bootstrap resamples of the training rows, then averages their predictions. Averaging cancels the errors that differ from copy to copy, cutting variance; the systematic error every copy shares survives, so bias stays.
solid answer
~50 sBagging, short for bootstrap aggregating, does three things: draw a resample of n rows with replacement from the n training rows, train the same learner on it, repeat B times, then combine the B predictions by averaging (regression) or voting (classification). The resamples differ, so each model makes somewhat different mistakes, and those idiosyncratic mistakes partly cancel in the average — that is a pure variance reduction. What does not cancel is anything all the models get wrong in the same direction, because averaging changes the spread of predictions around their mean, not the mean itself. So bagging pays off for high-variance, low-bias base learners such as deep unpruned trees, and pays almost nothing for stable ones such as a well-conditioned linear model or a depth-2 stump. The resamples overlap heavily, so the models stay correlated and the variance does not go to zero however many you add.
go deeper
Be ready to state the loop in one breath: resample rows with replacement, train the same learner, repeat, average or vote. Then say the one-line reason it helps — the individual models' random mistakes cancel.
Expect to be pushed into the bias-variance decomposition. Explain why the average of B models has a smaller spread but the same centre, and why overlapping resamples keep the models correlated so the variance does not fall to zero.
Show you can predict when bagging will pay before running it. Diagnose the base learner's stability first, and be able to say why a bagged stump ensemble or a bagged linear model was never going to move the number.
Own the cost side of the call. B times the training and serving compute, plus the loss of a single inspectable model, has to be justified against the measured variance gain — and against whether the real error is bias-shaped, in which case better features beat any ensemble.
## The procedure Bagging (bootstrap aggregating) is a recipe, not a model. Given a training table of `n` rows: 1. Draw `n` rows **with replacement** from the training table. Some rows appear twice or three times; roughly a third do not appear at all. This is one bootstrap resample, and it has the same size `n` as the original. 2. Train the base learner on that resample, using the same settings every time. 3. Repeat B times to get B fitted models. Nothing links them — no model sees another's output — so training is embarrassingly parallel. 4. Predict by combining: the mean of the B numeric predictions for regression, a vote or an average of predicted class probabilities for classification. The base learner is the same algorithm each time. The only thing that changes is which rows it saw and how often. ## Why averaging attacks variance Expected prediction error at a point decomposes into three pieces: `error = bias^2 + variance + irreducible noise`. **Bias** is how far the average prediction (averaged over hypothetical retrainings on fresh data) sits from the truth. **Variance** is how much the prediction jumps around when you retrain on different data. **Irreducible noise** is what no model can remove. Take B models whose predictions each have variance `s^2`. If they were statistically independent, the variance of their mean would be `s^2 / B` — that is the ordinary result that averaging B independent numbers shrinks the spread by a factor of B. They are not independent: two bootstrap resamples of the same table share most of their rows, so the fitted models resemble each other. With average pairwise correlation `rho` between their predictions, the variance of the average is `rho * s^2 + (1 - rho) * s^2 / B` As B grows the second term vanishes and the first stays. So the number of base models is a compute decision that buys you the full available reduction, and `rho * s^2` is the floor. How much you actually gain is set by how different the models are from one another, which in plain bagging comes only from row resampling. ## Why bias survives Each base model is trained on data drawn from the same distribution, so each one has approximately the same expected prediction. The average of B quantities that all have the same mean has that mean. Concretely: if every tree systematically under-predicts the score of very acidic wines because the feature encoding cannot express that relationship, then 200 such trees under-predict it in exactly the same way, and their average does too. Averaging is a smoothing operation over the *spread* of predictions; it has no mechanism for moving their centre. There is one small correction in the other direction. Each base model sees only about 63% of the distinct rows, so it is fit on slightly less information than a model trained on the whole table. That makes each base model marginally worse than a single full-data fit, so a bagged ensemble's bias can be a touch **higher**, not lower. On a few thousand rows the effect is negligible next to the variance gain; on a couple of hundred rows it can be enough to wipe the gain out. ## Where it pays and where it does not Take a 3,000-row wine-quality sensory-panel table — chemical measurements per bottle, a panel score as the target. A single unpruned decision tree grown on it is a textbook high-variance model: swap out a handful of rows and the top split can change, sending every downstream split somewhere else. Bag 200 such trees and the test error drops materially, because the wild disagreements between individual trees average out while their shared shape survives. Now bag 200 depth-2 stumps on the same table. A stump can express only two splits' worth of structure; it underfits badly and it barely changes when you resample rows. There is almost no variance to remove, and the shared underfitting is exactly what averaging cannot touch, so the ensemble lands roughly where a single stump did. This is the clearest practical statement of the rule: **bagging needs a base learner that is unstable.** The same reasoning explains why bagging a well-conditioned linear regression with plenty of rows, or a nearest-neighbour rule with a large neighbourhood, tends to be a waste of compute. ## What it does not fix Bagging does not fix a missing feature, a mis-specified target, systematically mislabelled rows, or a shifted deployment distribution — all of those are bias-shaped problems that every base model inherits. It also costs you B times the training compute and B times the prediction compute, and it takes a single readable tree and turns it into an object you can no longer trace by eye.
- You bagged 200 depth-2 stumps and the error barely moved — why?A depth-2 stump is high-bias and low-variance: it underfits, and resampling rows hardly changes what it learns. Bagging only removes the part of the error that differs between base models, and here that part was already tiny. The shared underfitting is untouched. The fix is more capacity per base model — grow the trees deep — so there is variance worth averaging away.
- Why can a bagged ensemble have slightly worse bias than a single model trained on all the data?Each base model is fit on one bootstrap resample, which contains only about 63% of the distinct rows. Every copy therefore learns from a little less information than a full-data fit, and averaging cannot recover what none of them saw. On a few thousand rows this is negligible; on a very small table it can cancel out the variance gain entirely.
- Which base learners are poor candidates for bagging?Stable ones. A linear or logistic model with many more rows than features, a heavily regularized model, or a nearest-neighbour rule with a large neighbourhood all produce nearly the same fit on any resample. The base models end up almost identical, their average is almost the same model, and you have paid B times the compute for nothing.
Ask 200 tasters who each sampled a different random subset of the same bottles. Their personal quirks cancel in the average score, but a preference they all share — everyone dislikes tannin — survives no matter how many tasters you add.
saying these in an interview costs you the question
- Says bagging reduces bias as well as variance
- Confuses resampling rows with sampling feature columns
- Thinks the resamples are disjoint partitions of the training data
- Believes the base models must be different algorithms
- Claims averaging helps even when the base models are near-identical