When you average 50 differently-seeded high-variance fits, what caps the variance reduction?
answer
- errors must disagree to cancel
- variance of a mean of correlated terms
- one term ignores M entirely
- rho*sigma^2 plus (1-rho)*sigma^2/M
- shared bias survives the average
basics
~20 sCorrelation between their errors. Averaging M fits whose errors have pairwise correlation rho leaves rho times the single-fit variance no matter how large M grows; only the independent share, (1 - rho)/M of the variance, is averaged away.
solid answer
~40 sAveraging helps only to the extent the fits make *different* mistakes. If each fit has error variance `sigma^2` and any two fits' errors correlate at `rho`, the average of M of them has variance `rho*sigma^2 + (1 - rho)*sigma^2/M`. The second term vanishes as you add members; the first does not. So the ensemble floor is `rho*sigma^2`, and with rho around 0.6 most of the available gain has arrived by five or ten fits — the fiftieth buys almost nothing. Two consequences follow. Effort is better spent decorrelating the members (different initialisations, different hyperparameters, a genuinely different model family) than adding more near-identical ones. And averaging attacks variance only: any error the fits share — a missing feature, a mis-specified functional form — sits in every member and survives the average untouched.
code
python · 17 linesimport random, statistics
random.seed(0)
M, rho, sigma = 50, 0.6, 1.0
def member_errors():
shared = random.gauss(0, sigma * rho ** 0.5)
return [shared + random.gauss(0, sigma * (1 - rho) ** 0.5) for _ in range(M)]
single, averaged = [], []
for _ in range(20000):
errs = member_errors()
single.append(errs[0])
averaged.append(sum(errs) / M)
print(round(statistics.pvariance(single), 3)) # ~1.00 = sigma^2
print(round(statistics.pvariance(averaged), 3)) # ~0.61 = rho + (1-rho)/Mgo deeper
Be ready to state the intuition in one line: averaging helps because the fits' mistakes point in different directions and partly cancel. Know that it removes the noisy part of the error and not a systematic one.
Explain the variance of an average of correlated errors and name both terms — the share that shrinks with the number of members and the share fixed by their correlation. Show why the gain saturates after a handful of fits.
Demonstrate the operating judgment: measure validation error against ensemble size, spend effort decorrelating members rather than cloning them, and weigh the latency and monitoring cost of many models against a small error gain.
Own the call between an ensemble and one well-specified model. Averaging cannot repair a bias every member shares, and each extra member is permanent training, serving and monitoring cost — argue when the ensemble is the wrong purchase.
## What averaging is for Some learning procedures are unstable: run one twice on the same data with a different random seed — a different initialisation, a different tie-break, a different order of presentation — and you get two noticeably different models. Each fit's error splits into two parts. One part is the same every run: what the model class, the feature set and the loss simply cannot express. The other changes from run to run: the seed-dependent, sample-dependent wobble. Averaging the predictions of many such fits attacks the second part and leaves the first exactly where it was. That is the whole idea, and it is why averaging is counted as a regularizer even though nothing was added to the loss function. ## The arithmetic that answers the question Let member j make error `e_j` on a given input, with each `e_j` having variance `sigma^2`, and let any two members' errors have correlation `rho`. The ensemble predicts the mean of the members, so its error is the mean of the `e_j`. The variance of a mean of M identically distributed, equally correlated terms is `Var(mean) = rho*sigma^2 + (1 - rho)*sigma^2 / M` Read the two terms separately. - The second term is the **independent** share of the error. It falls as 1/M and goes to zero. This is the part averaging actually removes. - The first term is the **shared** share. It does not contain M at all. No number of members touches it. So the floor is `rho*sigma^2`. If the members were perfectly independent (rho = 0) the variance really would fall as `sigma^2/M`; that is the textbook picture and it never holds in practice. If the members were perfectly correlated (rho = 1) averaging would achieve nothing, which is exactly what happens when you average a deterministic algorithm's output with itself. ## Why the returns saturate so fast The reducible part shrinks by a factor of `1 - 1/M`. At M = 2 you have half of it, at M = 5 four fifths, at M = 10 nine tenths, at M = 50 ninety-eight percent. The first few members do nearly all the work. This matters commercially: members two through five are usually worth their training and serving cost, and members twenty through fifty almost never are. The honest way to fix the number is empirical — plot validation error against ensemble size and stop where the curve flattens. ## Correlation is never small Fifty fits of one algorithm, on one dataset, with one feature set, differing only by seed, agree with each other on most inputs. They were shaped by the same rows. Their errors are correlated precisely because they share everything except the noise source you varied. That is why the practical lever is not M but rho: change the initialisation, change the hyperparameters, change the representation, or use a different model family altogether. A penalised linear fit and a tree-based fit disagree far more than two runs of the same procedure, so their errors correlate less and the average gains more. There is a tradeoff hiding in that advice. Making members more different usually makes each one individually worse, so you trade a smaller rho for a larger `sigma^2`. The ensemble improves only if the correlation falls faster than the individual quality does. Deliberately crippling members to decorrelate them is a real failure mode. ## Bias is untouched This is the point candidates most often miss. If every member omits an important feature, or forces a straight line through a curved relationship, every member is wrong in the same direction and the average is wrong by the same amount. Averaging is variance control, not a repair for a mis-specified model. When validation error stays high and stubbornly flat as you add members, that is the signature of shared bias, and the fix is a better model or better features, not a bigger ensemble. ## Parameter sharing is the same idea from the other end Averaging M fits pools information across models. Tying parameters together — forcing one coefficient block to serve several related sub-populations — pools information across data. Both cut variance by making the effective estimate depend on more evidence, and both risk a systematic error when the things being pooled are not really alike. ## What to say in an interview State the formula, name both terms, say which one M kills and which one it does not, and give the practical consequence: decorrelate rather than multiply, measure the saturation curve, and never expect an ensemble to fix a bias every member shares.
- Does averaging reduce bias as well as variance?No. Bias is the part of the error every member shares, and averaging a quantity that everyone gets wrong in the same direction leaves it exactly as wrong. If all fifty fits omit an important feature or force a linear shape on a curved relationship, the average inherits that error in full. Averaging cancels only the idiosyncratic, run-to-run part.
- How would you drive the correlation between the members' errors down?Change something that actually differs: the initialisation and seed, the hyperparameter settings, the feature representation, or the model family itself. A penalised linear fit averaged with a tree-based fit correlates far less than fifty runs of one procedure on one dataset. Watch the tradeoff — a decorrelated member is often individually weaker, so you are buying a smaller rho with a larger per-member variance.
- How many members before the extra compute stops paying?The reducible variance falls by a factor of 1 - 1/M, so five members capture about eighty percent of it and ten about ninety. Past that you pay linearly in training and serving for a few tenths of a percent. Fix the number empirically: plot validation error against ensemble size and stop where the curve flattens, then weigh the remaining gain against latency and monitoring cost.
Ask fifty people to guess the number of beans in a jar and the average is usually good — but only because their guesses miss in different directions. Fifty people who all applied the same wrong rule of thumb average to the same wrong number.
saying these in an interview costs you the question
- Says averaging reduces bias as well as variance
- Assumes members are independent, so variance falls as 1/M
- Believes adding more members always keeps helping
- Averages fifty near-identical fits and expects a large gain
- Cannot say why averaging a deterministic fit with itself does nothing