Why must a stacking ensemble's meta-learner be trained on out-of-fold base-model predictions?
answer
- Two levels trained on different information
- The meta-features are themselves predictions
- In-sample predictions flatter the base models
- Score rows only from excluding folds
basics
~20 sBase models memorise their own training rows, so in-sample predictions look far better than they ever will on new data. Out-of-fold predictions, where each row is scored by a model that never saw it, give the meta-learner honest inputs.
solid answer
~50 sA stack has two levels: several base models at level 0, and a small meta-learner at level 1 whose input features are the base models' predictions. If those predictions are made on rows the base models trained on, they are near-copies of the target, so the meta-learner learns to trust whichever model memorised hardest rather than whichever generalises best. The construction that avoids it: split the training data into K folds, train each base model on K-1 folds and predict the held-out fold, repeat until every training row has one prediction per base model from a model that excluded it, then fit the meta-learner on that matrix. Afterwards refit the base models on all the data for serving, and keep the meta-learner as trained. The meta-learner itself should be small and regularised — often a linear model with non-negative weights, so the stack is essentially a learned weighted average.
go deeper
Be able to say what the two levels are: base models that see the raw features, and one small model on top whose features are the base models' predictions. Know the phrase out-of-fold and that it means a row was scored by a model that did not train on it.
Explain the fold loop step by step and why in-sample predictions look almost like the target. Be ready to say what happens to the fold models afterwards and what gets deployed.
Show the diagnosis: a stack that validates far above every base model is the signature of leaked meta-features. Talk about how you would prove it, and about keeping an untouched hold-out to score the whole stack.
Own the call on whether the stack is worth its complexity at all. Weigh the honest gain over the best single model against five training pipelines, and set the team standard for how ensembles are validated before anyone ships one.
## The shape of a stack Stacking (stacked generalisation) is a two-level ensemble. **Level 0** holds several base models — say a regularised ridge regression, a nearest-neighbour model and a gradient-boosted tree ensemble, all trained on the same target. **Level 1** holds a single small model, the *meta-learner*, whose input features are not the raw columns at all but the **predictions of the level-0 models**. The meta-learner's job is to learn how much to trust each base model, and in which region of the input space. That is the whole idea. Everything difficult about stacking is about one question: *which* predictions the meta-learner gets to learn from. ## Why in-sample predictions poison it A model's predictions on rows it was trained on are systematically better than its predictions on new rows. A deep tree ensemble may reproduce its own training labels almost exactly; a nearest-neighbour model with k = 1 will reproduce them perfectly, because each row is its own nearest neighbour. If you feed those in-sample predictions to the meta-learner, it sees three columns that are nearly copies of the target. It learns the only rule that fits: trust the model that is most nearly perfect on training data — which is the model that memorised hardest, not the model that generalises best. Validation on those same kinds of features looks spectacular (a 0.97 AUC on weekly store-sales data is the classic shape of it), and the moment the stack scores rows nobody trained on, the base predictions stop being near-copies of the target and the learned weights are pointed at the wrong model. The score collapses to something at or below the best single base model. ## The out-of-fold construction The fix is to build **out-of-fold (OOF) meta-features**: 1. Split the training set into K folds. 2. For each fold k, train every base model on the other K-1 folds and predict fold k. 3. After K rounds, every training row carries one prediction per base model, and each of those predictions came from a model that never saw that row. 4. Stack those columns into an N-by-M matrix — probabilities rather than hard labels for classification, one column per class per model if it is multiclass — and fit the meta-learner on that matrix against the original targets. 5. **Refit each base model on the whole training set.** The fold models were scaffolding; the full-data models are what score live traffic. The meta-learner is kept exactly as trained on the OOF matrix. At serving time a row goes to all M base models, their outputs become the meta-features, and the meta-learner emits the final prediction. (The folds themselves must respect whatever structure the data has — grouping, time ordering — but that is a general splitting concern, not a stacking-specific one.) One asymmetry is worth naming out loud: the meta-learner is *trained* on predictions from models fitted on (K-1)/K of the data, but *serves* predictions from models fitted on all of it. The full-data models are slightly stronger and slightly differently calibrated than the fold models. This mismatch is usually benign and it shrinks as K grows, which is one reason people use a larger K here than they would for plain model selection, or repeat the OOF procedure with several fold seeds and average. ## What the meta-learner should be Small and heavily regularised. It has only a handful of inputs, those inputs are highly correlated with each other and with the target, and each of them carries fold noise. A regularised logistic or linear model — often with non-negative coefficients, which makes the stack a learned weighted average and keeps it interpretable — is the standard choice. A deep second-level tree ensemble over three correlated columns is a reliable way to fit fold noise and give back the gain. Adding a few raw features alongside the meta-features (so the meta-learner can learn *where* each base model is trustworthy) sometimes helps, and sometimes just gives it more rope. ## Blending: the cheap variant Blending holds out a single slice — 10% is typical — instead of running K folds. Base models train on the other 90%, predict the held-out slice, and the meta-learner is fitted on that one set of predictions. It is one round instead of K, it is much harder to get subtly wrong, and it costs you two things: the meta-learner sees only 10% of the rows, so its weights are noisy, and the base models either lose that slice permanently or must be refitted on everything, reintroducing the train/serve mismatch above. Blending is a reasonable choice under a deadline or on a very large dataset where 10% is still hundreds of thousands of rows; full OOF stacking is the choice when data is scarce. ## Scoring the stack honestly The meta-learner's own OOF score is not a clean estimate of the stack: the meta-learner was fitted using all of those OOF rows, and if you tried several meta-learners you also selected on them. Keep a final hold-out that no level of the stack — no base model, no meta-learner, no threshold choice — has touched, and report that number.
- How does blending differ from full out-of-fold stacking?Blending holds out one slice — often 10% — instead of running K folds: base models train on the rest, predict that slice, and the meta-learner is fitted on those predictions alone. It is one round instead of K and much harder to get subtly wrong, but the meta-learner sees far fewer rows, so its weights are noisy, and the held-out slice is either lost to the base models or forces a refit. Prefer it under a deadline or on very large data; prefer full out-of-fold stacking when rows are scarce.
- Which models actually score live traffic once the stack is trained?The base models refitted on the entire training set, plus the meta-learner exactly as it was fitted on the out-of-fold matrix. The per-fold base models existed only to manufacture honest meta-features and are thrown away. This does introduce a mild mismatch — the meta-learner learned weights for slightly weaker fold models — which shrinks as the number of folds grows.
- What should the level-1 meta-learner be, and why not something powerful?Something small and heavily regularised, typically a linear or logistic model, often constrained to non-negative weights. It has only a handful of inputs, they are strongly correlated with each other and the target, and each one carries fold noise. A deep second-level model over three correlated columns mostly fits that noise and hands back the gain the stack was supposed to deliver.
Grading forecasters on the days they already knew the outcome for tells you who has the best memory, not who is the best forecaster. You have to score them only on days they called in advance.
saying these in an interview costs you the question
- Feeds the meta-learner predictions from models that trained on those rows
- Claims a stack cannot overfit because it is an ensemble
- Uses a deep, heavily tuned meta-learner over three correlated inputs
- Reports the meta-learner's own out-of-fold score as the stack's final number
- Thinks out-of-fold simply means predicting on the test set