skip to content

Resampling Schemes

How the rows get divided — a single hold-out split, k-fold and leave-one-out, and folds that respect class balance or grouping. Interviewers start here because the split decides every number after it.

on this pageshow

explore

questions

16

How does a bootstrap resample estimate a model's prediction error, and which rows are held out?

level: juniorimportance: must knowfreq 54%

answer

  1. resample the rows, not the model
  2. draw n rows, replacement allowed
  3. some rows twice, some never
  4. (1 - 1/n)^n tends to 0.368
  5. the never-drawn rows do the scoring

basics

~20 s

Draw n rows with replacement from the sample of n and train on that draw. Roughly 37% of the original rows are never drawn; score the model on those. Repeat a few hundred times and average the errors.

solid answer

~50 s

A bootstrap replicate draws `n` rows with replacement from a sample of size `n`, so some rows appear twice or more and some are never drawn at all. You fit the model on the drawn rows and score it on the rows that draw missed - roughly 37% of the sample, and a different 37% every replicate. Repeat this a few hundred times and average the per-replicate errors to get a bootstrap estimate of prediction error. The appeal on small data is that no row is permanently sacrificed to a hold-out set: on a 400-application loan-default sample, carving off 80 rows you can never train on is expensive, whereas every row eventually plays both roles. The price is that each model is trained on only about 63% distinct rows, and that has consequences for the estimate's bias.

code

python · 18 lines
python
import random

random.seed(7)
rows = list(range(400))          # 400 loan applications, by index
n = len(rows)

# one bootstrap replicate: draw n rows WITH replacement
replicate = [random.choice(rows) for _ in range(n)]

in_bag = set(replicate)                       # distinct rows trained on
oob = [r for r in rows if r not in in_bag]    # rows this draw never took

print("rows in the replicate:", len(replicate))   # always 400 (with repeats)
print("distinct rows trained on:", len(in_bag))   # about 63% of 400
print("rows held out to score:", len(oob))        # about 37% of 400

# the model would be fitted on `replicate` and scored on `oob`,
# and the whole block repeated a few hundred times

go deeper

for a junior

Recall the mechanics: draw n rows with replacement, train on the draw, score on the roughly 37% of rows the draw missed, repeat and average. Be able to say why some rows appear twice.

for a middle

Explain where 63.2% and 36.8% come from via (1 - 1/n)^n, and why duplicates must be kept. Say what is refit inside each replicate and why the apparent error on the resample is not usable.

for a senior

Show judgment about when this beats a single hold-out - small samples where sacrificing rows is costly - and be explicit that each fit sees only about 63% distinct rows, so the estimate leans pessimistic.

for a principal

Own the reporting standard: which resampling scheme produced the headline number, how many replicates, and what sat inside the loop. Argue when the extra compute of hundreds of refits is worth the variance reduction and when it is not.

## What the bootstrap actually does here Start with a sample of `n` labelled rows. One **bootstrap replicate** is built by drawing `n` rows from it **with replacement**: each draw picks uniformly from all `n` rows, and a row that has already been picked can be picked again. The result is a new dataset the same size as the original, but with duplicates - some rows appear twice or three times, and some do not appear at all. Used for **error estimation**, the recipe is: 1. Draw a replicate of size `n` with replacement. 2. Fit the model on that replicate (duplicates and all). 3. Score it on the rows the replicate never drew. 4. Repeat B times (a few hundred is typical) and average the per-replicate errors. The rows a replicate missed are the held-out set for that replicate. They are sometimes called the left-out or out-of-bag rows. Crucially they are *free*: the model never saw them, so scoring on them is an honest hold-out - and because the draw is random, a different subset is held out each time. ## Where the ~37% comes from For a specific row, the chance of *not* being picked on one draw is `1 - 1/n`. There are `n` independent draws, so the chance the row is missed by the whole replicate is `(1 - 1/n)^n`. As `n` grows this converges to `1/e = 0.3679`, and it is already close at modest sizes: about 0.358 at n = 10, about 0.366 at n = 100. So on average roughly **36.8% of rows are left out** and **63.2% of distinct rows are in** each replicate. Note the 63.2% counts *distinct* rows - the replicate still has `n` rows in total, just with repeats. ## Why you would use it instead of one split Consider 400 small-business loan applications with a default flag. A conventional three-way split would hand 20% - eighty applications - to a test set that is never trained on, and eighty rows is both a large slice of your training signal and a small, noisy set to measure error on. The bootstrap avoids that trade: every row is trained on in most replicates and scored in some, and the error is averaged over hundreds of fits rather than read off one unlucky partition. Averaging over many replicates is exactly what makes the estimate low-variance - re-run the whole procedure with a different seed and the number barely moves, which is not true of a single hold-out. ## What people get wrong **"The bootstrap manufactures more data."** It does not. All the information is in the original `n` rows; resampling only lets you reuse them in many different arrangements. If the sample is unrepresentative, every replicate inherits the same distortion. **"Deduplicate the resample before fitting."** No - the duplication is the point. Repeated rows are how a replicate mimics a fresh draw from the population, and stripping them changes the training-set size and the effective weighting. **"Score on the resample you trained on."** That gives the apparent (resubstitution) error, which is optimistically biased and, for a flexible model, can be zero regardless of how bad the model is. **"The same rows are held out each time."** They are not; the held-out set is re-drawn with every replicate, and that is where the averaging power comes from. ## The part you must not skip Everything the model learns from data has to be re-learned *inside* each replicate: scaling statistics, imputation values, feature selection, encoding categories, target-derived transforms. If you fit those once on the full sample and then bootstrap only the final estimator, the left-out rows have already influenced the pipeline and the error comes out too low. The rule is that the replicate defines the training data for the entire procedure, not just the last step. ## The catch to be honest about Because each fit sees only about 63% distinct rows, the model is systematically trained on less data than the one you will eventually ship, which is trained on all `n`. Learning curves slope downward, so a model trained on less data usually predicts worse - and the error measured on the left-out rows therefore comes out **too high**. That upward (pessimistic) bias is real and is the reason corrected estimators exist. It also frames the comparison with k-fold: 10-fold trains on 90% of the rows and so is much less pessimistic, while the bootstrap buys lower variance by averaging over far more fits. ## Reporting it Quote the averaged error, say how many replicates you used, and say what was refit inside each replicate. An error estimate whose resampling scheme is unstated is not interpretable, because the same model can score noticeably differently under a single hold-out, 10-fold, and a bootstrap.

  • Where does the 37% figure come from?
    A given row survives one draw unpicked with probability `1 - 1/n`, and there are `n` independent draws, so it is missed entirely with probability `(1 - 1/n)^n`. That converges to `1/e`, about 0.368, and is already near it for small n - about 0.366 at n = 100. So roughly 37% out, 63% distinct rows in.
  • How many replicates do you need before the estimate settles?
    A few hundred is usually enough for an averaged point estimate, because you are averaging away Monte Carlo noise in the resampling, not sampling noise in the data. The practical check is to rerun the whole procedure with a different random seed: if the estimate moves by less than you would act on, you have enough replicates.
  • What has to be refit inside each replicate rather than once up front?
    Anything learned from data: scaling and centring statistics, imputation values, category encodings, feature selection, any target-derived transform. Fitting those on the full sample first lets the left-out rows influence the pipeline, so they are no longer held out and the error comes back too optimistic. The replicate must define the training data for the whole procedure.

Deal 400 cards from a 400-card deck, putting each card back before dealing the next. Some cards come out twice, and the ones that never came out at all are your surprise quiz.

saying these in an interview costs you the question

  • Claims the bootstrap creates new information from nowhere
  • Removes duplicate rows from the resample before fitting
  • Scores the model on the very rows it was trained on
  • Thinks the same rows are held out in every replicate
  • Draws without replacement, which just shuffles the sample
  • Fits scaling or feature selection once on all rows first

context

open as a page

Why shuffle the rows before splitting them into k-fold cross-validation folds?

level: juniorimportance: must knowfreq 64%

basics

~20 s

Exported files usually arrive sorted — by label, by date, by customer. Cutting such a file into contiguous blocks gives folds that are not representative, sometimes single-class, so every fold score misleads. Shuffling row order first breaks that ordering.

open as a page

What does stratified k-fold cross-validation preserve, and when does plain k-fold fail?

level: juniorimportance: must knowfreq 78%

basics

~20 s

Stratified k-fold gives every fold roughly the same class proportions as the whole dataset. Plain k-fold fails when a class is rare: some folds get very few or zero minority rows, so fold scores swing or go undefined.

open as a page

Why does tuning a model require a validation set separate from the test set?

level: juniorimportance: must knowfreq 82%

basics

~20 s

Every choice made by looking at a score - hyperparameters, features, thresholds - bends the model toward those rows. Validation absorbs that optimism; the untouched test set then gives an honest estimate of performance on new data.

open as a page

In k-fold cross-validation, how does the choice of k trade off bias against variance?

level: middleimportance: must knowfreq 78%

basics

~20 s

Larger k trains each fold's model on more rows, so the error estimate is less pessimistic, but the training sets overlap more and each held-out fold is smaller, so the estimate is noisier. k=5 or 10 balances both.

open as a page

When do you need group k-fold or leave-one-subject-out instead of ordinary k-fold?

level: middleimportance: must knowfreq 62%

basics

~20 s

Grouped folds are for data with repeated rows per entity — many windows from one wearer, many rows from one store. Every row of a group goes into one fold, so the score measures generalisation to unseen entities.

open as a page

How should dataset size change your train/validation/test split ratios?

level: middleimportance: must knowfreq 62%

basics

~20 s

Ratios are a heuristic; what matters is the absolute count in each partition. With a few hundred rows the estimate is too noisy to trust, so resample instead; with millions, a one-percent hold-out is plenty.

open as a page

How do you put a percentile band on a classifier's accuracy using bootstrap replicates?

level: middleimportance: should knowfreq 45%

basics

~20 s

Resample the evaluation rows with replacement, recompute accuracy on each resample, and take the 2.5th and 97.5th percentiles of those values as a 95% band. It captures evaluation-set sampling noise, with the trained model held fixed.

open as a page

When should a train/validation/test split be stratified rather than purely random?

level: middleimportance: should knowfreq 48%

basics

~20 s

Stratify whenever a partition is small enough that random assignment could land a class mix quite different from the full table. Sampling within each class keeps the proportions matched, so partitions are comparable and no class nearly disappears.

open as a page

When is leave-one-out cross-validation worth its n model fits instead of 5-fold?

level: seniorimportance: should knowfreq 55%

basics

~20 s

Only when rows are scarce and each fit is cheap. Leave-one-out trains on all but one row, so it is nearly unbiased and deterministic, but it costs one fit per row and averages near-identical models, leaving the estimate jumpy.

open as a page

Leave-one-store-out CV scores 0.71 where random 10-fold scores 0.94 — which do you trust?

level: seniorimportance: should knowfreq 42%

basics

~10 s

The two schemes answer different questions: random 10-fold describes new rows from branches already in training, leave-one-store-out describes a branch never seen. Report whichever matches who the model will actually score.

open as a page

Your team has scored the same hold-out test set 40 times this quarter - what now?

level: principalimportance: should knowfreq 40%

basics

~20 s

That hold-out is no longer a test set: forty scored comparisons made it a validation set, and its number is optimistic by an unmeasurable amount. Get fresh untouched rows for the real verdict, and budget future use.

open as a page

How do you stratify cross-validation folds when the regression target is continuous?

level: middleimportance: nice to knowfreq 30%

basics

~20 s

Turn the continuous target into a temporary discrete label — quantile bins such as quartiles — and stratify the folds on that bin. Every fold then carries a similar spread of target values, steadying fold metrics on small or skewed data.

open as a page

Why is bootstrap left-out error pessimistic, and what does the .632 estimator correct?

level: seniorimportance: nice to knowfreq 28%

basics

~20 s

Each bootstrap replicate trains on only about 63% distinct rows, so its models are weaker and the error on left-out rows comes out too high. The .632 estimator blends that pessimistic error with the optimistic training error, 0.632 to 0.368.

open as a page

What do you do when 5-fold cross-validation scores swing by 0.08 between random partitions?

level: seniorimportance: nice to knowfreq 38%

basics

~20 s

Treat that swing as a measurement of your own noise floor. Repeat the whole cross-validation with several different random partitions and average across all runs, and refuse to call any model difference smaller than the swing you just observed.

open as a page

Should you refit on train plus validation once the hyperparameters are chosen?

level: seniorimportance: nice to knowfreq 28%

basics

~20 s

Usually yes. Once the configuration is frozen, the validation rows are just labelled data, and training on more rows generally gives a better model. Report the test score for the refit model, and never redo selection after folding validation in.

open as a page