How does a bootstrap resample estimate a model's prediction error, and which rows are held out?
answer
- resample the rows, not the model
- draw n rows, replacement allowed
- some rows twice, some never
- (1 - 1/n)^n tends to 0.368
- the never-drawn rows do the scoring
basics
~20 sDraw n rows with replacement from the sample of n and train on that draw. Roughly 37% of the original rows are never drawn; score the model on those. Repeat a few hundred times and average the errors.
solid answer
~50 sA bootstrap replicate draws `n` rows with replacement from a sample of size `n`, so some rows appear twice or more and some are never drawn at all. You fit the model on the drawn rows and score it on the rows that draw missed - roughly 37% of the sample, and a different 37% every replicate. Repeat this a few hundred times and average the per-replicate errors to get a bootstrap estimate of prediction error. The appeal on small data is that no row is permanently sacrificed to a hold-out set: on a 400-application loan-default sample, carving off 80 rows you can never train on is expensive, whereas every row eventually plays both roles. The price is that each model is trained on only about 63% distinct rows, and that has consequences for the estimate's bias.
code
python · 18 linesimport random
random.seed(7)
rows = list(range(400)) # 400 loan applications, by index
n = len(rows)
# one bootstrap replicate: draw n rows WITH replacement
replicate = [random.choice(rows) for _ in range(n)]
in_bag = set(replicate) # distinct rows trained on
oob = [r for r in rows if r not in in_bag] # rows this draw never took
print("rows in the replicate:", len(replicate)) # always 400 (with repeats)
print("distinct rows trained on:", len(in_bag)) # about 63% of 400
print("rows held out to score:", len(oob)) # about 37% of 400
# the model would be fitted on `replicate` and scored on `oob`,
# and the whole block repeated a few hundred timesgo deeper
Recall the mechanics: draw n rows with replacement, train on the draw, score on the roughly 37% of rows the draw missed, repeat and average. Be able to say why some rows appear twice.
Explain where 63.2% and 36.8% come from via (1 - 1/n)^n, and why duplicates must be kept. Say what is refit inside each replicate and why the apparent error on the resample is not usable.
Show judgment about when this beats a single hold-out - small samples where sacrificing rows is costly - and be explicit that each fit sees only about 63% distinct rows, so the estimate leans pessimistic.
Own the reporting standard: which resampling scheme produced the headline number, how many replicates, and what sat inside the loop. Argue when the extra compute of hundreds of refits is worth the variance reduction and when it is not.
## What the bootstrap actually does here Start with a sample of `n` labelled rows. One **bootstrap replicate** is built by drawing `n` rows from it **with replacement**: each draw picks uniformly from all `n` rows, and a row that has already been picked can be picked again. The result is a new dataset the same size as the original, but with duplicates - some rows appear twice or three times, and some do not appear at all. Used for **error estimation**, the recipe is: 1. Draw a replicate of size `n` with replacement. 2. Fit the model on that replicate (duplicates and all). 3. Score it on the rows the replicate never drew. 4. Repeat B times (a few hundred is typical) and average the per-replicate errors. The rows a replicate missed are the held-out set for that replicate. They are sometimes called the left-out or out-of-bag rows. Crucially they are *free*: the model never saw them, so scoring on them is an honest hold-out - and because the draw is random, a different subset is held out each time. ## Where the ~37% comes from For a specific row, the chance of *not* being picked on one draw is `1 - 1/n`. There are `n` independent draws, so the chance the row is missed by the whole replicate is `(1 - 1/n)^n`. As `n` grows this converges to `1/e = 0.3679`, and it is already close at modest sizes: about 0.358 at n = 10, about 0.366 at n = 100. So on average roughly **36.8% of rows are left out** and **63.2% of distinct rows are in** each replicate. Note the 63.2% counts *distinct* rows - the replicate still has `n` rows in total, just with repeats. ## Why you would use it instead of one split Consider 400 small-business loan applications with a default flag. A conventional three-way split would hand 20% - eighty applications - to a test set that is never trained on, and eighty rows is both a large slice of your training signal and a small, noisy set to measure error on. The bootstrap avoids that trade: every row is trained on in most replicates and scored in some, and the error is averaged over hundreds of fits rather than read off one unlucky partition. Averaging over many replicates is exactly what makes the estimate low-variance - re-run the whole procedure with a different seed and the number barely moves, which is not true of a single hold-out. ## What people get wrong **"The bootstrap manufactures more data."** It does not. All the information is in the original `n` rows; resampling only lets you reuse them in many different arrangements. If the sample is unrepresentative, every replicate inherits the same distortion. **"Deduplicate the resample before fitting."** No - the duplication is the point. Repeated rows are how a replicate mimics a fresh draw from the population, and stripping them changes the training-set size and the effective weighting. **"Score on the resample you trained on."** That gives the apparent (resubstitution) error, which is optimistically biased and, for a flexible model, can be zero regardless of how bad the model is. **"The same rows are held out each time."** They are not; the held-out set is re-drawn with every replicate, and that is where the averaging power comes from. ## The part you must not skip Everything the model learns from data has to be re-learned *inside* each replicate: scaling statistics, imputation values, feature selection, encoding categories, target-derived transforms. If you fit those once on the full sample and then bootstrap only the final estimator, the left-out rows have already influenced the pipeline and the error comes out too low. The rule is that the replicate defines the training data for the entire procedure, not just the last step. ## The catch to be honest about Because each fit sees only about 63% distinct rows, the model is systematically trained on less data than the one you will eventually ship, which is trained on all `n`. Learning curves slope downward, so a model trained on less data usually predicts worse - and the error measured on the left-out rows therefore comes out **too high**. That upward (pessimistic) bias is real and is the reason corrected estimators exist. It also frames the comparison with k-fold: 10-fold trains on 90% of the rows and so is much less pessimistic, while the bootstrap buys lower variance by averaging over far more fits. ## Reporting it Quote the averaged error, say how many replicates you used, and say what was refit inside each replicate. An error estimate whose resampling scheme is unstated is not interpretable, because the same model can score noticeably differently under a single hold-out, 10-fold, and a bootstrap.
- Where does the 37% figure come from?A given row survives one draw unpicked with probability `1 - 1/n`, and there are `n` independent draws, so it is missed entirely with probability `(1 - 1/n)^n`. That converges to `1/e`, about 0.368, and is already near it for small n - about 0.366 at n = 100. So roughly 37% out, 63% distinct rows in.
- How many replicates do you need before the estimate settles?A few hundred is usually enough for an averaged point estimate, because you are averaging away Monte Carlo noise in the resampling, not sampling noise in the data. The practical check is to rerun the whole procedure with a different random seed: if the estimate moves by less than you would act on, you have enough replicates.
- What has to be refit inside each replicate rather than once up front?Anything learned from data: scaling and centring statistics, imputation values, category encodings, feature selection, any target-derived transform. Fitting those on the full sample first lets the left-out rows influence the pipeline, so they are no longer held out and the error comes back too optimistic. The replicate must define the training data for the whole procedure.
Deal 400 cards from a 400-card deck, putting each card back before dealing the next. Some cards come out twice, and the ones that never came out at all are your surprise quiz.
saying these in an interview costs you the question
- Claims the bootstrap creates new information from nowhere
- Removes duplicate rows from the resample before fitting
- Scores the model on the very rows it was trained on
- Thinks the same rows are held out in every replicate
- Draws without replacement, which just shuffles the sample
- Fits scaling or feature selection once on all rows first