skip to content

Why does a bootstrap resample of n rows contain only about 63% of the distinct original rows?

level: middleimportance: should knowfreq 46%

answer

  1. probability a row is missed each draw
  2. independence across n draws
  3. (1 - 1/n)^n has a famous limit
  4. 1/e is about 0.368
  5. the resample still holds n rows

basics

~20 s

Drawing n rows with replacement, a given row is missed on every draw with probability (1 - 1/n)^n, which converges to 1/e, about 0.368. So roughly 63.2% of the distinct rows land in-bag and 36.8% are out-of-bag.

solid answer

~50 s

A bootstrap resample makes `n` independent draws from `n` rows, each draw uniform. A given row survives one draw un-picked with probability `1 - 1/n`, and the draws are independent, so it is missed by all of them with probability `(1 - 1/n)^n`. That expression converges to `1/e = 0.3679` as n grows, so the expected fraction of distinct rows present is `1 - 1/e = 0.632`. For n = 1,000 the exact value is `1 - 0.999^1000 = 0.6323`. Two things candidates get wrong: the resample still has n rows — the missing 37% are made up by duplicates of in-bag rows — and the convergence is fast, so the figure is already about 65% at n = 10 and about 63.4% at n = 100. The roughly 37% of rows a given base model never saw are that model's out-of-bag rows.

code

python · 14 lines
python
import math
import random

n = 1000
trials = 200
in_bag = []

for _ in range(trials):
    resample = [random.randrange(n) for _ in range(n)]
    in_bag.append(len(set(resample)) / n)

print("simulated in-bag fraction:", round(sum(in_bag) / trials, 4))
print("formula 1 - (1 - 1/n)**n: ", round(1 - (1 - 1 / n) ** n, 4))
print("limit 1 - 1/e:            ", round(1 - 1 / math.e, 4))

go deeper

for a junior

Know the headline numbers: about 63% of distinct rows in, about 37% out, and the resample is still the same size as the original table because rows repeat.

for a middle

Be ready to derive it live. State the per-draw miss probability, multiply across n independent draws, and name the limit 1/e — the interviewer is checking that you can reason about the resampling scheme, not that you memorised 0.632.

for a senior

Turn the arithmetic into a design fact: the complement is the out-of-bag pool, and roughly 0.368 times the number of base models get to score any single row. That figure tells you whether an out-of-bag estimate on your ensemble size is stable or noise.

for a principal

Be able to argue when the default resample size should change at all. Shrinking it buys diversity and training speed on very large tables; on small ones it starves each base model, and the honest answer is usually to leave it alone.

## The derivation A bootstrap resample of a table with `n` rows is built by making `n` independent draws, each one picking a row uniformly at random from all `n` and putting it back. Fix one particular row, say row 7. - On any single draw, the chance of picking row 7 is `1/n`, so the chance of missing it is `1 - 1/n`. - The draws are independent, so the chance of missing it on all `n` draws is `(1 - 1/n)^n`. - Therefore the chance row 7 appears at least once is `1 - (1 - 1/n)^n`. Because the same argument applies to every row, the *expected* fraction of distinct rows in the resample is exactly that number. The limit is the classic one: `(1 - 1/n)^n -> e^-1 = 0.3679`, so the in-bag fraction tends to `1 - 1/e = 0.6321`. ## How fast it settles The limit is reached quickly, which is why practitioners quote 63% for any real dataset: - n = 5: out-of-bag `0.8^5 = 0.328`, in-bag 67.2% - n = 10: out-of-bag `0.9^10 = 0.349`, in-bag 65.1% - n = 100: out-of-bag 0.366, in-bag 63.4% - n = 1,000: out-of-bag 0.3677, in-bag 63.23% - limit: in-bag 63.21% Note the direction: small tables keep *more* of their distinct rows in-bag, and the in-bag fraction falls to 63.2% from above. ## The resample still has n rows This is the point most often muddled. The resample is not a 63%-sized subset. It contains exactly `n` row slots; about 63% of the distinct original rows fill them, and the shortfall is covered by duplicates. How many times does a given row appear? Its count is Binomial(n, 1/n), which for large n is very close to Poisson with mean 1: - about 36.8% of rows appear zero times - about 36.8% appear exactly once - about 18.4% appear twice - about 8% appear three or more times The duplication matters for the learner. A row that appears three times carries three times the weight in a split-quality calculation or a loss sum, which is part of why two base models trained on two resamples of the same table can end up structurally different. ## Why the number is useful The complement is the operationally interesting part. For each base model, the roughly 36.8% of rows that its resample missed are that model's **out-of-bag rows** — rows it has genuinely never been fit on. Aggregating each row's prediction over only those base models that missed it gives an error estimate without setting aside any data, which is the single most attractive property of bagging on small tables. It also tells you roughly how many base models get to vote on any one row's out-of-bag prediction: about `0.368 * B`. With B = 200 base models that is about 74 of them, which is enough for the estimate to be reasonably stable; with B = 20 it is about 7, which is not. ## Variations The 63% figure assumes the resample size equals the training size, which is the default. If you deliberately draw fewer than `n` rows — sometimes done to speed up training on very large tables — the in-bag fraction drops and the out-of-bag pool grows, giving more diversity between base models at the cost of a weaker fit per model. Drawing rows *without* replacement is a different scheme entirely (subsampling), and none of the arithmetic above applies to it: there, the in-bag fraction is simply whatever proportion you chose to draw.

  • Does the 63% figure depend on the size of the training table?
    Only for tiny tables. It is a limit: 67% in-bag at n = 5, 65% at n = 10, 63.4% at n = 100, and 63.2% from a few hundred rows onward. The in-bag fraction falls toward 63.2% from above, so small tables keep slightly more of their distinct rows and leave slightly fewer out-of-bag.
  • How many times does a typical in-bag row appear in one resample?
    Each row's count is Binomial(n, 1/n), which is close to Poisson with mean 1 for any realistic n. About 37% of rows appear zero times, another 37% appear exactly once, about 18% appear twice, and about 8% appear three or more times. Duplicated rows carry proportionally more weight in the base learner's loss.
  • What changes if you draw fewer than n rows per resample?
    The in-bag fraction drops below 63% and the out-of-bag pool grows, so the base models see less data each and disagree with each other more. You trade a weaker individual model for more diversity, plus faster training — sometimes worth it on very large tables, rarely worth it on small ones.

Draw 1,000 raffle tickets from a drum of 1,000, putting each ticket back after you read it. You end up with 1,000 slips of paper, but some names come up twice or three times and about a third of the names never come out at all.

saying these in an interview costs you the question

  • Says the resample is 63% of the rows drawn without replacement
  • Thinks the bootstrap sample contains only 0.63n rows
  • Confuses the 63% with a train/validation split ratio
  • Assumes every in-bag row appears exactly once
  • Cannot connect the missing 37% to out-of-bag rows

context