skip to content

When does the bootstrap fail, and what goes wrong for the sample maximum or for n = 5?

level: seniorimportance: should knowfreq 45%

answer

  1. empirical distribution must resemble the population
  2. resamples cannot exceed what you observed
  3. extremes and non-smooth statistics misbehave
  4. an atom of about 63% at one point
  5. m-out-of-n or subsampling as the repair

basics

~20 s

The bootstrap fails when the sample is a poor stand-in for the population or the statistic is not smooth. No resample can exceed the observed maximum, so extremes give degenerate intervals; n = 5 gives too little to resample.

solid answer

~50 s

The bootstrap assumes the empirical distribution of your sample is a serviceable stand-in for the population and that the statistic is a smooth function of it. Both can break. For the sample maximum it breaks badly: every resample's maximum is one of the observed values and can never exceed the observed maximum, so the bootstrap distribution piles roughly 63% of its mass on that single point. It is inconsistent for extreme order statistics and the interval never covers upward. Tiny samples break the other assumption: at n = 5 there are only 126 distinct resamples ignoring order, and a five-point empirical distribution does not resemble a continuous population, so real coverage falls far below nominal. Heavy tails with infinite variance and parameters on a boundary are the other classic failures. The repairs are an m-out-of-n bootstrap, subsampling, or an explicit extreme-value model.

go deeper

for a junior

Know the one-line limit: the bootstrap can only reshuffle what you already collected, so nothing it returns can reach beyond the range of the observed data.

for a middle

Explain the two assumptions behind it, that the empirical distribution stands in for the population and that the statistic varies smoothly with it, and name the sample maximum as the textbook counterexample.

for a senior

Show that you check rather than trust. Inspect the bootstrap distribution for atoms and spikes, watch the endpoints across seeds, and move to subsampling, an m-out-of-n scheme or an extreme-value model when the statistic is extremal.

for a principal

Own the standard for when a resampled interval may be published at all: a coverage check by simulation on a realistic generating model, and a documented rule about which sample sizes and which statistics your teams may bootstrap.

## The two assumptions that can break The ordinary bootstrap rests on two things being true at once. First, the empirical distribution, which places mass 1/n on each observed value, must be close enough to the true distribution that sampling from one behaves like sampling from the other. Second, the statistic must be a **smooth functional** of that distribution: a small perturbation of the distribution should move the statistic only a little. When either fails, resampling produces confident nonsense, and it does so silently, because the procedure always returns two numbers. ## The sample maximum: a clean counterexample Suppose you estimate the largest value a process can produce by the largest value observed. Bootstrap it and something disturbing happens: the maximum of any resample is one of the values already in your data, so the bootstrap distribution is capped at the observed maximum. It cannot look upward at all, and the upper end of the interval is pinned to a data point. Worse, the distribution is not merely capped, it is lumpy. Each of the n independent draws in a resample misses a specific observation with probability (1 - 1/n), so that observation is absent from the whole resample with probability (1 - 1/n) raised to the power n, which tends to 1/e, about 0.368, as n grows. The observed maximum is therefore present in a resample about 63% of the time, and whenever it is present it *is* the resample maximum. The bootstrap distribution has an atom of roughly 63% at a single point. Increasing B does nothing: the atom is a property of the resampling scheme, not of the simulation size. Formally, the bootstrap is **inconsistent** for extreme order statistics, and any interval you read off is wrong in a way that does not shrink with more computation. ## Tiny samples At n = 5 the empirical distribution is five spikes. The number of distinct resamples, counting multisets and ignoring order, is the number of ways to choose 5 items from 5 with repetition, which equals 126. The bootstrap distribution of any statistic is therefore supported on at most 126 values, often far fewer once ties collapse, and it is visibly discrete. More fundamentally, five points from a continuous distribution simply do not pin down its shape, especially its tails, and every percentile the bootstrap reports is a percentile of a distribution built from those five points. Coverage of a nominal 95% interval can be dramatically below 95%. The honest response to n = 5 is to say the data are insufficient, not to reach for a computational patch. ## The other classic failures **Heavy tails.** If the underlying distribution has infinite variance, the bootstrap for the sample mean is inconsistent: the limiting behaviour of the mean is governed by the tail, and the empirical distribution has, by construction, bounded support and finite variance. **Boundary parameters.** When the true parameter sits on the edge of its allowed range, for example a variance component that is genuinely zero and cannot be negative, the estimator's distribution is not smooth at that point and the bootstrap does not reproduce it. **Non-smooth statistics generally.** The median is a milder case, where the ordinary bootstrap works but converges slowly. Statistics defined by a hard threshold or a selection step, such as an estimate produced after choosing the best-looking of several candidates, break the smoothness assumption more seriously. **Dependence.** Rows that are grouped or ordered in time are not independent, and resampling them as if they were shrinks intervals sharply. This is the failure most likely to appear in real work, and the repair is to resample the independent unit, whole groups or contiguous blocks, rather than the row. ## Repairs The **m-out-of-n bootstrap** draws resamples of size m with m much smaller than n, letting m grow to infinity but m/n go to zero. Shrinking the resample loosens its dependence on the exact observed sample and restores consistency in several of the cases above, at the price of rescaling the resulting distribution to the right sample size and choosing m, which is itself a tuning problem. **Subsampling**, drawing subsets without replacement and rescaling, works under strikingly weak conditions and is the more general tool. For extremes specifically, the right answer is usually a model: extreme-value theory gives a parametric family for the behaviour of maxima, and fitting that is far more informative than resampling the maximum ever will be. ## How to notice in practice Do not read only the two endpoints. Plot the bootstrap distribution. Warning signs are a visible atom or spike, an endpoint sitting exactly on an observed data value, strong skew that the interval flavour is not designed for, or endpoints that move materially when you change the seed or B. The definitive check is a simulation study: generate data from a plausible model where you know the truth, run the entire procedure many times, and count how often the interval actually contains it. If nominal 95% coverage comes out at 71%, you have your answer, and no amount of resampling will change it.

  • Where does the 63% figure for the bootstrap distribution of the maximum come from?
    Each of the n draws in a resample misses a given observation with probability (1 - 1/n), so that observation is absent from the entire resample with probability (1 - 1/n) to the power n, which tends to 1/e, about 0.368. The observed maximum is therefore present, and hence is the resample maximum, roughly 63% of the time. That atom is built into the scheme and no larger B smooths it away.
  • What is the m-out-of-n bootstrap and why does it help?
    It draws resamples of size m with m much smaller than n, letting m grow to infinity while m/n goes to zero. The smaller resample is less tightly bound to the exact observed sample, which restores consistency in several cases where the ordinary bootstrap fails, including extremes and some non-smooth functionals. The price is rescaling the resulting distribution to the correct sample size and choosing m, which becomes its own tuning decision.
  • How would you tell in practice that a bootstrap interval is untrustworthy?
    Look at the whole bootstrap distribution rather than only the two percentiles. Warning signs are atoms or spikes, an endpoint landing exactly on an observed data value, extreme skew, or endpoints that shift materially when you change the seed or B. The decisive check is a simulation study: generate data from a plausible model with a known truth, run the whole procedure many times, and count actual coverage.

saying these in an interview costs you the question

  • Believes the bootstrap is assumption-free and always valid
  • Bootstraps a sample maximum or minimum without hesitation
  • Treats a tiny sample as fine because B is large
  • Confuses more resamples with more data
  • Never inspects the bootstrap distribution, only its percentiles

context