What is bootstrap resampling, and how does it produce a confidence interval for a correlation?
answer
- the sample stands in for the population
- same size n, drawn with replacement
- recompute the statistic every time
- middle 95% of the replicate values
basics
~20 sThe bootstrap treats your sample as a stand-in for the population: draw many new samples of the same size with replacement, recompute the statistic on each, and read the interval off the middle 95% of those values.
solid answer
~40 sThe bootstrap estimates the sampling distribution of a statistic when no formula for it exists. You hold one sample of size `n`; you draw B resamples, commonly 10,000, each also of size `n`, drawing with replacement so some observations appear twice and others not at all. Recompute the statistic on every resample and you get B values whose spread approximates how much the estimate would wobble across repeated samples from the population. The percentile interval is then simply the 2.5th and 97.5th percentiles of those B values. That is why it is reached for on statistics with no clean closed form, such as a correlation coefficient or a 10% trimmed mean. It adds no information: it reuses the sample you already have, so a biased or unrepresentative sample yields a confidently wrong interval.
code
python · 16 linesimport random, statistics
data = [12, 15, 9, 22, 31, 14, 8, 19, 27, 11, 16, 24]
B = 10000
def trimmed_mean(xs, frac=0.10):
xs = sorted(xs)
k = int(len(xs) * frac)
return statistics.fmean(xs[k:len(xs) - k] if k else xs)
random.seed(0)
n = len(data)
reps = sorted(trimmed_mean(random.choices(data, k=n)) for _ in range(B))
lo = reps[int(0.025 * B)]
hi = reps[int(0.975 * B) - 1]
print(round(trimmed_mean(data), 2), round(lo, 2), round(hi, 2))go deeper
Be ready to state the loop out loud: draw n values with replacement, recompute the statistic, repeat thousands of times, then take the 2.5th and 97.5th percentiles of what you collected.
An interviewer expects you to explain why replacement is essential and why the resample size must equal n, and to name statistics such as a correlation or a trimmed mean where no textbook standard-error formula exists.
Show you know what the interval inherits. Resampling the empirical distribution carries selection bias, clustering and time dependence straight through, so say how you would resample the unit that is actually independent rather than the row.
Own the call of when resampling is worth its compute at all, what B the organisation standardises on, and how results are made reproducible and reviewable through fixed seeds, a recorded B and a documented resampling unit.
## The problem it solves Classical inference hands you an algebraic standard error for a small set of statistics: a sample mean, a proportion, a regression slope under stated assumptions. For most quantities people actually report, no such formula is available, or the available one leans on assumptions the data violates. A correlation coefficient, a 10% trimmed mean, the ratio of two estimated quantities, the difference between two such ratios: writing down the sampling distribution analytically ranges from painful to impossible. The bootstrap replaces the missing algebra with computation. ## The plug-in idea The uncertainty you want is variation across repeated samples drawn from the population. You cannot draw those, because you have one sample and no population. The bootstrap substitutes the **empirical distribution**: the distribution that places probability 1/n on each of the n observed values. Drawing an observation from that distribution is exactly the same operation as picking one of your data points uniformly at random. Drawing n of them independently from it is exactly the same as sampling n values from your data *with replacement*. So resampling with replacement is not a trick or an approximation of convenience; it is literally sampling from the best non-parametric estimate of the population you have. The assumption being made is now visible: the empirical distribution must be a decent stand-in for the true one, and the statistic must respond smoothly to small changes in that distribution. When both hold, the variation of the statistic across resamples tracks its variation across genuine repeat samples. ## The algorithm 1. Compute the statistic once on the full sample. Call it the point estimate. 2. Draw a resample of size n from the sample, with replacement. 3. Recompute the statistic on that resample and store it. 4. Repeat steps 2 and 3 B times, with B typically 2,000 to 10,000 or more. 5. You now hold B replicate values. Their standard deviation estimates the standard error of the statistic. Their 2.5th and 97.5th percentiles form a 95% percentile confidence interval. Two details in step 2 carry weight. **With replacement** is essential: sampling without replacement from your own sample of size n just reorders it, every replicate is identical, and the estimated spread collapses to zero. **Size n** is essential too: the spread of a sampling distribution depends on how many observations went into the estimate, roughly shrinking like 1/sqrt(n) for well-behaved statistics. A resample of size n/2 would produce a distribution that is too wide and an interval that over-covers. ## Reading an interval off the replicates The percentile interval is the simplest recipe: sort the B replicates and take the alpha/2 and 1 - alpha/2 quantiles. For a 95% interval that is the 2.5th and 97.5th. It is attractive because it respects the parameter's natural range and any monotone transformation of it. A correlation interval computed this way cannot fall outside the interval from -1 to 1, because no replicate can. Other recipes exist and adjust these endpoints for bias and skew, but the percentile interval is the one to be able to describe on demand. ## What it does not buy you The most common misconception is that resampling manufactures information. It does not. Every replicate is built from the same n observations. Raising B from 1,000 to 100,000 makes the reported endpoints more stable across reruns with different seeds, and nothing else: it reduces **Monte Carlo noise** in the procedure, not **statistical uncertainty** about the population, which is fixed by n and by the data-generating process. Similarly, the bootstrap inherits every defect of the sample. If the sampling frame missed a segment of the population, every single resample misses it too, and the interval will be tight around the wrong value. Selection bias, measurement bias and non-response are fixed by design, weighting or better data collection, never by resampling. Dependence is the third trap. The ordinary bootstrap assumes the units you resample are independent. If rows are grouped, repeated or ordered in time, resampling rows treats correlated observations as if they were independent and produces intervals that are far too narrow. The repair is to resample the unit that actually is independent, whole groups or blocks rather than single rows. ## Practical notes Record and report the seed and the value of B so the numbers reproduce exactly. Look at the histogram of the replicates rather than only at two percentiles: a distribution with visible spikes, atoms sitting on individual data values, or heavy skew is telling you that either the statistic or the sample size is a poor fit for the ordinary bootstrap.
- Why must each bootstrap resample be the same size as the original sample?Because the spread of a sampling distribution depends on sample size, shrinking roughly like 1/sqrt(n) for well-behaved statistics. Resampling m values with m smaller than n gives a distribution that is too wide and an interval that over-covers; m larger than n gives one that is too narrow. Deliberately choosing m much smaller than n is a specialised repair for cases where the ordinary bootstrap is inconsistent, not a default.
- How many resamples should you draw, and what changes if you draw 1,000 instead of 10,000?B controls only simulation noise in the answer, never statistical uncertainty about the population, which is fixed by n. A thousand replicates is usually enough for a standard error, but interval endpoints sit out in the tails where estimates are noisiest, so 10,000 or more is the common choice for a 95% interval. Fix and record the seed so the endpoints reproduce.
- Can the bootstrap rescue a sample you know is unrepresentative?No. It resamples the empirical distribution, so it inherits every bias in the data. What it estimates is how much the estimate would move under repeated sampling from a population that looks like your sample. If the sampling frame missed a segment, every resample misses it too, and you get a tight interval around the wrong value. Selection and measurement bias are design problems, not resampling problems.
You have one photograph of a crowd and want to know how much a head count would vary. The bootstrap builds new crowds by drawing faces from the photo at random with repeats allowed, counts each one, and treats the spread of counts as the wobble you would have seen across real crowds.
saying these in an interview costs you the question
- Says resampling creates new data or new information
- Draws resamples without replacement, which merely reorders the sample
- Uses a resample size different from n as the default
- Claims the bootstrap corrects a biased or unrepresentative sample
- Thinks more resamples make the confidence interval narrower