skip to content

Multiplicity and Resampling

What happens once you run many tests, and what to do when no formula fits: family-wise error control, false-discovery rate, and bootstrap or permutation resampling. Analyses quietly break here.

on this pageshow

explore

questions

20

What is bootstrap resampling, and how does it produce a confidence interval for a correlation?

level: juniorimportance: must knowfreq 62%

answer

  1. the sample stands in for the population
  2. same size n, drawn with replacement
  3. recompute the statistic every time
  4. middle 95% of the replicate values

basics

~20 s

The bootstrap treats your sample as a stand-in for the population: draw many new samples of the same size with replacement, recompute the statistic on each, and read the interval off the middle 95% of those values.

solid answer

~40 s

The bootstrap estimates the sampling distribution of a statistic when no formula for it exists. You hold one sample of size `n`; you draw B resamples, commonly 10,000, each also of size `n`, drawing with replacement so some observations appear twice and others not at all. Recompute the statistic on every resample and you get B values whose spread approximates how much the estimate would wobble across repeated samples from the population. The percentile interval is then simply the 2.5th and 97.5th percentiles of those B values. That is why it is reached for on statistics with no clean closed form, such as a correlation coefficient or a 10% trimmed mean. It adds no information: it reuses the sample you already have, so a biased or unrepresentative sample yields a confidently wrong interval.

code

python · 16 lines
python
import random, statistics

data = [12, 15, 9, 22, 31, 14, 8, 19, 27, 11, 16, 24]
B = 10000

def trimmed_mean(xs, frac=0.10):
    xs = sorted(xs)
    k = int(len(xs) * frac)
    return statistics.fmean(xs[k:len(xs) - k] if k else xs)

random.seed(0)
n = len(data)
reps = sorted(trimmed_mean(random.choices(data, k=n)) for _ in range(B))
lo = reps[int(0.025 * B)]
hi = reps[int(0.975 * B) - 1]
print(round(trimmed_mean(data), 2), round(lo, 2), round(hi, 2))

go deeper

for a junior

Be ready to state the loop out loud: draw n values with replacement, recompute the statistic, repeat thousands of times, then take the 2.5th and 97.5th percentiles of what you collected.

for a middle

An interviewer expects you to explain why replacement is essential and why the resample size must equal n, and to name statistics such as a correlation or a trimmed mean where no textbook standard-error formula exists.

for a senior

Show you know what the interval inherits. Resampling the empirical distribution carries selection bias, clustering and time dependence straight through, so say how you would resample the unit that is actually independent rather than the row.

for a principal

Own the call of when resampling is worth its compute at all, what B the organisation standardises on, and how results are made reproducible and reviewable through fixed seeds, a recorded B and a documented resampling unit.

## The problem it solves Classical inference hands you an algebraic standard error for a small set of statistics: a sample mean, a proportion, a regression slope under stated assumptions. For most quantities people actually report, no such formula is available, or the available one leans on assumptions the data violates. A correlation coefficient, a 10% trimmed mean, the ratio of two estimated quantities, the difference between two such ratios: writing down the sampling distribution analytically ranges from painful to impossible. The bootstrap replaces the missing algebra with computation. ## The plug-in idea The uncertainty you want is variation across repeated samples drawn from the population. You cannot draw those, because you have one sample and no population. The bootstrap substitutes the **empirical distribution**: the distribution that places probability 1/n on each of the n observed values. Drawing an observation from that distribution is exactly the same operation as picking one of your data points uniformly at random. Drawing n of them independently from it is exactly the same as sampling n values from your data *with replacement*. So resampling with replacement is not a trick or an approximation of convenience; it is literally sampling from the best non-parametric estimate of the population you have. The assumption being made is now visible: the empirical distribution must be a decent stand-in for the true one, and the statistic must respond smoothly to small changes in that distribution. When both hold, the variation of the statistic across resamples tracks its variation across genuine repeat samples. ## The algorithm 1. Compute the statistic once on the full sample. Call it the point estimate. 2. Draw a resample of size n from the sample, with replacement. 3. Recompute the statistic on that resample and store it. 4. Repeat steps 2 and 3 B times, with B typically 2,000 to 10,000 or more. 5. You now hold B replicate values. Their standard deviation estimates the standard error of the statistic. Their 2.5th and 97.5th percentiles form a 95% percentile confidence interval. Two details in step 2 carry weight. **With replacement** is essential: sampling without replacement from your own sample of size n just reorders it, every replicate is identical, and the estimated spread collapses to zero. **Size n** is essential too: the spread of a sampling distribution depends on how many observations went into the estimate, roughly shrinking like 1/sqrt(n) for well-behaved statistics. A resample of size n/2 would produce a distribution that is too wide and an interval that over-covers. ## Reading an interval off the replicates The percentile interval is the simplest recipe: sort the B replicates and take the alpha/2 and 1 - alpha/2 quantiles. For a 95% interval that is the 2.5th and 97.5th. It is attractive because it respects the parameter's natural range and any monotone transformation of it. A correlation interval computed this way cannot fall outside the interval from -1 to 1, because no replicate can. Other recipes exist and adjust these endpoints for bias and skew, but the percentile interval is the one to be able to describe on demand. ## What it does not buy you The most common misconception is that resampling manufactures information. It does not. Every replicate is built from the same n observations. Raising B from 1,000 to 100,000 makes the reported endpoints more stable across reruns with different seeds, and nothing else: it reduces **Monte Carlo noise** in the procedure, not **statistical uncertainty** about the population, which is fixed by n and by the data-generating process. Similarly, the bootstrap inherits every defect of the sample. If the sampling frame missed a segment of the population, every single resample misses it too, and the interval will be tight around the wrong value. Selection bias, measurement bias and non-response are fixed by design, weighting or better data collection, never by resampling. Dependence is the third trap. The ordinary bootstrap assumes the units you resample are independent. If rows are grouped, repeated or ordered in time, resampling rows treats correlated observations as if they were independent and produces intervals that are far too narrow. The repair is to resample the unit that actually is independent, whole groups or blocks rather than single rows. ## Practical notes Record and report the seed and the value of B so the numbers reproduce exactly. Look at the histogram of the replicates rather than only at two percentiles: a distribution with visible spikes, atoms sitting on individual data values, or heavy skew is telling you that either the statistic or the sample size is a poor fit for the ordinary bootstrap.

  • Why must each bootstrap resample be the same size as the original sample?
    Because the spread of a sampling distribution depends on sample size, shrinking roughly like 1/sqrt(n) for well-behaved statistics. Resampling m values with m smaller than n gives a distribution that is too wide and an interval that over-covers; m larger than n gives one that is too narrow. Deliberately choosing m much smaller than n is a specialised repair for cases where the ordinary bootstrap is inconsistent, not a default.
  • How many resamples should you draw, and what changes if you draw 1,000 instead of 10,000?
    B controls only simulation noise in the answer, never statistical uncertainty about the population, which is fixed by n. A thousand replicates is usually enough for a standard error, but interval endpoints sit out in the tails where estimates are noisiest, so 10,000 or more is the common choice for a 95% interval. Fix and record the seed so the endpoints reproduce.
  • Can the bootstrap rescue a sample you know is unrepresentative?
    No. It resamples the empirical distribution, so it inherits every bias in the data. What it estimates is how much the estimate would move under repeated sampling from a population that looks like your sample. If the sampling frame missed a segment, every resample misses it too, and you get a tight interval around the wrong value. Selection and measurement bias are design problems, not resampling problems.

You have one photograph of a crowd and want to know how much a head count would vary. The bootstrap builds new crowds by drawing faces from the photo at random with repeats allowed, counts each one, and treats the spread of counts as the wobble you would have seen across real crowds.

saying these in an interview costs you the question

  • Says resampling creates new data or new information
  • Draws resamples without replacement, which merely reorders the sample
  • Uses a resample size different from n as the default
  • Claims the bootstrap corrects a biased or unrepresentative sample
  • Thinks more resamples make the confidence interval narrower

context

open as a page

If you run 20 independent hypothesis tests at alpha = 0.05, how likely is at least one false positive?

level: juniorimportance: must knowfreq 76%

basics

~20 s

About 64 percent. Each test independently avoids a false positive with probability 0.95, so all 20 stay clean with probability 0.95^20 = 0.36, leaving roughly a 64 percent chance that at least one comes back significant by luck.

open as a page

Model B beats model A by 0.3 accuracy points on a 1,000-example test set - is that real?

level: juniorimportance: must knowfreq 68%

basics

~20 s

Not on that evidence. On 1,000 examples, 0.3 accuracy points is three examples, and a net of three can never reach statistical significance in a paired comparison. The gap sits inside the test set's noise.

open as a page

How does a permutation test build a null distribution, and what does exchangeability require?

level: middleimportance: must knowfreq 55%

basics

~20 s

A permutation test shuffles the group labels thousands of times, recomputing the test statistic on each shuffle to build the null distribution from the data itself. The p-value is the share of shuffles at least as extreme as the observed statistic.

open as a page

Why is the Bonferroni correction called conservative when testing 15 endpoints at once?

level: middleimportance: must knowfreq 68%

basics

~20 s

Bonferroni tests each of the 15 endpoints at 0.05/15 = 0.0033 instead of 0.05. That keeps the family-wise error rate under 5 percent whatever the dependence between endpoints, but the tiny per-test threshold makes real effects far harder to detect.

open as a page

Why is comparing two models on one shared held-out set a paired problem, not two proportions?

level: middleimportance: must knowfreq 57%

basics

~20 s

Both models are scored on the very same examples, so their accuracy estimates are strongly positively correlated. The two-independent-proportions formula sets that correlation to zero, inflates the standard error of the difference, and wastes power.

open as a page

What is p-hacking, and why does it inflate the false-positive rate?

level: middleimportance: must knowfreq 68%

basics

~20 s

P-hacking is trying analysis variants (extra outcomes, outlier rules, subgroups, added data) and reporting only the one that reaches p < 0.05. Each attempt is a fresh chance at luck, so the true false-positive rate far exceeds 5%.

open as a page

What does McNemar's test compare when two classifiers are scored on the same test set?

level: middleimportance: should knowfreq 44%

basics

~20 s

It compares only the discordant examples: those one model gets right while the other gets them wrong, counted each way. Under the null those split like fair coin flips, so a lopsided split is the evidence.

open as a page

What is HARKing, and why does it invalidate the p-value it is reported with?

level: middleimportance: should knowfreq 45%

basics

~20 s

HARKing means Hypothesizing After the Results are Known: you find which comparison came out significant, then present it as the hypothesis you set out to test. The write-up hides the search, so the p-value is uninterpretable.

open as a page

When does the bootstrap fail, and what goes wrong for the sample maximum or for n = 5?

level: seniorimportance: should knowfreq 45%

basics

~20 s

The bootstrap fails when the sample is a poor stand-in for the population or the statistic is not smooth. No resample can exceed the observed maximum, so extremes give degenerate intervals; n = 5 gives too little to resample.

open as a page

When would you control false-discovery rate with Benjamini-Hochberg instead of family-wise error?

level: seniorimportance: should knowfreq 52%

basics

~20 s

Use false-discovery-rate control when you are screening thousands of hypotheses and want a bounded proportion of false hits among your discoveries rather than near-certainty of none. Family-wise control is right when a single false claim is costly.

open as a page

How do you build a paired permutation test for the log-loss gap between two models?

level: seniorimportance: should knowfreq 36%

basics

~20 s

Take the per-example loss differences on the shared test set. The null makes the two models interchangeable, so flip each difference's sign at random thousands of times and count how often the resampled mean is as extreme.

open as a page

What is the garden of forking paths, and how does it differ from deliberate p-hacking?

level: seniorimportance: should knowfreq 38%

basics

~20 s

The garden of forking paths (Gelman and Loken) is the point that one analysis can be biased even when only one test is run: choices made after seeing the data mean the analyses you would otherwise have run still count.

open as a page

How do you decide which set of tests counts as one family needing multiplicity correction?

level: principalimportance: should knowfreq 38%

basics

~20 s

Define the family as the set of tests over which you claim an error guarantee, fix it in the analysis plan before seeing results, and keep it as small as the claim allows. No formula supplies it.

open as a page

How would you preregister analyses on a team that must still explore data freely?

level: principalimportance: should knowfreq 32%

basics

~20 s

Do not ban exploration, separate it. Require a time-stamped plan naming the primary outcome, exclusions and model before outcome data is inspected; treat everything else as labelled exploratory work, and confirm on data that had no role in choosing it.

open as a page

What is the jackknife, and why does it break down for the sample median?

level: middleimportance: nice to knowfreq 26%

basics

~20 s

The jackknife recomputes a statistic n times, leaving out one observation each time, and uses the spread of those values to estimate bias and standard error. It fails for non-smooth statistics such as the median.

open as a page

How does the Holm-Bonferroni step-down procedure improve on plain Bonferroni?

level: middleimportance: nice to knowfreq 31%

basics

~20 s

Holm sorts the p-values smallest first and compares the i-th to alpha/(m - i + 1), stopping at the first failure. Thresholds loosen as you go, so Holm rejects at least as much as Bonferroni under the same guarantee.

open as a page

How does a BCa bootstrap interval differ from a percentile interval, and when is it worth it?

level: seniorimportance: nice to knowfreq 32%

basics

~20 s

A percentile interval reads the 2.5th and 97.5th percentiles of the bootstrap distribution as they are. BCa shifts which percentiles it reads, correcting for an off-centre distribution and a standard error that moves with the parameter.

open as a page

How do you put a confidence interval on the AUC difference between two models?

level: seniorimportance: nice to knowfreq 23%

basics

~20 s

Resample the test examples with replacement, recompute both models' AUC on each identical resample, and take percentiles of the differences. DeLong's method is the analytic alternative, giving a closed-form variance for the difference between two correlated AUCs.

open as a page

How does the file-drawer problem bias a published literature, and can you detect it?

level: seniorimportance: nice to knowfreq 26%

basics

~20 s

Null results stay unpublished in the file drawer, so the visible literature is selected on significance: too many positives, and effects too large. A funnel plot of effect size against precision exposes it: missing small null studies appear as asymmetry.

open as a page