skip to content

How does a permutation test build a null distribution, and what does exchangeability require?

level: middleimportance: must knowfreq 55%

answer

  1. the null says the label is arbitrary
  2. pool the groups, then reshuffle labels
  3. recompute the statistic on every shuffle
  4. tail count over shuffles gives the p-value
  5. structure must survive the shuffle

basics

~20 s

A permutation test shuffles the group labels thousands of times, recomputing the test statistic on each shuffle to build the null distribution from the data itself. The p-value is the share of shuffles at least as extreme as the observed statistic.

solid answer

~40 s

Under the null hypothesis that the label carries no information, every assignment of the observed values to the two groups is equally likely. That is exchangeability. So you compute the observed statistic, say the difference in group means, then pool all observations, reshuffle the labels 10,000 times, and recompute the statistic on each shuffle. Those 10,000 values are the null distribution, built with no normality and no large-sample assumption. The two-sided p-value is the fraction of shuffles whose statistic is at least as extreme as the observed one, conventionally reported as `(1 + count) / (B + 1)` so it can never come out exactly zero. Exchangeability is the load-bearing assumption: with paired, clustered or time-ordered data a free shuffle destroys structure that the null does not, and the test stops being valid.

go deeper

for a junior

Be ready to describe the loop out loud: pool the observations, reassign the labels at random, recompute the statistic, repeat thousands of times, then see where the observed value sits in the pile you collected.

for a middle

Explain why the procedure needs no normality assumption, what exchangeability under the null means in plain words, and why the p-value convention adds one to both the tail count and the number of shuffles.

for a senior

Show judgment about structure. Paired, clustered and time-ordered data require permuting within the design rather than across it, and unequal variances with unequal group sizes call for a studentised statistic before shuffling.

for a principal

Own the tradeoff between a resampling test that costs compute and buys assumption-freedom and a parametric one that is cheap and familiar, and set which the organisation reaches for by default and when it escalates.

## The idea A hypothesis test needs a null distribution: the set of values the test statistic could plausibly take if the effect you are looking for were absent. Classical tests get one from theory, at the price of assumptions about the shape of the data or about how large the sample is. A permutation test gets one by brute force, from the data in front of you. The null hypothesis is that the group label is arbitrary, that the observations would have looked the same had the labels been handed out differently. If that is true, the particular labelling you observed is just one of many equally likely labellings. So generate the others and see where your observed statistic falls among them. ## The procedure 1. Compute the statistic of interest on the data as labelled. For two groups this is typically the difference in means, but it can be a difference in medians, in trimmed means, or any other comparison, which is part of the appeal. 2. Pool all observations, discarding the labels. 3. Randomly reassign the labels, keeping the original group sizes, and recompute the statistic. Store it. 4. Repeat step 3 B times, commonly 10,000. 5. Count how many of the B replicate statistics are at least as extreme as the observed one. For a two-sided test, compare absolute values against the observed absolute value. 6. Report the p-value as (1 + count) / (B + 1). The plus-one convention exists because the observed labelling is itself one valid permutation and belongs in the reference set. It also keeps the reported value away from an impossible exact zero: with B = 10,000 the smallest value you can report is about 1/10,001. ## Exchangeability, precisely A set of random variables is **exchangeable** if their joint distribution is unchanged by any reordering of them. In a permutation test the requirement is narrower and conditional: under the null hypothesis, the operation you apply, shuffling labels across the pooled observations, must leave the joint distribution of the data unchanged. That has three practical consequences. First, the null being tested is stricter than most people say out loud. For a two-sample label shuffle, the null that makes labels exchangeable is that the two samples come from the **same distribution**, not merely that they have the same mean. If the two groups genuinely share a mean but differ in variance, and the group sizes are unequal, the shuffled null distribution has the wrong variance and the type I error rate drifts away from nominal. Studentising the statistic, dividing the mean difference by an estimate of its standard error before permuting, largely repairs this. Second, structure in the data restricts which shuffles are legal. Matched pairs are not freely exchangeable: swapping a label across pairs breaks the very link the design was built around. The exchangeable operation for paired data is flipping the sign of each pair's difference independently, each with probability one half. For clustered data, whole clusters are exchangeable while individual rows are not, so you permute cluster labels. For time series, a free shuffle destroys autocorrelation that exists under the null as well as under the alternative, so block-based schemes are needed. Third, exchangeability is an assumption about the null, not a property you can verify from the data. It comes from knowing how the data were generated: whether assignment was random, whether observations are nested, whether time matters. ## Exact versus Monte Carlo With small samples you can enumerate every possible labelling. Then the test is genuinely **exact**: the p-value is a precise combinatorial probability, valid at any sample size, with no approximation anywhere. Enumeration becomes infeasible quickly, so in practice you sample B labellings at random. The result is a Monte Carlo approximation to the exact test. It remains a valid test for any B, because the (1 + count) / (B + 1) construction accounts for the sampling, but the number carries simulation noise. A p-value sitting near a decision threshold can flip between seeds, which is a signal to raise B rather than to pick the seed you like. ## What it buys and what it costs The payoff is freedom from distributional assumptions and the ability to test statistics that have no tractable theory, along with genuine exactness in the enumeration case. The costs are compute, the need to think carefully about which shuffles are legal, and a common misreading: a permutation test gives you a p-value, not an interval or an effect size. Pair it with a separately computed estimate and interval when the size of the effect matters, which it usually does. ## Common failures Shuffling when the data are paired or clustered, and thereby answering a different question than intended. Reporting p = 0 when no shuffle beat the observation. Assuming a permutation test needs normal data, which is precisely the assumption it removes. And treating a permutation test as more powerful in principle than a parametric one: when the parametric assumptions hold, the two agree closely, and the resampling version buys robustness rather than power.

  • What exactly does a two-sample permutation test of a mean difference assume under the null?
    The null that makes labels exchangeable is that both samples come from the same distribution, which is stronger than equal means. If the groups share a mean but differ in spread and the group sizes are unequal, shuffling produces a null distribution with the wrong variance and the type I error rate drifts off nominal. Studentising the statistic before permuting largely repairs this.
  • How do you run a permutation test when the observations are matched pairs rather than independent groups?
    Permute within the structure rather than across it. For matched pairs the operation that is exchangeable under the null is flipping the sign of each pair's difference, each pair independently and with probability one half. Freely shuffling all labels would break the pairing and answer a different question. The same principle governs clustered data: permute whole clusters, because only clusters are exchangeable.
  • Is a permutation test with 10,000 random shuffles exact?
    Enumerating every possible labelling gives an exact test. Sampling 10,000 shuffles is a Monte Carlo approximation to it. The test stays valid for any B because the plus-one convention accounts for the sampling, but the reported number carries simulation noise, so a p-value near the decision threshold can move between seeds. Raise B when the answer is borderline rather than choosing a seed.

saying these in an interview costs you the question

  • Says a permutation test requires normally distributed data
  • Freely shuffles labels when the data are paired or clustered
  • States the null as equal means rather than identical distributions
  • Reports p equals zero when no shuffle beat the observation
  • Thinks more shuffles can turn a weak effect significant

context