Under a true null hypothesis, what distribution do p-values follow?
answer
- simulate pure noise 10,000 times and plot
- the histogram has no shape at all
- push a variable through its own CDF
- P(p below 0.05) equals 0.05 exactly
- flat on the interval 0 to 1
basics
~20 sWhen the null hypothesis is true and the test statistic is continuous, p-values are uniformly distributed between 0 and 1. Every interval of equal width is equally likely, so about 5% of null results land below 0.05 by construction.
solid answer
~50 sFor a correctly specified test with a continuous test statistic, the p-value under a true null is `Uniform(0, 1)`. The reason is the probability-integral transform: the p-value is the null cumulative distribution function evaluated at the observed statistic, and feeding a random variable through its own CDF yields a uniform variable. The practical consequence is direct — under the null, `P(p <= 0.05) = 0.05`, which is exactly why 0.05 behaves as a long-run error rate. Simulate 10,000 datasets with no effect at all and the histogram of p-values is flat, with roughly 500 of them below 0.05. Two caveats: discrete test statistics give a lumpy, conservative distribution rather than an exact uniform, and if assumptions are violated the null distribution is wrong, so the p-values are not uniform even when nothing real is going on.
go deeper
Remember the headline fact: with no real effect, p-values are spread evenly between 0 and 1, so about one null result in twenty falls below 0.05 purely by chance.
Explain the mechanism — the p-value is the null CDF evaluated at the statistic, and a continuous variable pushed through its own CDF is uniform — and derive the 5% consequence from it.
Use it as a tool: simulate or permute under the null, run the whole pipeline, and read the histogram's shape as a diagnosis of dependence, variance mis-specification, or leakage.
Make null simulation a standing validation step before an analysis pipeline is trusted for decisions, and set the expectation that a calibration check precedes any argument about thresholds.
## The claim Run a test on data generated with **no effect whatsoever**, record the p-value, and repeat 10,000 times. Plot the histogram. It is **flat**: roughly the same count in the bin from 0.00 to 0.05 as in the bin from 0.90 to 0.95. That flatness is the uniform distribution, and it is not an empirical accident — it follows from the definition of the p-value. ## Why it is uniform Let `T` be the test statistic and `F` the cumulative distribution function of `T` **under the null**. For a one-sided test in the upper tail, the p-value is ``` p = 1 - F(T) ``` The **probability-integral transform** says that if `T` is a continuous random variable with CDF `F`, then `F(T)` is uniformly distributed on the interval from 0 to 1. Since `1 - U` is uniform whenever `U` is, `p` is uniform too. The same argument works for two-sided tests, where the p-value is the total tail mass beyond the observed statistic in both directions. The crucial condition is that `F` must be the *correct* null distribution of `T`. Uniformity is a property of a well-specified test, not a universal law about p-values. ## Why it matters **It is what makes a threshold mean anything.** If null p-values are uniform, then `P(p <= 0.05 | H0) = 0.05` exactly. The rate at which a correct test produces small p-values on pure noise is known in advance and equals the threshold. Without uniformity, a threshold has no calibrated long-run meaning at all. **It gives you a diagnostic.** Uniformity is testable. Simulate under the null — permute labels, or generate synthetic data with the effect switched off — run the full analysis pipeline end to end, and look at the histogram. This catches an enormous class of bugs and assumption failures: - A histogram **skewed toward zero** means the test fires too readily on noise. Typical causes: dependence between observations treated as independent (clustered users, repeated measurements, time-correlated data), a variance estimate that is too small, or a leak between the units being compared. - A histogram **skewed toward one** means the test is conservative: it rarely produces small p-values even when it should. Discrete statistics with few possible outcomes do this naturally, as do some overly cautious variance estimators. - A **spike just below 0.05**, in a collection of *reported* results rather than simulated ones, is a fingerprint of selection: analyses were adjusted, subsets chosen, or results published preferentially until the number crossed the conventional line. Under a true null the density near 0.049 and near 0.051 is identical, so a discontinuity at the threshold cannot be produced by the test itself. ## The caveats a good answer includes **Discrete statistics.** With a discrete test statistic — small counts, exact tests on tables — the p-value can only take finitely many values, so it cannot be exactly uniform. The distribution is stochastically larger than uniform, meaning the test is conservative: the actual rate of p-values below 0.05 is at most 0.05, often noticeably less. **Composite nulls and approximations.** When the null does not pin down every parameter, or when the null distribution is an asymptotic approximation valid only for large samples, uniformity holds only approximately. In small samples the approximation can be poor in exactly the tail you care about. **Violated assumptions.** This is the practically important one. If observations are correlated but the test assumes independence, the true null distribution of the statistic is wider than the assumed one, the assumed CDF is wrong, and null p-values pile up near zero. The test then produces "significant" results on data containing no effect at all — and no amount of care in interpreting a single p-value will reveal it. Only the distribution of many null p-values will. ## How to say it in an interview State the result (uniform on 0 to 1 under a true null with a continuous statistic), name the mechanism (probability-integral transform: pushing a continuous variable through its own CDF gives a uniform), give the consequence (`P(p <= 0.05) = 0.05`, which is what makes a threshold calibrated), and then offer the diagnostic use — simulating under the null and inspecting the histogram is one of the most effective ways to validate an analysis pipeline before it is trusted.
- You simulate under the null and the p-value histogram piles up near zero. What do you suspect?That the assumed null distribution is wrong, so the test fires too readily on noise. The usual culprits are dependence between observations analysed as independent — clustered or repeated measurements — an understated variance estimate, or leakage between the units being compared. The fix is in the model or the pipeline, not in the threshold.
- Why can a p-value from a discrete test statistic not be exactly uniform?Because the statistic takes only finitely many values, so the p-value does too, and a distribution supported on a finite set cannot be continuous uniform. The resulting distribution is conservative: the probability of falling below a given cutoff is at most that cutoff, usually less, so such tests reject less often than the nominal rate suggests.
- In a histogram of many published p-values, what would a spike just under 0.05 indicate?Selection rather than nature. Under a true null the density is identical on either side of 0.05, and no genuine effect creates a discontinuity precisely at a conventional cutoff, so a spike points to analyses being adjusted or results being reported preferentially until the number crossed the line.
saying these in an interview costs you the question
- Says null p-values cluster near 1 because the null is true
- Claims p-values are normally distributed under the null
- Thinks uniformity holds regardless of model assumptions
- Cannot explain why 5% of null results fall below 0.05
- Treats a spike below the 0.05 line as a real effect