A KS test on 1,000,000 sessions returns p below 1e-16 for a shift nobody would notice - how do you decide whether to act?
answer
- significance is not magnitude
- the threshold shrinks like one over root n
- exact equality is never literally true
- D is already a 0-to-1 effect size
- set the materiality bar before looking
basics
~20 sAt a million observations the KS rejection threshold shrinks toward zero, so any real difference becomes significant. Judge the magnitude instead: read the statistic D as an effect size and compare it against a materiality threshold agreed beforehand.
solid answer
~50 sA tiny p-value at that sample size carries almost no information. The 5% threshold for the two-sample statistic is roughly `1.36 * sqrt((n + m) / (n * m))`, which shrinks like one over the square root of sample size - at a million per group it sits near 0.002, so a two-tenths-of-a-percentage-point gap in cumulative proportion clears it. The test is answering 'are these two distributions **exactly** identical?', and the honest answer to that is almost always no. So switch instruments: `D` itself is a scale-free effect size between 0 and 1, the largest difference in the fraction of sessions at or below any value. Quote it, plot the two empirical CDFs, and read off the difference at the thresholds the business actually cares about. Then compare against a materiality threshold fixed **before** the analysis, derived from what magnitude would change the decision. If `D` is 0.003 and the decision needs 3 percentage points, the finding is real and irrelevant, and you say so.
go deeper
Remember that a very small p-value means the difference is unlikely to be chance, not that it is big. With huge samples, tiny differences reach significance routinely.
Explain the mechanism: the rejection threshold for the maximum-gap statistic falls like one over the square root of sample size, so at a million observations a gap of a fraction of a percentage point clears it.
Show the working habit - report the statistic as an effect size with the value where it occurred, overlay the two empirical CDFs, and quote the difference at the cutoffs tied to the SLO or the decision.
Own the policy. Insist that every comparison ships with an effect size and a materiality threshold agreed in advance, and push back on sub-sampling and on significance-only alerting across many large-table metrics.
## Why the p-value stopped being informative A p-value answers one question: how surprising is data this extreme if the two distributions were identical? At a million observations per group, that question has almost no practical content, because two real-world populations are essentially never *identical*. Different weeks, different traffic mixes, different device populations - some difference always exists. With enough data the test finds it. The arithmetic is explicit. The large-sample 5% rejection rule for the two-sample maximum-gap statistic is ``` D > 1.36 * sqrt( (n + m) / (n * m) ) ``` With `n = m = 1,000,000` that threshold is about 0.0019. A gap of two-tenths of a percentage point in cumulative proportion - nothing anyone would design a change around - is enough to reject. Push to ten million and the threshold falls by another factor of about three. **Statistical significance is a statement about sample size at least as much as about the world.** ## Switch to the magnitude The statistic you already computed is the effect size. `D` lives on 0 to 1 and reads in plain language: the largest difference, at any cutoff, between the fraction of sessions in group A at or below that cutoff and the fraction in group B. `D = 0.003` means that at the worst point, the two builds disagree by three sessions in a thousand. That sentence is auditable by a product owner in a way that 'p < 1e-16' is not. Three things to put in the report: 1. **`D` with the value at which it occurred.** 'The biggest divergence is 0.003, at around 240 ms.' 2. **The two empirical CDFs overlaid.** With a million points the curves are smooth, and a reader can see instantly whether they are essentially the same line or genuinely separated. 3. **Differences at the decision-relevant cutoffs.** If the SLO is 'under 500 ms', the number that matters is the change in the fraction of sessions above 500 ms, not the maximum gap wherever it happens to fall. ## Pre-commit to a materiality threshold The defensible way to avoid arguing about a result after seeing it is to fix the bar before. Write down, in advance: what magnitude of change would alter the decision we are about to make? That number comes from the economics, not from statistics - the cost of rolling back, the revenue attached to a percentage point, the risk tolerance of the SLO. Once that threshold exists, the analysis becomes a comparison of two numbers rather than a debate about a p-value. You can go further and invert the test into an **equivalence** framing: instead of asking whether any difference exists, ask whether the data lets you rule out a difference larger than the threshold. That flips the burden of proof to where the decision actually needs it - a large p-value from a test of exact equality is not evidence of sameness, whereas an equivalence conclusion is. ## The move that looks clever and is not Someone will suggest sub-sampling to a few thousand rows until the p-value 'behaves'. Refuse it. That does not make the difference smaller, it makes the estimate of the difference noisier. You have discarded information in order to lose the ability to detect the very thing you were measuring, and the p-value you produce is a function of an arbitrary sub-sample size you chose. If the p-value is embarrassing, the fix is to stop using the p-value as the decision rule, not to degrade the data until it agrees with you. The mirror-image error is silent multiplicity: running the test across dozens of metrics on huge tables and reporting whichever ones came back significant. At this sample size, essentially all of them will. ## What to tell stakeholders The useful sentence has this shape: *'The distributions are statistically distinguishable - at this volume they always will be - and the largest difference anywhere is 0.3 percentage points, well under the 3 points we agreed would matter. Recommendation: no action.'* It reports the result honestly, defuses the tiny p-value, and connects to a pre-agreed bar. The converse case matters just as much. If `D` is 0.14 with a gap concentrated beyond the 95th percentile, the same discipline says act, and act on the tail - even though the p-value looks no different from the trivial case. ## The organisational lesson At large scale, hypothesis tests of exact equality stop being decision procedures and become sample-size detectors. A team that keeps them as its default alerting rule will drown in true-but-irrelevant findings and gradually learn to ignore all of them, including the important ones. The durable fix is cultural: **every comparison ships with an effect size and a pre-agreed threshold for what counts as material**, and significance is at best a gate, never the verdict.
- What does the statistic itself tell you that the p-value does not?It is the magnitude. The maximum-gap statistic sits between 0 and 1 and reads as the largest difference in the fraction of observations at or below any cutoff - three in a thousand, or fourteen in a hundred. That number is comparable across studies and independent of sample size, which is exactly what the p-value is not.
- Would sub-sampling down to a few thousand rows fix the problem?No. It leaves the true difference untouched and only makes your estimate of it noisier, while making the p-value a function of an arbitrary sub-sample size you chose. It trades a precision problem for a power problem. Keep all the data and judge the magnitude instead.
- How do you set a materiality threshold defensibly?Derive it from the decision, before looking at results: what change in the metric would alter what we do, given the cost of acting and the value attached to a percentage point? Write it down with the owner who would act on it, so the analysis becomes a comparison against an agreed bar rather than an after-the-fact argument.
- Is a large p-value at this sample size evidence that the distributions are the same?Not on its own, but at this volume it is genuinely informative: the test could have detected minute differences and did not. The clean way to make that claim is an equivalence framing - show the data rules out any difference larger than the agreed threshold - rather than presenting failure to reject as proof of sameness.
A scale accurate to a milligram will tell you two supposedly identical bags of flour differ. It is right, and it is useless for deciding whether to reject the shipment - for that you need the tolerance written on the contract.
saying these in an interview costs you the question
- Treats a tiny p-value as evidence of a large difference
- Sub-samples the data until the p-value is non-significant
- Reports significance without ever quoting the effect size
- Sets the materiality threshold after seeing the result
- Runs the test across dozens of metrics and reports the hits