Why can a negligible effect give p = 0.4 at n = 100 but p = 0.001 at n = 1,000,000?
answer
- the statistic has a numerator and a denominator
- only one of the two depends on n
- standard error shrinks like one over root n
- distinguishable from zero, not big
- nothing is exactly zero at a million rows
basics
~20 sA p-value blends how big an effect is with how precisely it was measured. Standard error shrinks roughly as one over the square root of the sample size, so at a million observations even a trivial difference becomes statistically detectable.
solid answer
~50 sThe p-value is driven by the test statistic, which for a mean difference is roughly `(estimated difference) / SE`, and the standard error scales like `s / sqrt(n)`. Hold the true difference fixed and multiply n by 10,000: the standard error falls by a factor of 100, the statistic grows by the same factor, and the tail probability collapses. So a difference too small to matter to anyone can produce p = 0.001 with enough data, while the identical difference sits at p = 0.4 in a small sample. The right reading is that the p-value answers "is this distinguishable from exactly zero?", not "is this big enough to act on?". At very large n almost nothing is exactly zero, so significance stops being informative on its own and the estimated magnitude in real units, with an interval around it, becomes the quantity that carries the decision.
go deeper
Know that p-values depend on sample size as well as on the size of the difference, and that a very small p-value does not by itself mean the difference is worth acting on.
Be able to write the statistic as an estimate divided by a standard error, state that the standard error shrinks like one over the square root of n, and derive the consequence out loud.
Show how you would stop this from misleading a stakeholder: lead with the estimate and its interval in real units, and set the magnitude that matters before the data arrive.
Own the reporting standard for high-traffic experiments, where significance is nearly automatic. Decide what magnitude threshold a result must clear and make that, not the p-value, the thing teams argue about.
## The mechanics For a comparison of two group means, the test statistic has the shape ``` statistic = (estimated difference) / SE ``` and the standard error of a mean behaves like ``` SE = s / sqrt(n) ``` where `s` is the sample standard deviation. Two things follow immediately. **The numerator is about the world.** It estimates how far apart the groups actually are. Collecting more data does not systematically push it up or down; it just estimates the same underlying quantity more stably. **The denominator is about the measurement.** It shrinks with the square root of the sample size. Going from n = 100 to n = 1,000,000 is a 10,000-fold increase in sample size and therefore roughly a 100-fold shrink in the standard error. So the statistic grows in proportion to `sqrt(n)` for a fixed true difference, and the p-value — a tail area beyond that statistic — falls away rapidly. A difference that yields a statistic of about 0.8 at n = 100 (p near 0.4) yields a statistic near 80 at n = 1,000,000, and the tail area is astronomically small. ## What this means for interpretation The p-value answers a very narrow question: **are these data hard to reconcile with the hypothesis that the difference is exactly zero?** In most real settings, nothing is exactly zero. Two checkout flows, two email subject lines, two model variants — they almost never produce *identical* population means to infinite decimal places. Given enough data, the test will eventually notice the difference, however tiny, and return a small p-value. This produces the two symmetric failure modes: - **Small n, real effect, large p.** A difference that genuinely matters can sit at p = 0.4 simply because the sample cannot resolve it. Reading that as "no effect" throws away a real finding. - **Huge n, trivial effect, tiny p.** A difference of 0.02% in conversion can arrive at p = 0.001 and be announced as a "highly significant win". Reading that as "an important effect" imports a claim the number never made. Both mistakes come from the same root: treating the p-value as if it measured *magnitude*. It does not. It measures *distinguishability from a specific null value*, and distinguishability is a joint function of magnitude and precision. ## How to report so the confusion cannot happen 1. **Lead with the estimate in the units of the problem.** "Checkout completion rose by 0.03 percentage points" is a statement a product owner can evaluate. "p = 0.001" is not. 2. **Attach an interval.** An interval estimate shows the range of differences the data are consistent with. At n = 1,000,000 that interval is narrow — and if the whole narrow interval lies inside the range of differences nobody cares about, the tiny p-value is irrelevant to the decision. 3. **Decide in advance what magnitude would change behaviour.** If a lift below 0.5 percentage points would not justify a rollout, say so before the analysis. Then a statistically distinguishable 0.03-point lift is simply a precisely measured non-event, not a win. 4. **Report the exact p-value.** "p < 0.05" hides the difference between borderline and overwhelming. ## The reverse direction: big p in a small sample The symmetric habit is just as important. When a small study returns a large p-value, the honest summary is "this sample cannot separate the effect from zero", and the interval will usually be wide enough to contain differences that would be commercially decisive. Announcing "no effect" collapses a wide interval into a point claim the data never supported. ## The interview shape Interviewers like this question because it exposes whether a candidate has internalised the composition of the test statistic. A strong answer names the `sqrt(n)` scaling explicitly, shows the numerator/denominator split, and then draws the practical conclusion: at very large sample sizes, significance is nearly free and the estimate's magnitude becomes the only quantity that carries information for a decision. A weaker answer says "more data means more power" and stops, which is the observation without the mechanism.
- Roughly how does the test statistic change if you quadruple the sample size, holding the true difference fixed?It roughly doubles. The standard error scales like `s / sqrt(n)`, so quadrupling n halves the standard error, and halving the denominator of `(estimated difference) / SE` doubles the statistic. The p-value drops correspondingly, even though the quantity being estimated has not changed at all.
- At n = 1,000,000, what would you report instead of the p-value?The estimated difference in the units the business uses, with an interval showing the range of values the data are consistent with, and an explicit comparison against the magnitude that was agreed beforehand to be worth acting on. At that sample size the interval is narrow, so it usually settles the decision on its own and the p-value adds nothing.
- Does a large p-value in a small study mean the effect is small?No. It means the data cannot separate the effect from zero at that sample size. The interval around the estimate will typically be wide and may include differences large enough to matter, so the honest summary is that the study is uninformative about magnitude rather than that the effect is negligible.
A stronger microscope does not make a speck of dust bigger; it just makes the speck impossible to miss. Sample size is the magnification, not the size of the thing.
saying these in an interview costs you the question
- Treats a smaller p-value as evidence of a bigger effect
- Says more data makes an effect stronger rather than better measured
- Announces a trivial difference as important because p is tiny
- Reads a large p in a small sample as proof of no effect
- Cannot name what sits in the denominator of the test statistic
- Reports only p < 0.05 with no estimate in real units