What does a p-value of 0.03 from a two-sample t-test actually mean?
answer
- which side of the conditional bar?
- a probability of data, not of a hypothesis
- everything is computed assuming the null
- at least as extreme, not exactly this
- P(data|H0), never P(H0|data)
basics
~20 sA p-value of 0.03 says that if the null hypothesis of no difference were true, data at least this extreme would arise about 3% of the time. It is not the probability that the null hypothesis is true.
solid answer
~50 sThe p-value is a tail probability computed *assuming the null hypothesis is true*. For a two-sample t-test with a null of equal means, p = 0.03 means: if the two population means really were equal, and the test's assumptions held, then in about 3 out of 100 repeated samples of this size we would see a t statistic at least as far from zero as the one we got. The conditioning runs one way only — it is `P(data at least this extreme | H0)`, never `P(H0 | data)`. Two other details matter: it uses *at least as extreme*, not the probability of the exact observed value, because any single value of a continuous statistic has essentially zero probability; and it is computed under the whole model, so a small p can also be produced by a violated assumption rather than a real difference in means.
go deeper
Be ready to state the definition in one clean sentence, with the conditioning the right way round, and to say plainly that it is not the probability the null is true.
Explain where the number comes from: a test statistic, a null sampling distribution, and a tail area. Say why 'at least as extreme' is needed rather than the exact outcome.
Show that a small p-value indicts the whole model, not just the effect. Point out that dependence between observations or a wrong variance model can manufacture small p-values from nothing.
Own how p-values are reported across a team: exact values rather than inequalities, always paired with an estimate and its uncertainty, and phrased so a non-specialist reader cannot slide into the inverted reading.
## What a p-value is A **hypothesis test** starts from a *null hypothesis* (`H0`) — a specific, fully-specified statement about the population, such as "the mean of group A equals the mean of group B". From the sample you compute a **test statistic**, a single number summarising how far the data sit from what `H0` predicts. In a two-sample t-test that statistic is roughly ``` t = (mean_A - mean_B) / SE ``` where `SE` is the standard error of the difference — the typical size of the sampling wobble in that difference. The **p-value** is then defined as: > the probability, computed *under the assumption that `H0` is true*, of obtaining a test statistic at least as extreme as the one actually observed. So p = 0.03 is a statement about how surprising the data are in a hypothetical world where the null holds. It is a probability of *data*, conditional on a *hypothesis*. ## The three parts people drop **1. "Assuming H0 is true."** The whole calculation lives inside that assumption. The p-value cannot tell you the probability that the assumption itself is true — that quantity was fixed at 1 for the purposes of the computation. **2. "At least as extreme."** For a continuous statistic, the probability of any exact value is essentially zero, so the exact result carries no usable information on its own. What is computed is a tail area: the probability mass beyond the observed statistic under the null sampling distribution. In a two-sided test, both tails count. **3. "Under the model."** The null sampling distribution comes from a set of assumptions — independence of observations, the sampling scheme, a variance model, sometimes approximate normality of the sampling distribution of the mean. A small p-value says *something in the whole package is unusual*. Usually people want that something to be the effect, but a dependence between observations, a mis-specified variance, or contaminated data can produce a small p just as well. ## The inverse-probability fallacy The single most common error is reading p = 0.03 as "there is a 3% chance the null hypothesis is true". That inverts the conditional. Written out: ``` p-value = P(data at least this extreme | H0) the claim = P(H0 | data) ``` These are different quantities and they are not numerically interchangeable. Consider a familiar analogue: `P(the animal has four legs | it is a dog)` is close to 1, while `P(it is a dog | it has four legs)` is nowhere near 1. Getting from one to the other requires how common each hypothesis was to begin with, and a p-value contains no such information — nothing in the calculation ever asked how plausible `H0` was before the data arrived. A related slip is the **complement fallacy**: "p = 0.03, so there is a 97% chance the alternative is true". Both halves are `P(H0 | data)` claims in disguise, and neither is what was computed. ## What a p-value can legitimately be used for - **As a measure of incompatibility.** Small p means the data are hard to reconcile with the null *and* the surrounding model assumptions. Large p means the data are compatible with the null — which is not the same as the null being true. - **As a continuous summary.** p = 0.001 is stronger evidence against `H0` than p = 0.04; the number carries information beyond a pass/fail label. Report the actual value, not just an inequality. - **Alongside an estimate.** The p-value says nothing about *how big* the difference is. A difference of 0.01 seconds and a difference of 10 seconds can both yield p = 0.03. Always pair the p-value with the estimated difference and an interval around it, so a reader can see the magnitude as well as the surprise. ## Phrasing that survives scrutiny Good: "If there were truly no difference between the groups, we would see a difference at least this large about 3% of the time." Bad: "There is a 3% chance the groups are the same", "There is a 97% chance our new version is better", "The result is 97% reliable", "p = 0.03, so the effect is important". Interviewers ask this question because the sloppy phrasings are everywhere in dashboards and reports, and because a candidate who states the conditional correctly usually goes on to read results correctly too.
- Why does the definition say 'at least as extreme' instead of 'exactly this result'?For a continuous test statistic, the probability of landing on any single exact value is essentially zero, so that probability would be uninformative for every dataset. Using a tail area — all outcomes as far from the null prediction as the observed one, or further — gives a quantity that varies meaningfully with how unusual the data are, and it is what makes the p-value comparable across tests.
- If a p-value is not the probability the null is true, what would you need to compute that?You would need how plausible each hypothesis was before seeing the data, plus how likely the data are under every competing hypothesis, not just under the null. A p-value uses only the null's sampling distribution, so the ingredients for the reverse conditional are simply absent from the calculation.
- A colleague reports p = 0.03 and concludes the effect is large. What do you say?That the p-value carries no magnitude information. It blends the size of the estimated difference with how precisely it was measured, so a trivially small difference measured very precisely and a large difference measured roughly can both return p = 0.03. Ask for the estimated difference in real units and an interval around it before judging whether the result matters.
Almost every dog has four legs, but most four-legged animals in the world are not dogs. A p-value is the first kind of statement; people routinely read it as the second.
saying these in an interview costs you the question
- Says p = 0.03 means a 3% chance the null is true
- Says 1 - p is the probability the alternative is true
- Calls the p-value the probability the result was due to chance
- Treats a small p-value as evidence of a large effect
- Forgets the p-value is computed assuming the null holds
- Describes the p-value as the probability of the exact observed result