skip to content

Two studies report identical effect estimates with p = 0.049 and p = 0.051 — how should a team treat them?

level: principalimportance: should knowfreq 36%

answer

  1. evidence is continuous, the line is a convention
  2. 0.002 apart with identical estimates
  3. the label destroys what the number carried
  4. pre-commit the rule, report the exact value
  5. borderline is itself a legitimate outcome

basics

~10 s

Two p-values 0.002 apart carry indistinguishable evidence. With identical estimates, any qualitative gap between 0.049 and 0.051 comes from the reporting convention, not from the data. Report magnitudes and uncertainty, not a significance label.

solid answer

~40 s

Evidence is continuous; the 0.05 line is a convention laid over it. Two studies with identical estimates and p-values of 0.049 and 0.051 carry indistinguishable information, so writing one up as a discovery and the other as a null finding manufactures a difference the data never contained. What I want from a team is: report the estimate in real units with an interval and the exact p-value, never as `p < 0.05`; fix the decision rule before the data are seen so a borderline number cannot be renegotiated afterwards; and let the size of the estimate relative to what matters drive the write-up. If a decision must be binary, the honest framing for a borderline result is that it is borderline, and the useful next step is usually more data.

go deeper

for a junior

Understand that 0.049 and 0.051 carry nearly identical evidence, and that a threshold is a reporting convention rather than a boundary in the data.

for a middle

Be able to explain that the p-value is a continuous measure and to report the exact value alongside the estimate instead of a significant or not-significant tag.

for a senior

Show how you keep a borderline result honest in practice: exact values, an interval, a rule fixed before the data, and a write-up that names the result as unresolved.

for a principal

Own the reporting and decision culture — remove the binary label from templates, require pre-committed rules, and make sure no launch narrative can rest on which side of a line a number happened to fall.

## The arithmetic first Two studies, same design, same estimated effect, same precision to a rounding error. One returns p = 0.049, the other p = 0.051. The difference in the underlying evidence is negligible — the test statistics differ in the third decimal place. Yet under a strict threshold convention, the first is written up as a positive finding and the second as a failure, and those two write-ups travel through an organisation as if they said opposite things. Nothing in the data justifies that. The p-value is a **continuous** measure of incompatibility with the null, and 0.05 is a **convention** — a widely-adopted round number, not a property of nature. There is no discontinuity in the world at 0.05. ## Why the dichotomy is seductive anyway Thresholds exist for a reason: decisions are often binary (ship or do not ship), and a pre-committed rule stops results being reinterpreted after the fact to suit whoever is arguing. That value is real, and the answer is not to abolish rules. The failure is in **collapsing the report to the label**. Once a result is stored as "significant" or "not significant", the magnitude, the precision and the closeness to the line are all thrown away, and downstream readers cannot recover them. The collapse also creates bad incentives. When a career, a launch or a quarterly narrative hangs on which side of a line a number falls, the pressure to nudge the analysis until it crosses becomes enormous — dropping a subgroup, changing an outlier rule, switching the metric definition. The 0.049 study and the 0.051 study are indistinguishable on the evidence, but only one of them faces that pressure, and the resulting distortion is invisible in the final report. ## What a lead should actually institute **1. Report exact p-values, never inequalities.** `p = 0.051` and `p = 0.049` are informative; `p < 0.05` and `n.s.` destroy the information that they were nearly the same. A number is cheap to print and impossible to misread as a category. **2. Lead with the estimate and its uncertainty.** The write-up should open with the change in the units of the business and the range of values the data are consistent with. If two studies show the same estimate with the same precision, that identity is visible immediately, and no reader can be misled by the labels. **3. Fix the rule before the data arrive.** Write down what will be measured, how, and what will be done at each outcome, *before* seeing results. This is what makes a threshold honest: it can no longer be argued about once the number is known. A borderline result then does not get renegotiated; it gets reported as borderline. **4. Treat borderline as its own outcome.** In many settings the useful response to a near-threshold result is neither "ship" nor "abandon" but "this is not resolved" — extend the study, replicate it, or accept the decision on other grounds and say so openly. Pretending the number settled the question when it plainly did not is the actual failure. **5. Separate the statistical question from the decision.** Whether the data are hard to reconcile with no effect is one question. Whether the estimated change is large enough to be worth the cost and risk of acting is another, and it is usually the one that should dominate. A borderline p-value attached to an estimate far below what would justify the work is an easy call; the same p-value on a transformative estimate deserves more data, not a coin flip. ## Talking to stakeholders The hardest part is cultural rather than technical. Stakeholders often want a verdict, and "it's borderline" sounds like an analyst hedging. The framing that works is to give the decision, the reason, and the uncertainty together: "our best estimate is a 1.2-point lift, the data are consistent with anything from 0.0 to 2.4 points, and the exact p-value is 0.051 — that is the same evidence as a study reporting 0.049, so I would not treat it as a null result. I recommend running two more weeks, which would roughly halve the width of that range." Nobody in that conversation is arguing about a line. ## What an interviewer is listening for That you refuse the false dichotomy without becoming nihilistic about rules; that you can explain *why* the pre-commitment matters while still insisting the report carry the continuous evidence; and that you can name the organisational consequences — pressure to cross the line, information destroyed by labels, narratives built on a 0.002 difference. A candidate who says only "0.05 is arbitrary" has spotted the problem but not shown how to run a team around it.

  • If thresholds are conventions, why not let each analyst judge borderline cases case by case?
    Because a rule decided after the data are seen is not a rule. Case-by-case judgment invites the result to be reinterpreted to suit whoever is arguing, and the flexibility is invisible to readers. Pre-commit the rule so it constrains, and keep the continuous evidence in the report so a borderline case can still be described honestly as borderline.
  • How would you change a results template so borderline findings cannot be misread?
    Put the estimate in business units and its interval in the first line, require the exact p-value rather than an inequality or a significance flag, and add an explicit field for the decision taken and what would change it. Removing the binary label from the template is the highest-leverage edit, because it is what downstream readers copy.
  • A stakeholder asks for a yes or no on a p = 0.051 result. What do you say?
    Give the decision and the reason without pretending the data settled it: state the estimated effect and its range, note that the evidence is indistinguishable from a result reported at 0.049, and recommend either more data or a decision made on cost and strategy with the uncertainty stated openly. The answer is a recommendation, not a verdict from the number.

A pass mark of 60 makes 59 and 61 different outcomes on a transcript, but nobody believes the two students learned different amounts. The line is administrative; the knowledge is continuous.

saying these in an interview costs you the question

  • Treats 0.049 and 0.051 as qualitatively different findings
  • Reports results only as significant or not significant
  • Reruns or reshapes an analysis until the number crosses the line
  • Says 0.05 is arbitrary with no alternative practice to offer
  • Rounds 0.051 down to report it as under the threshold
  • Lets the label rather than the magnitude drive the recommendation

context