skip to content

P-Hacking

Trying subgroups, stopping points and outcome definitions until something crosses 0.05 is the garden of forking paths. Preregistration and honest reporting are the answers interviewers want to hear.

on this pageshow

questions

5

What is p-hacking, and why does it inflate the false-positive rate?

level: middleimportance: must knowfreq 68%

answer

  1. many chances, one reported number
  2. researcher degrees of freedom
  3. outcome, exclusions, subgroups, sample size
  4. the procedure, not the arithmetic, breaks
  5. nominal 5% is not the real rate

basics

~20 s

P-hacking is trying analysis variants (extra outcomes, outlier rules, subgroups, added data) and reporting only the one that reaches p < 0.05. Each attempt is a fresh chance at luck, so the true false-positive rate far exceeds 5%.

solid answer

~50 s

P-hacking is exploiting the flexibility in how an analysis could be run, then reporting the version that produced significance as if it had been the plan. The usual levers are the researcher degrees of freedom catalogued in Simmons, Nelson and Simonsohn's 'False-Positive Psychology': measuring several dependent variables and reporting the one that worked, deciding outlier and exclusion rules after seeing their effect, collecting more subjects until the p-value dips below 0.05, adding or dropping covariates, and slicing subgroups. The arithmetic of any single test is untouched; what breaks is the error rate of the *procedure*. A pre-specified test at alpha = 0.05 rejects a true null 5% of the time. Give yourself four independent shots and report the best, and that rises to about 1 - 0.95^4, roughly 19%. The same authors demonstrated it by producing a straight-faced significant result that listening to 'When I'm Sixty-Four' made listeners 1.5 years younger.

code

python · 20 lines
python
import random, math

def phi(x): return 0.5 * (1 + math.erf(x / math.sqrt(2)))

def p_value(a, b, n):
    se = math.sqrt(2.0 / n)              # sd is known to be 1 in each group
    z = (sum(a) / n - sum(b) / n) / se
    return 2 * (1 - phi(abs(z)))

random.seed(0)
n, trials, preplanned, best_of_four = 20, 20000, 0, 0
for _ in range(trials):
    ps = []
    for _ in range(4):                   # four candidate outcomes, all truly null
        a = [random.gauss(0, 1) for _ in range(n)]
        b = [random.gauss(0, 1) for _ in range(n)]
        ps.append(p_value(a, b, n))
    preplanned += ps[0] < 0.05           # the outcome named in advance
    best_of_four += min(ps) < 0.05       # report whichever outcome "worked"
print(preplanned / trials, best_of_four / trials)   # ~0.05  vs  ~0.19

go deeper

for a junior

Be ready to define p-hacking in one sentence and name two concrete examples, such as measuring several outcomes and reporting the significant one, or dropping outliers after seeing the result.

for a middle

An interviewer expects the mechanism: alpha = 0.05 is a property of a fixed procedure, and taking several shots at it drives the real false-alarm rate well above 5%. Be able to do the 1 - 0.95^k arithmetic out loud.

for a senior

Show you can spot it in someone else's work: ask what was decided before the data arrived, how many outcomes were collected, and why the exclusion rule looks the way it does. Also explain why the reported effect size is inflated, not just the p-value.

for a principal

Own the framing that this is a process problem, not a character problem. Talk about how analysis plans, confirmatory-versus-exploratory labelling and holdout data change what your organisation is able to claim, and what those disciplines cost in speed.

## The idea A hypothesis test is a *procedure*, not a formula. When you say a result is significant at alpha = 0.05, you are making a claim about how often that procedure would produce a false alarm across repeated studies where nothing is going on: 5% of the time. That guarantee holds only if the whole procedure — which outcome you measure, which observations you keep, which covariates you adjust for, when you stop collecting data, which subgroup you look at — was fixed before you saw the data. P-hacking is what happens when it wasn't. The analyst tries several defensible versions of the analysis and reports the version that crossed the threshold, usually without any intent to deceive. Every individual step looks reasonable in isolation; the reported number is nevertheless wrong, because the number describes a procedure that was never actually run. ## Researcher degrees of freedom Simmons, Nelson and Simonsohn's paper 'False-Positive Psychology' named the levers and showed how ordinary they look: - **Choice of dependent variable.** Two or three plausible outcome measures were collected; the write-up names the one that reached significance. - **Exclusion and outlier rules.** "Remove responses more than 3 SD from the mean" is a fine rule — decided in advance. Decided after seeing that it moves p from 0.07 to 0.04, it is a data-dependent choice. - **Sample size decided as you go.** Running 20 subjects, checking the p-value, and adding 10 more if it is close, repeating until it crosses. - **Covariates.** Adding, dropping, or interacting a control variable and keeping the specification that worked. - **Conditions and subgroups.** Reporting the two conditions that differed out of three, or the half of the sample where the effect appeared. The paper's own demonstration is the memorable part: using nothing but these ordinary-looking choices, the authors produced a statistically significant finding that listening to 'When I'm Sixty-Four' made participants 1.5 years *younger* — an effect that is impossible by construction, obtained with an analysis that would have survived review. ## Why the error rate explodes Suppose nothing is real: the null holds for every outcome you measured. One pre-specified test rejects with probability 0.05. Now suppose you measured four roughly independent outcomes and report whichever produced the smallest p-value. The chance that at least one falls under 0.05 is about 1 - 0.95^4 = 0.185. You have quietly quadrupled your exposure while still printing a single p-value that claims a 5% error rate. Real p-hacking is worse than this arithmetic, because the choices are not made once and independently; they are made adaptively, in the direction that helps. And it is rarely visible in the write-up — the discarded branches leave no trace, so a reader cannot count how many chances were taken. ## Selection is not just about significance One consequence worth stating plainly: when a result survives because it crossed a threshold, the *magnitude* it reports is not an unbiased estimate either. You selected the noisiest, most favourable version of the analysis, so the estimate you publish sits in the lucky tail. That is why p-hacked findings tend to shrink or vanish when someone repeats the study exactly. ## What is not the fix - **Being honest.** P-hacking is mostly unwitting. Sincerity does not restore an error rate. - **Reporting the exact p-value.** Printing "p = 0.037" instead of a star communicates precision about a number that was chosen from a menu. - **A stricter alpha alone.** Moving to 0.005 lowers every branch's threshold but does nothing about the fact that the branch was chosen after the fact. ## What is the fix Decide the analysis before you see the outcome data and write it down: primary outcome, exclusion rules, stopping rule, model and covariates. Then honour the split between *confirmatory* claims (pre-specified, p-values interpretable) and *exploratory* ones (everything else, reported as hypothesis-generating and labelled as such). Exploration is not the sin — presenting exploration in confirmatory clothing is. When a result matters, the strongest move is to re-run the chosen analysis, unchanged, on data that had no part in choosing it. ## In an interview Expect to be asked to define the term, name three or four concrete degrees of freedom, and say precisely what is broken. The crisp formulation: the p-value is a property of the procedure, and p-hacking reports a p-value for a procedure that was never actually followed.

  • If the flexibility was completely innocent, is the reported p-value still wrong?
    Yes. The p-value describes how often a fixed procedure produces a false alarm, and the procedure that was actually followed included choosing among several analyses. Intent does not enter the calculation, which is exactly why p-hacking is dangerous: the people doing it usually believe every step they took was justified, and each step, taken alone, was.
  • Is exploratory analysis itself illegitimate then?
    No. Exploration is how hypotheses get generated, and refusing to look at data would be worse. The requirement is labelling: report exploratory findings as exploratory, describe the search that produced them, and do not attach a confirmatory p-value to them. The claim becomes confirmatory only when a pre-specified analysis is run on data that had no role in choosing it.
  • How would you rescue a finding you suspect came from a p-hacked analysis?
    Treat it as a hypothesis, not a result. Write down the exact analysis that produced it — outcome, exclusions, model, subgroup — freeze it, and run it unchanged on fresh data or a held-back sample. If the effect survives at a comparable magnitude, you now have a real finding; if it shrinks or disappears, you have learned what the original p-value could not tell you.

It is like a marksman who fires ten shots at a blank wall, paints a target around the tightest cluster, and shows you a photograph of one bullseye.

saying these in an interview costs you the question

  • Calls p-hacking deliberate fraud; most of it is unwitting
  • Says the result is fine because each analysis step was defensible
  • Thinks reporting one p-value per paper makes p-hacking impossible
  • Claims a stricter alpha alone repairs flexible analysis
  • Treats after-the-fact outlier removal as a neutral cleaning step
  • Assumes the reported effect size is unbiased even if the p-value is not

context

open as a page

What is HARKing, and why does it invalidate the p-value it is reported with?

level: middleimportance: should knowfreq 45%

basics

~20 s

HARKing means Hypothesizing After the Results are Known: you find which comparison came out significant, then present it as the hypothesis you set out to test. The write-up hides the search, so the p-value is uninterpretable.

open as a page

What is the garden of forking paths, and how does it differ from deliberate p-hacking?

level: seniorimportance: should knowfreq 38%

basics

~20 s

The garden of forking paths (Gelman and Loken) is the point that one analysis can be biased even when only one test is run: choices made after seeing the data mean the analyses you would otherwise have run still count.

open as a page

How would you preregister analyses on a team that must still explore data freely?

level: principalimportance: should knowfreq 32%

basics

~20 s

Do not ban exploration, separate it. Require a time-stamped plan naming the primary outcome, exclusions and model before outcome data is inspected; treat everything else as labelled exploratory work, and confirm on data that had no role in choosing it.

open as a page

How does the file-drawer problem bias a published literature, and can you detect it?

level: seniorimportance: nice to knowfreq 26%

basics

~20 s

Null results stay unpublished in the file drawer, so the visible literature is selected on significance: too many positives, and effects too large. A funnel plot of effect size against precision exposes it: missing small null studies appear as asymmetry.

open as a page