What is HARKing, and why does it invalidate the p-value it is reported with?
answer
- hypothesis written after seeing results
- the write-up sin, not the arithmetic sin
- eight subgroups, one promoted to a prediction
- erases the record of the search
- reader cannot discount for chances taken
basics
~20 sHARKing means Hypothesizing After the Results are Known: you find which comparison came out significant, then present it as the hypothesis you set out to test. The write-up hides the search, so the p-value is uninterpretable.
solid answer
~50 sHARKing, a term from Norbert Kerr, is presenting a post-hoc hypothesis as if it had been stated in advance. The classic shape: you measured eight subgroups, one of them showed an effect, and the paper's introduction now motivates exactly that subgroup with theory written afterwards. Nothing in the arithmetic is falsified — the test on that subgroup was computed correctly. What is falsified is the *narrative*, and the narrative is what tells a reader how many chances the finding had. A reader who believes a single hypothesis was named up front reads p = 0.03 as a 3% false-alarm risk; a reader who knows eight subgroups were scanned reads the same number as unremarkable. HARKing is the reporting sin that makes p-hacking invisible: it erases the record of the search, so no reader and no reviewer can discount for it.
go deeper
Be able to expand the acronym and give the one-line story: the hypothesis was written after seeing which comparison worked, then presented as if it came first.
Explain why the arithmetic being right does not save the result. The p-value's meaning depends on the procedure, and HARKing misrepresents the procedure to the reader while leaving the number untouched.
Demonstrate that you interrogate reports for it: ask what was written down before the data arrived, which comparisons went unreported, and whether a surprising finding is labelled as exploratory or smuggled into the introduction.
Own the incentive problem. Explain how review norms and reporting templates in your organisation either reward tidy predicted stories or make the full comparison set mandatory, and what it costs to switch.
## The definition HARKing stands for **Hypothesizing After the Results are Known**, a term introduced by Norbert Kerr in 1998. It describes a specific move in scientific writing: after seeing which analysis produced a significant result, the author writes the paper's hypothesis section so that this result appears to have been predicted from the start. It is not lying about the numbers. The data are real, the test was computed correctly, the p-value is arithmetically right for the comparison shown. What has been rewritten is the *order of events* — and the order of events is exactly the information a reader needs to interpret the number. ## The canonical shape A study collects an outcome and eight ways to slice the sample: age bands, gender, region, tenure, device, channel, first-time versus returning, weekday versus weekend. The overall effect is null. One slice — returning users on mobile, say — shows p = 0.02. The HARKed write-up says: *"We hypothesised that the intervention would be most effective for returning mobile users, because habit strength moderates response to interface changes."* Then it reports p = 0.02. The theory is plausible, the citation is real, and the sentence would never have been written if that slice had come out null. ## Why the p-value is no longer interpretable A p-value is a statement about a procedure: if the null is true, a procedure that names one comparison in advance and tests it at alpha = 0.05 produces a false alarm 5% of the time. A procedure that scans eight comparisons and reports the best is a different procedure with a much larger false-alarm rate. HARKing does not change which procedure was followed — the search happened either way. It changes which procedure the *reader believes* was followed. And because interpretation of a p-value is entirely relative to the procedure, the reader's inference is wrong even though the arithmetic is right. The number in the paper is correct for a test that was never pre-specified. A useful way to say this in an interview: p-hacking is what you do to the analysis; HARKing is what you do to the write-up. They usually travel together, and HARKing is the half that makes the other half undetectable. ## The near-neighbours HARKing gets confused with - **Reporting an unexpected finding.** Entirely legitimate — and it is legitimate precisely because it is labelled. "We did not predict this; the effect appeared only among returning mobile users and should be treated as exploratory" is honest science. The same sentence with the prediction moved to the introduction is HARKing. - **Generating hypotheses from data.** Also legitimate, and the normal engine of research. The problem starts when the hypothesis born from a dataset is then tested — and reported as confirmed — on that same dataset. - **Tidying the narrative for readability.** Papers are written after the work, so some retrospective ordering is unavoidable. The line is whether the reordering hides how many comparisons were in play. Cutting a paragraph is editing; promoting the winning subgroup to a prediction is HARKing. ## Why it is so common The incentives are brutally aligned against disclosure. A paper with a clean predicted effect reads as competent; a paper reporting "we looked at eight subgroups and one was significant, which may be noise" reads as a failure, even though it is more informative. Kerr's original point was partly sociological: HARKing survives because the literature rewards tidy stories, not because anyone decided to deceive. ## How to detect and prevent it Detection, from the outside, is mostly about asking what is missing: - Does the hypothesis section name a subgroup or moderator suspiciously specific to the significant result? - Are all the measured outcomes and slices reported, including null ones? Silence about a slice that was obviously collected is the tell. - Does the theory feel reverse-engineered — plausible for the finding, but equally plausible for its opposite? Prevention is structural. A time-stamped analysis plan written before the outcome data is inspected fixes the order of events in a way that cannot be quietly rewritten. Reporting standards that require every collected measure and every planned comparison to be listed remove the option of silent omission. And separating a report into a confirmatory section (pre-specified, p-values meaningful) and an exploratory section (labelled, treated as hypothesis-generating) lets a genuine surprise be published without dressing it as a prediction. ## The correct move when you find a surprise Say it plainly: this was not predicted, here is the full set of comparisons examined, here is the one that stood out, and here is the pre-specified test we will run on new data to find out whether it is real. That report is more useful than the HARKed version, and it is the version that survives replication.
- How is HARKing different from honestly reporting an unexpected finding?Only by disclosure, and that is the whole difference. Reporting a surprise as a surprise tells the reader how many comparisons were examined, so they can discount the result appropriately and treat it as hypothesis-generating. HARKing moves the same finding into the introduction as a prediction, which strips out exactly the information needed to interpret its p-value.
- You inherit an analysis whose stated hypothesis names one oddly specific segment. What do you ask?Ask three things: when was that hypothesis written down relative to seeing the outcome data, what other segments were computed, and where are their results. If the plan predates the data and the null segments are reported, the specificity is fine. If the answer is that the segment 'emerged from the analysis', treat the finding as exploratory and propose a confirmatory run on fresh data.
- Does HARKing bias the reported effect size as well as the p-value?Yes. The promoted comparison was chosen because it looked strong, so its estimate sits in the favourable tail of the sampling variation rather than at the true value. Even if the effect is real, the magnitude reported after selection is typically too large, which is one reason HARKed findings shrink when someone repeats the study exactly.
saying these in an interview costs you the question
- Says HARKing is fine because the calculation itself is correct
- Confuses HARKing with any post-hoc or exploratory analysis
- Thinks a plausible theory retroactively justifies the promoted subgroup
- Claims disclosure is only a style preference, not a validity issue
- Assumes reviewers can detect HARKing from the paper alone