What is the garden of forking paths, and how does it differ from deliberate p-hacking?
answer
- one test can still be biased
- choices contingent on the data
- branches not taken still count
- Gelman and Loken's metaphor
- intent is irrelevant to the error rate
basics
~20 sThe garden of forking paths (Gelman and Loken) is the point that one analysis can be biased even when only one test is run: choices made after seeing the data mean the analyses you would otherwise have run still count.
solid answer
~50 sGelman and Loken's garden of forking paths describes a researcher who runs exactly one analysis and reports exactly one p-value, yet whose result is still untrustworthy. The reason is counterfactual: at each junction — which outcome to use, how to bin a continuous variable, which control to include, which comparison is the interesting one — the choice was made after seeing the data and would have been different under a different dataset. The error rate of a procedure is defined over hypothetical repetitions, so the branches never taken still belong to the procedure. This is why it differs from p-hacking as usually imagined: no one tried twenty analyses and hid nineteen. The analyst was sincere, ran one test, and can honestly say so — and the p-value is still too small. It is also why asking 'how many tests did you run?' is the wrong audit question; the right one is 'would you have made the same choices under different data?'
go deeper
Know the headline: an analysis can be biased even when only one test was run, because the choices along the way were made after looking at the data.
Explain the counterfactual mechanism in your own words. A p-value assumes the analysis is a fixed function of the data, and a specification chosen adaptively breaks that assumption regardless of how many tests were computed.
Show the audit instinct in a real review: ask what was fixed before outcomes were seen, probe what would have been reported under a different result, and downgrade adaptively-specified findings to exploratory rather than arguing about intent.
Own the uncomfortable conclusion that there is no adjustment factor here, only process. Be ready to defend spending on pre-specified analysis plans, held-back data and replication against the objection that the team already only runs one test.
## The problem it names The standard picture of p-hacking is an analyst grinding through twenty variants and reporting the one that hit. Andrew Gelman and Eric Loken pointed out that the same statistical damage occurs in a far more common and far more innocent case: the analyst runs **one** analysis, reports **one** p-value, and would pass a lie detector when saying so. Their metaphor is a garden of forking paths. Every analysis walks a path through a large number of junctions. Which of the collected measures is the outcome. Whether age enters as a continuous variable or in bands, and where the bands fall. Whether to log-transform. Which observations are usable. Which of several natural comparisons is 'the' comparison. Whether to control for the obvious confounder. The analyst walked one path — but the path was chosen *while looking at the data*, and under different data they would have walked a different one. ## Why counterfactual branches count A p-value is defined by a thought experiment over repeated datasets: if the null were true, and I applied *this procedure* to fresh data, how often would I see a result this extreme? The definition therefore requires the procedure to be a fixed function of the data. If the function itself changes depending on what the data look like, the reference distribution used to compute the p-value is the wrong one — it corresponds to a rigid procedure that was never in force. The branches not taken are not hypothetical noise; they are part of what the procedure would do across the repetitions the p-value is defined over. So a single test, chosen adaptively, carries a real error rate above its nominal one, even though nothing was hidden and nothing was discarded. ## How it differs from p-hacking | | classic p-hacking | garden of forking paths | | --- | --- | --- | | tests actually run | many | one | | anything concealed | yes, the failed variants | no | | analyst's account of events | inaccurate | truthful | | error rate | inflated | inflated | The last row is the point: the consequence is identical. Which is why 'I only ran one test' is not the defence people think it is, and why the honest-analyst objection ('I would never fish') misses the mechanism entirely. The forking-paths framing removes intent from the discussion and makes it about whether the analysis specification was fixed before the data were seen. ## Recognising it in practice Warning signs in your own work or someone else's: - A specification that is unusually well suited to the data — bins that land exactly where the effect lives, an outcome definition that appears in no prior work. - A comparison described as "the natural one" that you cannot find written down anywhere before the analysis. - An analyst who cannot answer "what would you have done if the effect had appeared in the other direction, or in a different measure?" — if the honest answer is "I'd have reported that instead", the branch was live. - Many defensible modelling choices in a domain with weak theory and noisy, small data. That combination is where forking paths do the most damage, because noise decides which path looks compelling. ## What to do about it Since there is no list of tests to count, there is nothing to mechanically adjust, and that is the uncomfortable part of the diagnosis. The workable responses are procedural: 1. **Fix the specification before seeing outcomes.** Write down the outcome definition, transformations, bins, exclusions, model and the target comparison in advance. A specification you cannot bend after the fact has no forks left. 2. **Split the data.** Explore freely on one part; run the frozen specification once on a held-back part. Only the second run gets a confirmatory reading. 3. **Report the whole garden.** Show that the result survives across the reasonable choices you might have made — several binnings, several outcome definitions, with and without the covariate. A finding that appears only on one path was probably grown by that path. 4. **Treat single studies as weak evidence.** Independent replication with a pre-specified analysis is what actually settles it. ## In an interview The crisp statement is: the multiplicity that matters is not the number of tests you ran, but the number of analyses that were live given how you made your choices. Then give the concrete tell — an analyst who ran one test but chose the specification after seeing the data has the same inflated error rate as one who ran twenty.
- If the analyst genuinely ran one test, what exactly is inflated?The error rate of the procedure, not the arithmetic of that one calculation. A p-value is defined over repeated datasets under the null, and across those repetitions this analyst's specification would change with the data. The reference distribution used assumes a rigid specification, so the reported probability is smaller than the procedure's true false-alarm rate.
- How would you audit a colleague's analysis for forking paths?Do not ask how many tests they ran. Ask what was decided before the outcome data was inspected, and then ask the counterfactual: if the effect had shown up in a different measure, a different bin, or the opposite direction, what would the report have said? If a different result would have produced a different specification, the branches were live and the p-value should be read as exploratory.
- Does a robustness table across many specifications solve the problem?It helps but does not fully solve it. Showing the effect survives across binnings, outcome definitions and covariate sets rules out the case where noise on one arbitrary path created the result. It does not restore the nominal error rate, and the choice of which specifications to display is itself a fork. The honest use is as evidence of stability, with confirmation left to fresh data.
A hiker who takes a single route through a maze can still be said to have searched it, if at every junction they turned toward whichever corridor happened to look brightest.
saying these in an interview costs you the question
- Says one test cannot be biased because there is nothing to correct
- Treats sincerity of the analyst as a statistical defence
- Confuses it with running many tests and hiding the failures
- Thinks a numeric correction factor exists for unrun branches
- Ignores that weak theory plus noisy data multiplies the forks