How would you preregister analyses on a team that must still explore data freely?
answer
- separate exploration, do not ban it
- confirmatory versus exploratory tracks
- plan must name outcome, exclusions, decision rule
- held-back sample as the enforcement
- log deviations rather than forbid them
basics
~20 sDo not ban exploration, separate it. Require a time-stamped plan naming the primary outcome, exclusions and model before outcome data is inspected; treat everything else as labelled exploratory work, and confirm on data that had no role in choosing it.
solid answer
~50 sThe failure mode is a policy that treats every look at data as a sin; teams route around it and the discipline dies. I would split reporting into two tracks. **Confirmatory**: a short written plan filed before outcome data is inspected — primary outcome, exclusion rules, model and covariates, the specific comparison, and what result would change the decision — with p-values and decisions attached only to that. **Exploratory**: everything else, unrestricted, but reported without confirmatory language and never as the basis for a launch or a publication claim. The enforcement mechanism people accept is a held-back sample: explore all you like on one part, then run the frozen plan once on the reserve. I would also require deviations to be recorded and explained rather than forbidden, since a rule with no legitimate escape hatch just gets broken silently. The cost is real — slower cycles, more nulls on record — and I would name it rather than pretend the discipline is free.
go deeper
Know what a preregistration is: a plan written and time-stamped before the data are examined, saying which outcome will be tested and how, so the analysis cannot be adjusted afterwards.
Be able to list what a binding plan contains, especially the primary outcome definition, exclusion rules and the model, and explain why an explore-then-confirm split on separate data is stronger than a promise to be careful.
Show you have made this work in practice: held-back samples, deviation logs, labelling exploratory findings honestly in a report, and pushing back when a post-hoc segment result is presented as a launch justification.
Own the tradeoff and the incentives. Decide where the threshold for full rigour sits, name the cost in cycle time and headline findings, and explain how you reward pre-specified nulls so the discipline survives delivery pressure.
## The organisational problem Every analytics or research team faces the same tension. Flexible, curious analysis is how you find things; flexible analysis is also what makes claims unreliable. A leader who resolves the tension by banning exploration will get compliance theatre — plans written after the fact, or exploration relabelled as something else. A leader who resolves it by ignoring the problem will get a stream of confident findings that evaporate on contact with the next dataset. The workable answer is not less exploration. It is a **hard boundary between two kinds of claim**, with different rules and different rhetoric attached to each. ## Two tracks **Confirmatory.** Governed by a written, time-stamped plan filed before the outcome data are inspected. This is the only track allowed to produce decisions, launches or headline claims, and the only track where a p-value carries its stated meaning. **Exploratory.** Unlimited. Slice, model, transform, hunt. Its output is *candidate hypotheses*, reported in explicitly hypothesis-generating language, never as an established effect, and never as a launch justification on its own. The rhetorical discipline matters as much as the statistical one. If exploratory findings can be written up in the same voice as confirmatory ones, the boundary has no teeth. ## What a plan must contain to actually bind A plan that says "we will analyse the effect of the intervention" constrains nothing. To close the forks it needs: 1. **The primary outcome**, defined precisely enough that two analysts would compute the same number — including the window, the unit of analysis and the aggregation. 2. **Exclusion and data-quality rules**, written before anyone knows what removing a group does to the result. 3. **The model and covariates**, including transformations and how continuous variables are binned. 4. **The specific comparison** that constitutes the test, and the direction expected. 5. **The decision rule**: what result leads to what action. This is the item most often omitted and the one that most reliably prevents post-hoc reinterpretation of an ambiguous result. 6. **Secondary and exploratory questions**, listed as such, so that later interest in them is documented rather than promoted. Keep it short. A one-page plan that is actually filed beats a template so heavy that nobody fills it in. ## Making it enforceable - **Held-back data.** The most persuasive mechanism, because it does not rely on anyone's memory of what they intended. Explore on one portion; freeze the specification; run it once on a reserve that had no influence on the choice. Analysts accept it because it costs them nothing during exploration. - **Timestamps in a place the team does not control.** A commit, a ticket, a registry entry. The point is not distrust; it is that a plan that can be silently edited provides no evidence. - **Deviation logs, not deviation bans.** Real analyses hit surprises — a broken instrument, an unusable measure. Forbidding deviation guarantees silent deviation. Require that a change be recorded with the reason and the date, and that the report show both the planned and the executed analysis. - **Review at plan time, not only at result time.** Reviewing a design before data collection is where the effort actually improves the study; reviewing after the result mostly negotiates the story. - **Publish the nulls internally.** If a pre-specified analysis that finds nothing disappears from the record, the team has recreated the file drawer inside the company, and the internal knowledge base becomes as selected as any biased literature. ## The costs, named honestly A principal-level answer does not pretend this is free. - **Speed.** Writing a plan and holding back data adds days and reduces the effective sample for the confirmatory run. - **Fewer headline findings.** The visible output gets less impressive, and if leadership rewards striking results the discipline will be quietly abandoned. - **Genuine loss of some real effects.** Restricting confirmatory claims to pre-specified comparisons means some true findings sit in the exploratory bucket for a cycle before they can be confirmed. - **Overhead falls unevenly.** A large, expensive, decision-critical study justifies the full apparatus. A quick diagnostic look does not, and mandating identical rigour for both is how the policy loses credibility. ## Where I would set the threshold Scale the requirement to the stakes. Anything that will drive an irreversible decision, an external claim or a publication gets a full plan and a held-back confirmation. Routine internal analysis gets the lightweight version: name your primary outcome before you look, and label the rest exploratory. That asymmetry is what keeps the heavy process credible for the cases that need it. ## In an interview The signals being looked for are: you know that pre-specification is the fix rather than more sophisticated arithmetic; you know a plan only binds if it names outcomes, exclusions, model and decision rule; you protect exploration instead of criminalising it; and you can state what the discipline costs and who pushes back.
- What is the single item teams most often leave out of an analysis plan?The decision rule. Plans routinely specify the outcome and the model but not what result triggers what action, which leaves an ambiguous finding open to post-hoc reinterpretation. Writing down in advance that a given result means ship, kill, or run again is what stops the discussion after the numbers arrive from becoming a negotiation about what the study was for.
- A genuine problem forces you to deviate from the filed plan. How do you handle it?Record the deviation, the reason and the date, then report both analyses: what the plan specified and what was actually run. A policy that forbids deviation produces silent deviation, which is strictly worse. If the change was driven by anything you saw in the outcome data, say so plainly and downgrade the result to exploratory rather than defending its confirmatory status.
- How do you keep this discipline from being abandoned under delivery pressure?Scale it to the stakes so it is not uniformly expensive: full plans and held-back confirmation for irreversible or externally-visible decisions, a one-line primary-outcome declaration for routine work. Then change what gets rewarded. If only striking positive findings earn credit, the process loses; if a well-executed pre-specified null is treated as a real contribution, it holds.
saying these in an interview costs you the question
- Proposes banning exploratory analysis outright
- Treats a vague plan naming no outcome as preregistration
- Omits any decision rule from the analysis plan
- Forbids deviations instead of requiring them to be logged
- Claims the discipline has no cost in speed or findings
- Applies identical heavy process to every trivial analysis