skip to content

In an A/B test, what is the difference between a pre-registered segment cut and a post-hoc one?

level: juniorimportance: should knowfreq 58%

answer

  1. the data is the same either way
  2. when was the cut chosen?
  3. written down before launch
  4. confirmatory versus exploratory
  5. one generates hypotheses, one tests them

basics

~20 s

A pre-registered cut is fixed in the design doc before launch, so its error rate holds. A post-hoc cut is chosen after seeing results, so its p-value was selected and cannot be read at face value.

solid answer

~50 s

The data can be identical; what differs is when the cut was chosen. If the design doc says before launch "we will also read mobile separately, because the change only affects the mobile layout", that cut is confirmatory: one planned test, a stated significance threshold, and ideally enough traffic in the mobile arm to detect the effect size you care about. If instead the all-up read is flat and you then hunt through device, country and tenure until something looks good, the cut is exploratory: the number you report was picked for being extreme, so its p-value no longer means what a p-value means. Post-hoc cuts are still worth doing — that is how real heterogeneity gets discovered — but they generate hypotheses, not decisions. The honest readout labels each cut as planned or exploratory and states how many exploratory cuts were run.

go deeper

for a junior

Be ready to say that a pre-registered cut is decided and written down before the experiment starts, while a post-hoc cut is chosen after seeing the numbers, and that this changes how much the result is worth.

for a middle

Explain why the timing matters mechanically: the significance guarantee applies to a test fixed in advance, and choosing a cut because it looked good already spent the data.

for a senior

Demonstrate you can run a readout that carries both kinds without confusion, including powering the planned segment properly and refusing to let an exploratory cut drive a launch.

for a principal

Own where your organisation draws the line. Decide which segments are pre-registered by default, what a team must show to promote an exploratory cut, and how to keep the rule from stifling genuine discovery.

## Same numbers, different evidential status A segment cut is just a filter on the experiment's users — mobile only, Brazil only, users with more than 90 days of tenure. Nothing about the arithmetic changes based on when you decided to apply it. What changes is the **error rate you are entitled to quote**. A significance test at alpha = 0.05 promises that, when the treatment truly does nothing, it will declare a difference at most 5% of the time. That promise holds for a test you committed to before looking. It does not hold for a test you chose *because* its numbers looked interesting, because the choosing step already used the data. ## Pre-registered (confirmatory) cuts A pre-registered cut is one that appears in the experiment's design document before the first user is bucketed. A well-formed one carries four things: - **The cut itself** — mobile versus desktop, stated exactly, including how the boundary is defined. - **A reason** — a mechanism you believe in advance. "The treatment changes a layout that only mobile users see" is a reason; "mobile is a segment we have" is not. - **A sample plan** — enough traffic inside the segment to detect an effect worth acting on. Splitting an experiment powered for the whole population usually leaves each half underpowered, so this often means running longer or targeting the experiment at the segment. - **A decision rule** — what launch or no-launch action the segment read licenses, written down before it can be argued about. Because the cut was fixed in advance, its result is a genuine test. If you pre-register several, they still form a family and the threshold has to reflect that, but each one is at least an honest test rather than a selected one. ## Post-hoc (exploratory) cuts A post-hoc cut is anything decided after the data are visible: the flat result you sliced afterwards, the dimension you added because a colleague asked, the tenure band you moved from 30 days to 14 days because the split looked better there. These are not illegitimate — most useful heterogeneity is first noticed this way. The error is in reporting them with the vocabulary of a confirmatory test. Three things go wrong when a post-hoc cut is treated as a finding: 1. **The stated error rate is fiction.** Twenty cuts on a null effect give roughly one significant result by chance, so "p = 0.04" from a hunt carries far less evidence than the same number from a planned test. 2. **The effect size is inflated.** The cut surfaced because its estimate was large; selecting on largeness biases the estimate away from zero, so a replication comes back smaller. 3. **The story writes itself.** Any segment can be rationalised after the fact — a mechanism invented to explain a result is not evidence for it, because you would have invented a different mechanism for whichever cell happened to win. ## How to handle both in one readout A readout that stays honest usually looks like this. The primary all-up metric with its interval is the decision. Below it, a short block of pre-registered segment reads, each labelled as planned. Below that, an explicitly labelled exploratory block that states how many cuts were examined and presents the interesting ones as candidates for a follow-up rather than as results. Nothing in the exploratory block moves a launch decision on its own. The practical discipline that makes this work is small: agree the segment list at design review, and let the analysis tool record which cuts were actually pulled. A cut count in the readout costs nothing and changes how a reader weighs everything under it. ## The junior-level version If you remember one line: **planned cuts test a hypothesis, discovered cuts generate one**. Both belong in the readout; only the first belongs in the decision.

  • What makes a pre-registered segment cut well-formed rather than just a name in a document?
    It states the cut precisely, gives a mechanism for why the effect should differ there, plans enough traffic inside the segment to detect an effect worth acting on, and fixes the decision rule in advance. Without the sample plan the cut is honest but underpowered, so a flat segment read tells you nothing.
  • Are post-hoc segment cuts something a team should avoid entirely?
    No. Exploratory slicing is how genuine heterogeneity is first noticed, and forbidding it wastes information you already paid to collect. The rule is about status, not permission: label the cuts as exploratory, disclose how many you ran, and require a confirmatory follow-up before any of them changes a launch or targeting decision.

saying these in an interview costs you the question

  • Says a cut is valid because the data supports it
  • Calls any segment split pre-registration without a written plan
  • Believes exploratory cuts should never be run
  • Reads a post-hoc p-value at face value

context