skip to content

How do you tell a real subgroup effect in an A/B test from post-hoc story fitting?

level: seniorimportance: should knowfreq 45%

answer

  1. a story explains any cell
  2. was the mechanism written first?
  3. test the difference, not two halves
  4. neighbouring cells should agree
  5. confirmatory run, powered for the segment

basics

~20 s

Demand evidence the story cannot supply: a mechanism stated before the data, a formal test that the effect differs between segments rather than two separate within-segment tests, consistency across neighbouring cells, and replication in a confirmatory experiment.

solid answer

~50 s

A post-hoc story explains whichever cell happened to win, so I look for evidence a story cannot manufacture. First, a mechanism argued before the cut was chosen — if the reasoning was written afterwards, it carries no weight, since a different narrative would have been written for a different winner. Second, the right test: a claim that the effect *differs* by segment is a test of the interaction between treatment and segment, not two separate tests. Significant in one arm and non-significant in the other is not evidence of a difference, because the difference between significant and non-significant is not itself significant. Third, coherence: a real effect usually shows a gradient across ordered bands and consistent signs in neighbouring cells, not one isolated cell in a sea of noise. Fourth, replication in a fresh experiment powered for that segment. Effects that survive are actionable; the rest cost only the follow-up.

go deeper

for a junior

Be ready to say that an explanation invented after seeing which segment won is not evidence, and that a segment finding should be confirmed by a second experiment before anyone acts on it.

for a middle

Explain the mechanics: a claim that the effect differs between groups is a test on the difference of effects, and separate within-group verdicts cannot establish it.

for a senior

Show the full discipline: interaction test with an interval, coherence across neighbouring cells, a powered confirmatory run, and an expectation that a real effect shrinks on replication.

for a principal

Own what evidence unlocks a targeted launch in your organisation, and balance the cost of a confirmatory run against the cost of shipping personalised experiences built on noise.

## The problem with stories Human beings are extremely good at explaining any pattern after seeing it. Told that a checkout change worked best for long-tenure users on tablets in one market, a competent product person will produce a plausible mechanism within a minute — and would have produced an equally plausible one for the opposite cell. That is why a narrative attached to a discovered segment is worth almost nothing as evidence. The question is what evidence *does* discriminate. ## Four discriminating tests ### 1. Was the mechanism stated before the data? A mechanism argued in the design doc — "this only changes the mobile layout, so we expect the effect to concentrate on mobile" — is a prediction, and predictions can fail. A mechanism written after the readout is a description, and descriptions cannot fail. Ask when the sentence was written. Timestamped design docs make this checkable rather than a matter of memory. ### 2. Test the interaction, not the two halves The most common technical error in subgroup claims is comparing significance across segments: "it was significant for new devices, p = 0.02, and not significant for old ones, p = 0.11, so the effect is device-dependent". That reasoning is invalid. **The difference between significant and non-significant is not itself statistically significant.** Two estimates can sit either side of a threshold while their difference has an interval comfortably containing zero — especially when the segments have different sample sizes, which alone shifts p-values without any change in the underlying effect. The correct claim is about the *difference of effects*, so the correct test is on that difference — the interaction between the treatment indicator and the segment indicator. Estimate the difference in lifts with its own interval and judge that. Interaction tests are also inherently less powerful than tests of a main effect: detecting an interaction of a given size typically needs substantially more data than detecting a main effect of the same size. That is why so many subgroup claims are underpowered even when they are correctly specified. ### 3. Look for coherence across the grid A real heterogeneity usually leaves a pattern. If the effect grows with tenure, tenure bands should show a gradient rather than one band lit up and its neighbours flat. If the effect is about screen size, phones and tablets should point the same way. A single isolated cell surrounded by cells with the opposite sign is the signature of noise, because noise has no reason to respect adjacency and a real mechanism usually does. Ordered segments give you this check for free; unordered ones (country, channel) give you less, which is another reason to be more sceptical there. ### 4. Replicate before you act The decisive test is a fresh experiment: declare the segment in advance, power the run for that segment specifically rather than for the whole population, and pre-commit to the decision rule. Two things typically happen. Genuine effects replicate but come back smaller, because the discovering estimate was selected for being extreme and the replication is not. Manufactured effects vanish. Both outcomes are informative and both are cheap relative to shipping a targeted change on a phantom. ## Practical guardrails for a readout - Report the number of cuts examined alongside any segment claim, so a reader can weigh it. - Quote intervals on the *difference* between segments, not two separate lifts side by side. - Keep exploratory findings in a clearly labelled section that cannot, by policy, carry a launch decision. - Write the mechanism down at design time even when you are unsure — a wrong prediction that you honour is more valuable than a right story invented later. - Prefer a small list of pre-declared, well-powered cuts to a large grid, since power, not curiosity, is the binding constraint. ## Answering this in an interview The answer that lands names the interaction-test point explicitly, because most candidates do not. Say the sentence — significant here and not there is not evidence of a difference — and then give the replication requirement as the operational rule. Mention that you expect a real effect to shrink on replication rather than hold at its discovered size; that single detail signals you have actually watched subgroup claims survive contact with a second experiment.

  • An analyst says the effect is significant on mobile and not on desktop, so it is mobile-specific. What is wrong with that?
    Comparing significance is not comparing effects. Two estimates can land on opposite sides of a threshold while the interval on their difference easily contains zero, and unequal segment sample sizes shift p-values on their own. The claim is about the difference of effects, so it needs a test of that difference, with its own interval, not two separate verdicts.
  • Why do subgroup differences need more data than main effects to detect?
    A main effect is estimated from the whole sample, while a difference of effects combines the noise of two smaller estimates, so its standard error is larger. For the same practical size, detecting an interaction typically demands substantially more traffic than detecting an overall effect, which is why most incidental subgroup readouts are badly underpowered.
  • The confirmatory run reproduces the effect but at a third of the original size. How do you read that?
    As a success with a corrected magnitude. The first estimate was selected for being extreme, so shrinkage on replication is expected, not evidence against the effect. I would plan on the replication's number for any forecast or targeting decision and discard the original figure entirely, since it was never an unbiased estimate.

saying these in an interview costs you the question

  • Offers a mechanism invented after seeing the winner
  • Compares significance across segments instead of effects
  • Claims a lone cell proves the effect without neighbours agreeing
  • Ships a targeted change without a confirmatory run
  • Expects a replicated effect to match the discovered size

context