skip to content

Why does slicing a flat A/B test result into many segments so often produce a false winner?

level: middleimportance: must knowfreq 72%

answer

  1. one experiment, but how many tests?
  2. 5% of nothing is still something
  3. 1 - 0.95^k grows fast
  4. small cells, wide intervals
  5. only extreme cells clear the bar

basics

~20 s

Every segment cut is another hypothesis test. At a 5% false-positive rate, about twenty cuts of a genuinely flat experiment throw up roughly one significant segment by chance alone, so the winner you find is usually noise.

solid answer

~50 s

Reading segments is running many tests, not one. If the treatment truly does nothing and each cut is judged at alpha = 0.05, each cut has a 5% chance of a spurious `significant` result; across twenty roughly independent cuts the chance of at least one is about `1 - 0.95^20`, near 64%, with about one false win expected. Cut device by country by tenure into a thirty-cell grid and a false winner is close to guaranteed. Two things make it worse. Each cell has a fraction of the sample, so its interval is wide and its estimate noisy; and because only cells with large observed effects clear the threshold, the surviving lift is systematically overstated. So an unplanned segment win is a hypothesis, not a result: report how many cuts you ran, widen the bar for multiplicity, and confirm it in a fresh, adequately powered test before shipping.

go deeper

for a junior

Be ready to state that looking at many segments means running many tests, and that some will look significant purely by chance even when the treatment does nothing.

for a middle

Explain the arithmetic out loud: alpha times the number of cuts as expected false wins, one minus 0.95 to the k as the chance of at least one, and why smaller cells give wider intervals.

for a senior

Show the operating discipline. Report the cut count, refuse to ship on a discovered segment without a confirmatory run, and explain that the discovered lift is inflated so any forecast built on it will miss.

for a principal

Own the tradeoff between discovery and discipline. Too strict a rule kills real heterogeneity findings; too loose a one fills the roadmap with phantom wins, so decide where the team's segment readouts sit and what evidence unlocks a targeted launch.

## The setup An A/B test randomises users between a control and a treatment and compares one primary metric. Suppose the all-up read comes back flat: the estimated lift is near zero and its confidence interval comfortably contains zero. A natural next move is to ask *for whom* it might have worked, and to cut the same data by device, country, tenure, acquisition channel, plan tier and so on. That move is where most invented findings on an experimentation team are born. ## Each cut is another test A significance test at level alpha = 0.05 is calibrated so that, **when there is truly no effect**, it wrongly declares a difference 5% of the time. A p-value is the probability of seeing a difference at least as large as the one observed **if the null hypothesis of no true difference were true**. That 5% is a per-test guarantee. It says nothing about a family of tests. Run `k` cuts on a truly flat experiment and, if the cuts were independent, the probability that at least one comes back significant is `1 - (1 - 0.05)^k`: - k = 5 -> about 23% - k = 10 -> about 40% - k = 20 -> about 64% - k = 30 -> about 79% The expected number of false wins is simply `k x 0.05`: one in twenty cuts, one and a half in thirty. Real segment cuts are not independent — Android overlaps with mobile, Brazil overlaps with Latin America, tenure bands are nested — so the exact family-wise error rate is somewhere below these numbers when the cuts are positively correlated. The arithmetic is an approximation, but the direction is not in doubt: the more cells you look at, the closer to certain a spurious winner becomes. A device x country x tenure grid of thirty cells is a machine for manufacturing one. ## Small cells make it worse in two ways **Wider intervals.** The standard error of a mean scales like `s / sqrt(n)`. A cell holding 1/30 of the sample has an interval roughly `sqrt(30)`, about 5.5 times, wider than the all-up read. Segment estimates are therefore dominated by noise even when the underlying effect is real. **Inflated winners.** Because the cell is noisy, only a *large* observed effect clears the significance threshold. Conditioning on having cleared it selects the draws that happened to land high. So the segments that look like wins carry estimates biased away from zero — often by a factor, not a few percent. This is the winner's-curse or exaggeration problem: even when a real effect exists in that segment, the number you would quote from the discovering test is an overstatement, and a replication will come back smaller. Sign errors are possible too: a cell can be significant in the wrong direction. ## What honest practice looks like 1. **Decide the cuts before the data.** A segment named in the design doc — say, a mobile-versus-desktop split written down before launch because the treatment changes a mobile layout — is a confirmatory test with a real error rate, provided you also planned the sample and the reporting rule. Everything decided after seeing results is exploratory. 2. **Count and disclose the cuts.** The readout should say how many segment views were examined, not just the one that paid off. A p-value of 0.03 out of thirty cuts is unremarkable; the same p-value on a single pre-planned cut means something. 3. **Adjust the bar for multiplicity.** When several segment readouts are treated as decisions, the significance threshold has to account for the family of tests rather than each test alone; the specific adjustment procedures are a topic of their own. 4. **Look at the whole grid, not the winner.** If the treatment really helps heavy users, you expect a gradient across tenure bands and a consistent sign across neighbouring cells, not one lone cell surrounded by noise. 5. **Require replication before shipping.** Treat the segment result as a hypothesis and run a fresh experiment powered for that segment specifically. Effects that survive a confirmatory test are worth acting on; the ones that evaporate cost you only the follow-up. ## The sentence to say in the interview "The flat all-up read is the result. The segment win is a hypothesis whose false-positive rate I cannot state without knowing how many cuts I ran, and whose effect size is inflated by the selection that surfaced it." That framing shows you understand both the multiplicity arithmetic and the estimation bias, which is what separates a candidate who has read about p-hacking from one who has actually had to say no to a product manager holding a slide.

  • You ran thirty segment cuts on a flat experiment and one came back at p = 0.03. What is your honest read?
    With thirty cuts on a truly null effect you expect roughly one and a half results below 0.05 by chance, so a single p = 0.03 is exactly what noise looks like. I would report it as an exploratory observation with the cut count attached, not as a finding, and propose a confirmatory test powered for that segment if the mechanism is plausible enough to be worth the traffic.
  • Why is the lift estimate in a segment that reached significance usually overstated?
    Small cells have wide sampling error, so only draws that land far from zero clear the threshold. Selecting on that condition selects the high draws, which biases the surviving estimate away from zero. The published number is therefore the top of a noisy distribution rather than its centre, and a replication will typically come back materially smaller.
  • Does the same problem apply when the all-up result is a clear win?
    Yes. Cutting a winning experiment to find where it worked best is the same multiplicity exercise, and the best-looking cell is still selected on noise. The difference is that the launch decision rests on the all-up read, which is sound, so the risk is a wrong targeting or rollout story rather than a wrong ship decision.

Fire thirty arrows at a blank wall, then paint a target around the one nearest the centre. The bullseye is real; the marksmanship is not.

saying these in an interview costs you the question

  • Says the segment is significant so the effect is real
  • Never mentions how many cuts were examined
  • Treats a small cell's estimate as precise
  • Reports only the winning slice and drops the rest
  • Assumes a 5% error rate survives twenty tests unchanged

context