An A/B test shows a +40% conversion lift in one locale and flat elsewhere — what do you do?
answer
- too good to be true, usually is
- base rates favour defects over discoveries
- check units before checking mechanisms
- one locale, one deploy, one integration
- Twyman's law
basics
~20 sTreat it as a suspected defect before treating it as a finding. Twyman's law says any figure that looks unusually interesting is usually wrong, so check instrumentation, currency and units, bot traffic and the assignment split in that locale first.
solid answer
~50 sA +40% lift in one locale when everything else is flat is far outside what the change plausibly does, and Twyman's law — any figure that looks interesting or different is usually wrong — says the first hypothesis should be a defect, not a discovery. I would check, in order: whether the conversion event is logged the same way in that locale, whether revenue or price values are being read in the wrong currency or scale, whether bot or internal traffic is concentrated in that market, whether the assignment split in that locale matches the intended ratio, and whether some other release shipped only there during the window. I would also check whether the locale's sample is small enough that a 40% swing is inside its ordinary noise. Only if every one of those comes back clean does the result graduate to a hypothesis worth a confirmatory test.
go deeper
Be ready to say that an unusually large result should be checked for data problems before it is believed, and to name a couple of concrete checks such as event logging and currency units.
Explain why the base rates favour a defect: real product effects are small, while instrumentation errors produce huge apparent effects, so an extreme number is evidence about the pipeline.
Give an ordered investigation with the highest-yield checks first, tie in the segment's interval width, and state exactly what evidence would move the result from suspected defect to testable hypothesis.
Own the norm. Decide what validation is mandatory before any segment result enters a readout, and build the culture where surprising good news is interrogated as hard as surprising bad news.
## Why the reflex should be suspicion, not excitement Twyman's law is the old data-analysis maxim that **any figure that looks interesting or different is usually wrong**. It is not cynicism; it is base rates. Genuine treatment effects of the size product changes actually produce are typically small — a few percent on a conversion metric is a good outcome. Data defects, by contrast, produce enormous apparent effects routinely: a mis-scaled currency multiplies revenue by a hundred, a duplicated event doubles a count, a filter that silently drops one arm's users inflates the other's rate. So when a single locale shows +40% and the rest of the world shows nothing, the likelihood ratio strongly favours a defect over a real effect. This is also a segmentation problem. The locale cut is one of many cells you could have looked at, and the cell that stands out was selected for standing out — so its estimate is inflated by the selection even before you ask whether the data are right. ## The checklist, roughly in order of yield **1. Metric definition and logging.** Is the conversion event emitted by the same code path in that locale? Localised flows often use a different template, a different payment provider or a legacy page. A conversion that fires twice, or fires on page view rather than purchase, produces exactly this pattern. **2. Units and currency.** For revenue-shaped metrics, a locale is the classic place for a unit error: minor units versus major units, a missing conversion to a reporting currency, a decimal separator parsed wrong. Check the raw distribution, not the mean — a unit bug usually shows as a cluster of values in the wrong order of magnitude. **3. Population and filters.** Bots, load tests, internal QA traffic and affiliate traffic concentrate in specific markets. So do users whose consent settings suppress some events. If treatment and control are affected unequally by any of these, the comparison is broken. **4. Assignment.** Compare the observed split of users between arms in that locale against the intended ratio. A ratio that deviates more than sampling noise allows means the randomisation or the exposure logging is compromised there, and no effect estimate from that cell can be trusted. **5. Concurrent changes.** Did anything else ship to that market during the window — a price change, a promotion, a partner outage, a regulatory banner? A locale is precisely the scope at which such changes land, and they will masquerade as a treatment effect. **6. Sample size and noise.** Compute the interval, not just the point estimate. A locale holding 2% of traffic can throw a 40% swing with a interval spanning zero. If the interval is enormous, there is nothing to explain: the number is noise, and no defect hunt is needed. **7. Consistency.** Does the effect appear in secondary metrics that should move with it? A real +40% in conversions should show in downstream revenue, session counts and funnel steps. A defect usually breaks one metric and leaves its neighbours untouched. ## If it survives all of that Then you have an interesting hypothesis, not a result. State the mechanism you believe: something about that market — payment methods, language quality, shipping cost, a competitor's absence — makes the treatment much more valuable there. Then test it as a confirmatory experiment, powered for that locale, with the segment named in advance. If the effect is real it will replicate, though almost certainly smaller than +40%, because the discovering estimate was selected on being extreme. ## What good sounds like in an interview Candidates who impress here do three things: they name Twyman's law or the underlying base-rate argument explicitly, they give an *ordered* checklist rather than a scattered list of possibilities, and they say plainly what evidence would move them from "defect" to "finding". Candidates who do badly start designing a rollout for that locale. The instinct being tested is whether you can be handed a flattering number and stay sceptical of it. ## The failure mode this prevents The expensive version of this mistake is not the wasted analysis; it is shipping a targeted change to the locale, watching the metric fail to move in production, and then spending a quarter explaining the gap between the experiment's promise and reality. Every hour spent validating an implausible number is cheap compared with defending a launch built on a logging bug.
- The locale's interval on that +40% spans from -5% to +90%. Does that change your response?Yes, substantially. An interval that wide and containing zero means the cell is too small to say anything, so the point estimate needs no explanation at all. I would report it as uninformative rather than launch a defect investigation, though I would still sanity-check units if the metric is revenue-shaped, since a magnitude error can survive inside a wide interval.
- Everything checks out and the number is real. What do you tell the product team?That we have a credible hypothesis about that market and nothing more. I would name the mechanism we think explains it, propose a confirmatory experiment powered for that locale with the segment declared up front, and warn explicitly that the replicated effect will very likely be smaller than +40% because the original estimate was selected for being extreme.
- Does Twyman's law apply to a surprisingly bad segment result too?Yes, and teams apply it far less consistently there. A catastrophic drop in one segment deserves the same instrumentation and assignment checks as a spectacular gain, because the same defect classes produce both. The asymmetry is cultural: nobody wants to interrogate good news, and everybody wants to explain away bad news.
A thermometer reading 60 degrees in one room of the house tells you more about the thermometer than about the room.
saying these in an interview costs you the question
- Accepts the lift and proposes a locale rollout
- Invents a market explanation before checking the data
- Never checks currency, units or event definitions
- Ignores the segment's interval width entirely
- Applies scepticism only to disappointing results