An A/A run flags 1 of 20 metrics at alpha 0.05 - do you block the launch?
answer
- count how many tests you ran
- twenty times 0.05 is one expected flag
- at least one flag two runs in three
- borderline versus astronomically small p-value
- calibration over many runs beats one verdict
basics
~20 sNo. Scanning 20 metrics at a 0.05 threshold when nothing truly differs gives one flag on average, and at least one flag about 64% of the time. Investigate the flagged metric, but a single borderline result is not evidence of a broken split.
solid answer
~50 sOne flag out of twenty is the expected outcome, not a finding. With 20 metrics tested at a 0.05 threshold under no real difference, the expected number of flags is `20 x 0.05 = 1`, and the chance of at least one is `1 - 0.95^20`, roughly 64%. So the default answer is: do not block. What would change my mind is a pattern rather than a count. Is the p-value borderline at 0.04, or one in a million? Is the gap between two identical experiences large in absolute terms, which points at instrumentation? Does the *same* metric flag again on a repeat run? And is the flagged quantity a downstream behaviour metric, or the number of users assigned to each arm - because a mismatch in the split itself is a blocking defect diagnosed on its own terms. Decide the acceptance rule before you run, not after you see the result.
go deeper
Remember the arithmetic: twenty metrics tested at a 0.05 threshold with nothing truly different give about one flag on average. A single flag in an A/A run is expected rather than alarming.
Explain why the chance of at least one flag among twenty independent metrics is about 64%, and why that makes a lone borderline result uninformative about whether the split works.
Show the triage that separates noise from defect: how extreme the p-value is, how big the gap is in the metric's own units, whether the same metric flags on repeat runs, and whether the flag sits on behaviour or on assignment counts.
Set the acceptance policy before anyone runs the test - which metrics are scanned, what magnitude blocks a launch regardless of significance, and how repeated A/A runs are used to demonstrate that the platform's error rate matches what it claims.
## Why one flag is unremarkable A significance threshold is a deliberate false-positive rate. Setting alpha to 0.05 means that when there is truly no difference, the procedure will still declare one about 5% of the time. Run that procedure over 20 metrics on data where, by construction, no treatment exists, and the arithmetic is immediate: - Expected number of flagged metrics: `20 x 0.05 = 1`. - Probability that at least one flags, if the metrics were independent: `1 - 0.95^20 = 0.64`. So a healthy platform running an A/A across twenty metrics will show a flag on roughly two runs in three. Reporting zero flags every time would be the surprising outcome - it would suggest the analysis is over-conservative rather than that the split is unusually clean. In practice the metrics are correlated with one another, which pulls the at-least-one probability somewhat below 64%, but the qualitative conclusion is unchanged. This is why `one metric out of twenty came back significant` is, on its own, information-free. The number of tests you performed is part of the result and has to be reported with it. ## What actually distinguishes noise from a defect **How extreme, not merely whether.** Chance produces p-values just under the threshold; it does not produce astronomically small ones. A metric at 0.04 among twenty is exactly what the arithmetic predicts. A metric at one in a million is not reachable by sampling variation on a null comparison, so it points to a defect in assignment or logging. **Absolute size, independent of the p-value.** Two identically treated arms differing by 15% on an event count is an instrumentation problem even if the sample is small enough that the p-value looks unimpressive. Always look at the magnitude in the metric's own units, not only at the verdict. **Reproducibility.** Noise does not repeat on the same metric. Re-run the A/A - or look at the last several A/A runs - and see whether the same metric keeps flagging. A metric that flags three runs in a row is not chance; it has a cause. **Where the flag sits.** There is an important asymmetry between downstream behaviour metrics and the split itself. A difference in a conversion rate between identical arms is a candidate for noise. A mismatch between the number of users assigned to each arm and the ratio you configured is not a metric result at all - it says the assignment mechanism is not doing what you asked, and that is a blocking defect handled on its own terms rather than filed under expected variation. ## The right acceptance criterion A single A/A run is a weak instrument, so do not build the gate on one verdict. The stronger criterion is calibration across many runs: over repeated A/A tests, the share that flag any given metric should sit near the alpha you set, and the p-values should be scattered across the whole range rather than piling up near zero. A platform that flags 15% of the time at a 0.05 threshold has a calibration problem no matter how comfortable any individual run looked. This criterion has the advantage of being falsifiable and of not depending on a judgement call each time a run comes back. It also detects the failure mode a single run cannot: an analysis whose variance is systematically understated, which makes every future experiment more likely to declare a win that is not there. ## Decide the rule before you look The organisational failure here is arguing about the criterion after seeing the result. Write down beforehand: which metrics are in the A/A panel, what magnitude of discrepancy blocks a launch regardless of significance, what pattern across repeat runs blocks a launch, and who makes the call. Otherwise the decision is made by whoever is most invested in launching. ## Answering in an interview Open with the arithmetic - one expected flag, about a 64% chance of at least one - so it is clear you are not talking yourself out of a real problem. Then give the escalation criteria: extremity of the p-value, absolute size of the gap, whether it repeats, and whether the flag is on assignment counts rather than behaviour. Finish with the calibration-across-runs criterion. Candidates who either block on any flag or wave away every flag are showing the same weakness from opposite sides.
- What p-value from a flagged A/A metric would actually worry you?One that is not borderline. A metric at 0.04 among twenty is what chance produces; a p-value of one in a million is not reachable by sampling variation on a null comparison and points to a defect in assignment or logging. Magnitude matters independently: a 15% gap between two identical experiences is a bug even when the sample is too small for the p-value to look impressive.
- How do you make the A/A acceptance criterion trustworthy rather than a single verdict?Run many A/A tests, not one. Across repeated runs the share flagging any given metric should sit near your alpha, and the p-values should be scattered across the full range rather than piling up near zero. A platform that flags 15% of the time at a 0.05 threshold has a calibration problem regardless of what any single run said.
- The flagged quantity is the number of users assigned to each arm. Same reasoning?No. That is not a behaviour metric, it is the split itself, and a mismatch there says the assignment mechanism is not producing the ratio you configured. Treat an imbalance in arm membership as a blocking defect diagnosed on its own terms rather than filing it under expected noise.
saying these in an interview costs you the question
- Expects a healthy A/A to flag zero metrics out of twenty
- Blocks every launch on any A/A flag regardless of size
- Dismisses an astronomically small p-value as A/A noise
- Reports the flagged metric as a real difference between arms
- Ignores how many metrics were scanned when reading the result