In a Bayesian A/B test, what does 'probability B beats A is 96%' actually mean?
answer
- a claim about the unknown rates
- conditioned on data and prior
- share of posterior mass, not error rate
- direction only, never magnitude
basics
~20 sIt is the posterior probability that variant B's true rate is higher than variant A's, given the data and the prior. It is a claim about the unknown rates, and says nothing about how large the difference is.
solid answer
~40 sIt means that, after combining the prior with the data observed, 96% of the posterior belief about the pair of true rates falls in the region where B's rate exceeds A's. Two things follow. First, it is a direct statement about the parameters, not about the data or about how often an experiment procedure is right — so the leftover 4% is the posterior chance that A is at least as good as B, not a false-positive rate. Second, it is purely directional: it ranks the arms and carries no information about magnitude. A 96% probability of superiority can sit on top of a posterior lift centred on 0.05%, which is near-certain and still not worth acting on. That is why a probability of superiority is almost never the whole decision rule.
go deeper
Be ready to state it in one sentence: the posterior probability that B's true rate exceeds A's, given the data and the prior. Then add that it says nothing about how big the gap is.
Explain where the number comes from mechanically, and why the prior and the amount of data both move it. Show that it compresses the whole posterior for the difference into a single direction.
Demonstrate that you never report it alone. Pair it with the posterior for the difference and a magnitude-aware statement, and be explicit that the complement is not a launch error rate.
Own the reporting standard across teams: which quantities every experiment readout must carry, how priors are chosen and disclosed, and how you stop a single high-sounding percentage from driving launches.
## What the number is In a Bayesian analysis of an experiment you do not end with a single estimate of each arm's conversion rate; you end with a posterior distribution over each rate. The posterior expresses, after seeing the data, how plausible each possible value of the true rate is. Write the two unknown rates as `rate_A` and `rate_B`. The probability of superiority is then simply ``` P(rate_B > rate_A | data, prior) ``` the share of posterior belief that lands in the region where B's true rate is larger. When a tool reports "96% probability B beats A", that number is what it means. ## Why the conditioning matters Every part of the conditioning bar carries weight. - **Given the data**: the number is a summary of the evidence in hand right now. Collect more users and it moves. - **Given the prior**: the same data with a sceptical prior centred on "no difference" yields a lower probability of superiority than with a flat prior, especially early in a test when the data is thin. Two teams can honestly report different numbers from identical data if their priors differ, and a candidate who cannot say this has not internalised what a posterior is. - **About the parameters**: the randomness being described lives in the unknown rates, not in a hypothetical stream of repeated experiments. This is the structural difference from a p-value, which is a probability computed about data under an assumed no-difference world, not a probability attached to the hypothesis itself. ## The two standard misreadings **"So there is a 4% chance we are wrong."** Nearly, but be careful about what "wrong" means. The 4% is the posterior probability that A's rate is at least as high as B's. It is not the long-run error rate of your shipping process, and it is not a guarantee that four out of every hundred launches decided this way will regress. The operating characteristics of a shipping *procedure* — how often it ships a loser across many experiments — depend on the whole rule, including your threshold, your stopping behaviour and how often you test genuinely null ideas. **"96% is high, so the lift is big."** This is the more expensive mistake. Probability of superiority answers only "which arm is on top?" It compresses the entire posterior for the difference into a single direction bit weighted by belief. Precision drives it as much as effect size: with enough traffic, a true lift of 0.02% will eventually push the probability of superiority above 99%, because the posterior for the difference becomes narrow enough to sit almost entirely on the positive side of zero even though it is hugging zero. A tiny, certain win and a large, uncertain win can both report 96%. ## What to report alongside it Because of that second point, a probability of superiority should never travel alone. The same posterior gives you: - **The posterior for the difference itself**, `rate_B - rate_A`, with a credible interval — a range that holds a stated share of posterior belief, so a 95% credible interval is the range within which the difference lies with 95% posterior probability. - **A magnitude-aware probability**, such as the posterior probability that the lift exceeds some amount you actually care about rather than merely exceeds zero. - **The expected downside of choosing wrongly**, expressed in units of the metric. A candidate who reports "96% to beat control, posterior lift 1.2% with a 95% credible interval from 0.3% to 2.1%" has said something a decision-maker can act on. A candidate who reports only "96%" has said which arm won a coin-ranking and nothing about whether the win is worth having. ## Answering it out loud The compact interview answer is three beats: it is a posterior probability about the true rates, given data and prior; the complement is the posterior chance the control is at least as good, not an error rate; and it is directional only, so pair it with the size of the difference before deciding anything.
- How does that differ from a p-value of 0.04 in the same comparison?A p-value is a probability about data: how likely a result at least this extreme would be if the two arms were truly identical. The probability of superiority is a probability about the parameters themselves, given the data you actually collected and your prior. One is a statement under an assumed no-difference world; the other attaches belief directly to the hypothesis that B is better.
- Can the probability of superiority be 96% while the practical difference is negligible?Yes, and it happens constantly at high traffic. Probability of superiority rises with precision as well as with effect size, so a true lift of a few hundredths of a percent will eventually clear 96% once the posterior for the difference is narrow enough to sit mostly above zero. Always read it next to the posterior for the difference.
- Two teams report different probabilities of superiority from the same counts. How?Different priors. A sceptical prior centred on no difference pulls early posteriors toward zero and reports a lower probability of superiority than a flat prior does; with enough data the two converge. It is worth stating the prior alongside the number, because the number is only defined relative to it.
It is like asking which of two runners is genuinely faster, given the races you have watched. Answering 'almost certainly the second one' is not the same as saying she wins by a wide margin.
saying these in an interview costs you the question
- Calls it the probability that the experiment was run correctly
- Reads the remaining 4% as a false-positive rate
- Assumes a high probability implies a large lift
- Describes it as the probability that the data would repeat
- Reports it without ever mentioning the prior