In a Bayesian A/B test, what is the expected loss of shipping the variant?
answer
- probability ignores magnitude
- average only the downside
- zero on the draws where you were right
- divide by all draws, not the losing ones
- compare it to a tolerance in metric units
basics
~20 sExpected loss of shipping B is the posterior average of how much worse B is than A, counting zero where B is better. It is in metric units, so you compare it to a tolerance.
solid answer
~50 sTake the paired posterior draws and, for each pair, compute how far short of A the variant falls — `max(rate_A - rate_B, 0)` — then average that over *all* draws, not only the losing ones. That average is the expected loss of shipping B: the conversion you would give up, in expectation, by choosing B when B is in fact worse. The stopping rule becomes "ship when expected loss drops below a tolerance you fixed beforehand", for example two hundredths of a percentage point of checkout conversion. It beats a bare probability threshold because it is magnitude-aware in the metric's own units. A 30% posterior chance of being worse by 0.001 percentage points is harmless; a 3% chance of being worse by five points is not, and a rule keyed on probability alone cannot tell those apart. You compute the same quantity for shipping A, and both can be small at once.
go deeper
Know that a Bayesian readout can report more than 'B probably wins': it can report how much you would expect to lose by shipping B if B turns out to be worse, in units of the metric.
Be able to compute it from paired posterior draws and state the formula: average the shortfall of the chosen arm across all draws, counting zero wherever the choice was right.
Show why it makes a better stopping rule than a bare probability threshold, using the contrast between a likely-but-tiny downside and an unlikely-but-severe one, and read both directions of the loss.
Own the tolerance: what expected cost the business will accept, why it differs for reversible and irreversible changes, and when an asymmetric loss weighting is the honest model of the risk.
## The quantity Expected loss answers the question a probability of superiority dodges: *if I am wrong, how much does it cost me?* For a decision between control A and variant B, the expected loss of choosing B is ``` E[ max(rate_A - rate_B, 0) ] ``` where the expectation is taken over the posterior for the pair of rates. In words: over the whole posterior, average the amount by which the control beats the variant, treating that amount as zero on every part of the posterior where the variant is genuinely ahead. It is sometimes called risk, or expected regret. With paired posterior draws it is a one-liner: sum `rate_A - rate_B` across the draws where A is larger, and divide by the *total* number of draws. Dividing by the count of losing draws instead is the single most common implementation error, and it produces a much larger number that answers a different question — the conditional loss *given* that you were wrong, rather than the unconditional expected cost of the decision. ## Why it makes a better stopping rule A probability threshold compresses the posterior into one bit of direction weighted by belief. Expected loss keeps the units. Consider two experiments: - **Experiment 1**: the posterior says there is a 30% chance the variant is worse, but in the worst plausible case it is worse by 0.001 percentage points. Expected loss is negligible. Ship it; the downside is rounding error. - **Experiment 2**: the posterior says there is a 3% chance the variant is worse, but if it is worse it is worse by five percentage points. Expected loss is around 0.15 percentage points — potentially a serious hit. A 97% probability of superiority looks like an easy green light and is not. A rule keyed on "probability above 95%" ships the second and blocks the first. A rule keyed on expected loss gets both right, because it prices the tail rather than merely counting it. Expected loss also degrades gracefully as evidence accumulates. Early in a test, both arms carry large expected loss because the posterior is wide. As data accumulates, the expected loss of the better arm falls monotonically toward zero in the typical case, which gives you a natural, interpretable stopping quantity: stop when the cost of a mistake has become small enough to accept. ## Setting the tolerance The threshold has to be in the same units as the metric, which is exactly the property that makes it discussable outside the analytics team. On a checkout funnel with a baseline near 5% conversion, a tolerance of 0.02 percentage points says: we accept a decision rule whose expected cost, if it is wrong, is two hundredths of a point of conversion. Product and finance can argue about that number on its merits, which they cannot do with "95%". Because the tolerance encodes willingness to lose, it should be tighter for changes that are expensive or slow to reverse and looser for a copy tweak that can be rolled back in an afternoon. ## Both directions Expected loss is computed for each candidate decision. The expected loss of keeping A is `E[max(rate_B - rate_A, 0)]` — the upside you forgo by not shipping. Two structural cases follow: - **Both losses small.** The posterior for the difference is concentrated near zero. Neither choice can cost you much; decide on other grounds — simplicity, maintenance cost, strategic direction. - **Both losses large.** The posterior is wide and straddles zero. You do not have enough evidence, and either choice carries real exposure. That two-sided reading is the practical payoff: expected loss makes "we genuinely cannot lose much either way" an explicit, defensible conclusion instead of an awkward silence. ## Asymmetric costs The plain formula treats a unit of loss on one side as a unit of gain on the other. When that is false — a regression damages trust or triggers a support burden out of proportion to the revenue involved — weight the shortfall. Multiplying the negative side by a factor greater than one encodes "a point lost hurts more than a point gained helps" and shifts the rule toward the incumbent. Being able to say that a loss function is a modelling choice, not a fixed formula, is what separates a senior answer from a memorised one. ## Common failures Quoting expected loss without units, so nobody can judge whether it is small. Averaging only over the losing draws. Comparing expected loss across metrics measured on different scales. And shipping on a probability threshold while showing the expected loss on the dashboard as decoration.
- What is the expected loss of keeping the control instead?It is the mirror image: the posterior average of `max(rate_B - rate_A, 0)`, the upside you give up by not shipping. You compute both and act on whichever falls under the tolerance. If both are small, the posterior for the difference is concentrated near zero and the choice can be made on cost or simplicity; if both are large, you simply do not have enough evidence yet.
- What if a regression hurts more than an equal-sized gain helps?Then the symmetric formula misprices the decision. Weight the downside — multiply the shortfall by a factor above one — so the rule demands more evidence before moving off the incumbent. The loss function is a modelling choice you should make explicit, not a fixed formula, and stating the weight forces the business to say how asymmetric the risk really is.
- Why is dividing by the number of losing draws wrong?That gives the loss conditional on being wrong, which is a much larger number and answers a different question. Expected loss is the unconditional cost of the decision, so the draws where the choice was right must contribute zero to the average. Mixing the two makes a healthy experiment look dangerous and breaks any tolerance calibrated on the correct definition.
A probability of superiority tells you the odds the umbrella stays in the bag; expected loss tells you how wet you get on average when it does not.
saying these in an interview costs you the question
- Treats expected loss as a probability between 0 and 1
- Averages only over the draws where the chosen arm loses
- Quotes the number with no units attached
- Compares expected loss across metrics on different scales
- Ships on a probability threshold and ignores magnitude entirely