skip to content

Posterior Decision Rules

Once you hold a posterior you still have to act: compute the chance a variant wins, weigh the loss of being wrong, or route traffic adaptively. Interviewers ask what rule triggers the call.

on this pageshow

explore

questions

11

In a Bayesian A/B test, what does 'probability B beats A is 96%' actually mean?

level: juniorimportance: must knowfreq 76%

answer

  1. a claim about the unknown rates
  2. conditioned on data and prior
  3. share of posterior mass, not error rate
  4. direction only, never magnitude

basics

~20 s

It is the posterior probability that variant B's true rate is higher than variant A's, given the data and the prior. It is a claim about the unknown rates, and says nothing about how large the difference is.

solid answer

~40 s

It means that, after combining the prior with the data observed, 96% of the posterior belief about the pair of true rates falls in the region where B's rate exceeds A's. Two things follow. First, it is a direct statement about the parameters, not about the data or about how often an experiment procedure is right — so the leftover 4% is the posterior chance that A is at least as good as B, not a false-positive rate. Second, it is purely directional: it ranks the arms and carries no information about magnitude. A 96% probability of superiority can sit on top of a posterior lift centred on 0.05%, which is near-certain and still not worth acting on. That is why a probability of superiority is almost never the whole decision rule.

go deeper

for a junior

Be ready to state it in one sentence: the posterior probability that B's true rate exceeds A's, given the data and the prior. Then add that it says nothing about how big the gap is.

for a middle

Explain where the number comes from mechanically, and why the prior and the amount of data both move it. Show that it compresses the whole posterior for the difference into a single direction.

for a senior

Demonstrate that you never report it alone. Pair it with the posterior for the difference and a magnitude-aware statement, and be explicit that the complement is not a launch error rate.

for a principal

Own the reporting standard across teams: which quantities every experiment readout must carry, how priors are chosen and disclosed, and how you stop a single high-sounding percentage from driving launches.

## What the number is In a Bayesian analysis of an experiment you do not end with a single estimate of each arm's conversion rate; you end with a posterior distribution over each rate. The posterior expresses, after seeing the data, how plausible each possible value of the true rate is. Write the two unknown rates as `rate_A` and `rate_B`. The probability of superiority is then simply ``` P(rate_B > rate_A | data, prior) ``` the share of posterior belief that lands in the region where B's true rate is larger. When a tool reports "96% probability B beats A", that number is what it means. ## Why the conditioning matters Every part of the conditioning bar carries weight. - **Given the data**: the number is a summary of the evidence in hand right now. Collect more users and it moves. - **Given the prior**: the same data with a sceptical prior centred on "no difference" yields a lower probability of superiority than with a flat prior, especially early in a test when the data is thin. Two teams can honestly report different numbers from identical data if their priors differ, and a candidate who cannot say this has not internalised what a posterior is. - **About the parameters**: the randomness being described lives in the unknown rates, not in a hypothetical stream of repeated experiments. This is the structural difference from a p-value, which is a probability computed about data under an assumed no-difference world, not a probability attached to the hypothesis itself. ## The two standard misreadings **"So there is a 4% chance we are wrong."** Nearly, but be careful about what "wrong" means. The 4% is the posterior probability that A's rate is at least as high as B's. It is not the long-run error rate of your shipping process, and it is not a guarantee that four out of every hundred launches decided this way will regress. The operating characteristics of a shipping *procedure* — how often it ships a loser across many experiments — depend on the whole rule, including your threshold, your stopping behaviour and how often you test genuinely null ideas. **"96% is high, so the lift is big."** This is the more expensive mistake. Probability of superiority answers only "which arm is on top?" It compresses the entire posterior for the difference into a single direction bit weighted by belief. Precision drives it as much as effect size: with enough traffic, a true lift of 0.02% will eventually push the probability of superiority above 99%, because the posterior for the difference becomes narrow enough to sit almost entirely on the positive side of zero even though it is hugging zero. A tiny, certain win and a large, uncertain win can both report 96%. ## What to report alongside it Because of that second point, a probability of superiority should never travel alone. The same posterior gives you: - **The posterior for the difference itself**, `rate_B - rate_A`, with a credible interval — a range that holds a stated share of posterior belief, so a 95% credible interval is the range within which the difference lies with 95% posterior probability. - **A magnitude-aware probability**, such as the posterior probability that the lift exceeds some amount you actually care about rather than merely exceeds zero. - **The expected downside of choosing wrongly**, expressed in units of the metric. A candidate who reports "96% to beat control, posterior lift 1.2% with a 95% credible interval from 0.3% to 2.1%" has said something a decision-maker can act on. A candidate who reports only "96%" has said which arm won a coin-ranking and nothing about whether the win is worth having. ## Answering it out loud The compact interview answer is three beats: it is a posterior probability about the true rates, given data and prior; the complement is the posterior chance the control is at least as good, not an error rate; and it is directional only, so pair it with the size of the difference before deciding anything.

  • How does that differ from a p-value of 0.04 in the same comparison?
    A p-value is a probability about data: how likely a result at least this extreme would be if the two arms were truly identical. The probability of superiority is a probability about the parameters themselves, given the data you actually collected and your prior. One is a statement under an assumed no-difference world; the other attaches belief directly to the hypothesis that B is better.
  • Can the probability of superiority be 96% while the practical difference is negligible?
    Yes, and it happens constantly at high traffic. Probability of superiority rises with precision as well as with effect size, so a true lift of a few hundredths of a percent will eventually clear 96% once the posterior for the difference is narrow enough to sit mostly above zero. Always read it next to the posterior for the difference.
  • Two teams report different probabilities of superiority from the same counts. How?
    Different priors. A sceptical prior centred on no difference pulls early posteriors toward zero and reports a lower probability of superiority than a flat prior does; with enough data the two converge. It is worth stating the prior alongside the number, because the number is only defined relative to it.

It is like asking which of two runners is genuinely faster, given the races you have watched. Answering 'almost certainly the second one' is not the same as saying she wins by a wide margin.

saying these in an interview costs you the question

  • Calls it the probability that the experiment was run correctly
  • Reads the remaining 4% as a false-positive rate
  • Assumes a high probability implies a large lift
  • Describes it as the probability that the data would repeat
  • Reports it without ever mentioning the prior

context

open as a page

How do you compute P(B > A) in a Bayesian A/B test from posterior draws?

level: middleimportance: must knowfreq 68%

basics

~20 s

Sample a large set of values from each arm's posterior, pair the draws index by index, and take the fraction of pairs in which B's draw exceeds A's. That fraction is a Monte Carlo estimate of the probability of superiority.

open as a page

In Thompson sampling, how is the arm to serve chosen on a single request?

level: middleimportance: must knowfreq 72%

basics

~20 s

Thompson sampling draws one random value from each arm's posterior over its reward rate, then serves the arm whose draw came out largest. Every request repeats the draw, so allocation follows the posteriors on its own.

open as a page

Why does Thompson sampling keep serving an arm with only 1 success in 2 trials?

level: middleimportance: should knowfreq 58%

basics

~20 s

Two trials leave a very wide posterior, so a draw from that arm lands above the leader's draw often enough to win some requests. Uncertainty, not a tuned exploration parameter, is what buys the arm its traffic.

open as a page

In a Bayesian A/B test, what is the expected loss of shipping the variant?

level: seniorimportance: should knowfreq 50%

basics

~20 s

Expected loss of shipping B is the posterior average of how much worse B is than A, counting zero where B is better. It is in metric units, so you compare it to a tolerance.

open as a page

How does a region of practical equivalence change a Bayesian ship decision?

level: seniorimportance: should knowfreq 42%

basics

~20 s

A region of practical equivalence is a band of differences you would call 'no real difference', fixed before analysis. If the posterior for the lift sits entirely inside it, the arms are equivalent and you choose on cost.

open as a page

How do you keep Thompson sampling responsive when an arm's true rate changes over time?

level: seniorimportance: should knowfreq 45%

basics

~20 s

A long-running Thompson sampler locks in because its posteriors become extremely tight and the changed arm barely gets served. Make it forget on purpose: discount old observations each round, or keep only a sliding window of recent ones.

open as a page

How does Thompson sampling's cumulative regret compare with a fixed 50/50 split?

level: seniorimportance: should knowfreq 50%

basics

~20 s

A fixed 50/50 split's cumulative regret grows linearly with traffic: 50,000 impressions parked on a 4% arm instead of a 5% arm costs about 500 conversions. Thompson sampling's grows sublinearly, roughly like the logarithm of the horizon.

open as a page

How do you set the ship threshold on probability of superiority for a Bayesian experiment program?

level: principalimportance: should knowfreq 37%

basics

~20 s

The threshold should encode the cost of being wrong, not a borrowed convention. Set it tighter for expensive, hard-to-reverse changes and looser for cheap reversible ones, pair it with a magnitude requirement, and fix it before the experiment runs.

open as a page

Does checking P(B > A) every morning inflate error the way repeated p-value checks do?

level: seniorimportance: nice to knowfreq 33%

basics

~20 s

The posterior itself needs no correction for how often you look: it summarises the data in hand at any moment. But a rule that stops the first time the probability crosses a bar still selects lucky data and overstates the lift.

open as a page

How would you run Thompson sampling on a news feed where articles arrive and expire daily?

level: principalimportance: nice to knowfreq 28%

basics

~20 s

Thompson sampling handles arms coming and going without modification: each request draws over whatever arms are live, and a new arm's wide posterior earns it exploration. The real problems are cold-start priors, delayed clicks, and exploration swamping exploitation.

open as a page