skip to content

How do you compute P(B > A) in a Bayesian A/B test from posterior draws?

level: middleimportance: must knowfreq 68%

answer

  1. simulate rather than integrate
  2. one draw per arm makes one scenario
  3. count the pairs where B wins
  4. error shrinks like one over sqrt of draws
  5. the same draws give the difference posterior

basics

~20 s

Sample a large set of values from each arm's posterior, pair the draws index by index, and take the fraction of pairs in which B's draw exceeds A's. That fraction is a Monte Carlo estimate of the probability of superiority.

solid answer

~50 s

For a conversion metric with a Beta prior, each arm's posterior is Beta(prior successes + observed conversions, prior failures + observed non-conversions). Take 20,000 draws from each arm's posterior, line them up as 20,000 pairs, and count the pairs where B's draw is larger; that count over 20,000 estimates `P(rate_B > rate_A)`. The Monte Carlo error is roughly `sqrt(p*(1-p)/draws)`, so 20,000 draws pin the estimate to about three thousandths near 0.5 — cheap enough that you should over-draw rather than argue about it. The same paired draws give far more than direction: subtract them element-wise and you have a full posterior for the difference, from which you read credible intervals and the expected downside of a wrong call. If the arms share parameters in a hierarchical model, draw from the joint posterior instead of pairing independent marginals.

code

python · 23 lines
python
import random

random.seed(0)

# Beta(1, 1) prior + observed data -> Beta posterior for each arm
# control: 620 conversions of 12000; variant: 690 of 12000
a_post = (1 + 620, 1 + 12000 - 620)
b_post = (1 + 690, 1 + 12000 - 690)

draws = 20000
b_wins = 0
shortfall = 0.0

for _ in range(draws):
    a = random.betavariate(*a_post)
    b = random.betavariate(*b_post)
    if b > a:
        b_wins += 1
    else:
        shortfall += a - b  # how much we lose if we ship B and B is worse

print("P(B > A) =", b_wins / draws)
print("expected loss of shipping B =", shortfall / draws)

go deeper

for a junior

Know the shape of the procedure: draw many values from each arm's posterior, pair them up, and count how often the variant's draw is bigger. That fraction is the answer.

for a middle

Explain the conjugate Beta update, why pairing draws approximates integrating over the joint posterior, and quote the Monte Carlo error so you can defend the number of draws you chose.

for a senior

Show that you take the difference posterior off the same draws and decide on it, and that you know when independent pairing breaks — correlated arms in a hierarchical model need joint sampling.

for a principal

Set the house standard: how many draws, which derived quantities every readout carries, how priors are chosen, and how you keep teams from decoding overlapping marginal intervals as 'no difference'.

## Why simulate at all The quantity you want is an integral: the posterior probability mass over the region where one rate exceeds the other. For two independent Beta posteriors there is a closed form, but nobody reaches for it, because the moment your model grows a covariate, a hierarchical layer or a non-conjugate likelihood, the closed form evaporates while the sampling recipe stays identical. Monte Carlo is the recipe that does not change shape as the model does. ## The recipe 1. **Get a posterior per arm.** For a binary metric with a Beta(a0, b0) prior and observed data, the posterior is Beta(a0 + conversions, b0 + non-conversions). Beta(1, 1) is the flat starting point; a sceptical team uses something tighter around the historical baseline. 2. **Draw.** Take N draws from each arm's posterior. N = 20,000 is a common working choice. 3. **Pair them.** Treat draw i from A and draw i from B as one joint scenario — one plausible state of the world. 4. **Count.** The share of the N scenarios in which B's draw exceeds A's is your estimate of `P(rate_B > rate_A)`. The pairing step is the conceptual heart of it. Each pair is a coherent sample from the joint posterior over both rates, so counting pairs is a direct approximation to integrating over that joint distribution. When the two posteriors are independent — which they are for separate arms with separate priors and separate data — pairing arbitrary independent draws is legitimate, because independent draws paired by index *are* draws from the product distribution. ## How many draws The estimator is a sample proportion, so its standard error is `sqrt(p*(1-p)/N)`. At N = 20,000 and p near 0.5 that is about 0.0035; near 0.95 it is about 0.0015. Two consequences worth stating in an interview: - The error shrinks like one over the square root of N, so quadrupling the draws halves it. Going from 20,000 to 80,000 buys you a factor of two, not four. - The Monte Carlo error is *not* the uncertainty in the experiment. Adding draws makes your reading of the posterior sharper; it does not narrow the posterior, which is fixed by the data and the prior. If someone proposes running more draws to get a more decisive result, they have conflated the two. Because draws are computationally cheap and the number is being used to make a launch call, err high. Reporting a probability of superiority to three decimals off 500 draws is unjustifiable precision. ## What else the same draws buy you Having the paired draws in hand is worth much more than the single headline number. From the same array of pairs: - **Posterior for the difference**: subtract element-wise to get N draws of `rate_B - rate_A`, and take quantiles for a credible interval on the lift. - **Posterior for relative lift**: divide element-wise, which is usually the quantity the business actually discusses. - **Magnitude-aware probabilities**: the share of pairs where the difference exceeds a threshold you care about, rather than merely exceeds zero. - **Expected downside of a wrong choice**: average the shortfall across all pairs, counting zero on the pairs where the chosen arm is better. All of these come from one simulation. This is why the sampling framing dominates in practice: one set of draws answers every question the decision needs. ## Mistakes that show up **Comparing marginal intervals instead of the difference.** Two credible intervals that overlap can still leave a difference posterior almost entirely on one side of zero. Always form the difference from paired draws; never eyeball two separate intervals. **Reusing one random draw for both arms.** If the same underlying random value drives both arms, you have induced a dependence that is not in your model, and the difference posterior collapses toward a deterministic gap. **Sorting the draws.** Sorting each arm's draws before pairing destroys the independence structure and produces a wildly optimistic or pessimistic comparison depending on the sort order. **Ignoring a shared structure.** In a hierarchical model where arms borrow strength from a common hyperparameter, the marginal posteriors are correlated. Pairing independently drawn marginals then misstates the difference — you must draw from the joint posterior instead. ## The interview version Say: posterior per arm, N paired draws, count the share where B wins, quote the Monte Carlo error, and note that the same draws give you the whole posterior for the difference, which is what you actually decide on.

  • How many draws are enough, and what does adding more actually buy?
    The estimate is a sample proportion, so its Monte Carlo error is about `sqrt(p*(1-p)/N)` — roughly 0.0035 at 20,000 draws near 0.5. Quadrupling the draws halves that. What it does not do is narrow the posterior itself, which is fixed by the data and prior. More draws sharpen your reading of the answer, not the answer.
  • What changes if the two arms are correlated in the posterior?
    You must draw from the joint posterior rather than pairing independent draws from each marginal. In a hierarchical model where arms share a hyperparameter, the marginals are correlated, and index-pairing independent marginal draws would overstate the spread of the difference. Sampling the joint posterior keeps the dependence intact.
  • Why not just check whether the two arms' credible intervals overlap?
    Overlapping marginal intervals are entirely compatible with a difference posterior that sits almost wholly above zero, so the overlap check is systematically conservative. Form the difference from the paired draws and interval-check that instead — it is the quantity the decision is about.

saying these in an interview costs you the question

  • Compares the two posterior means and stops there
  • Judges the comparison by whether the marginal intervals overlap
  • Uses a few hundred draws and quotes three decimals
  • Thinks more draws make the experiment more conclusive
  • Reuses one random value for both arms in a pair

context