skip to content

Bayesian Statistics

You will learn the Bayesian way of updating beliefs with data — priors, posteriors, MAP vs MLE, and credible intervals — and where it beats the frequentist toolkit. Interviewers use it to test whether you truly understand Bayes' theorem beyond plugging into the formula, and it underpins Bayesian A/B testing and probabilistic ML.

on this pageshow

explore

questions

page 1 of 2

In a Bayesian A/B test, what does 'probability B beats A is 96%' actually mean?

level: juniorimportance: must knowfreq 76%

answer

  1. a claim about the unknown rates
  2. conditioned on data and prior
  3. share of posterior mass, not error rate
  4. direction only, never magnitude

basics

~20 s

It is the posterior probability that variant B's true rate is higher than variant A's, given the data and the prior. It is a claim about the unknown rates, and says nothing about how large the difference is.

solid answer

~40 s

It means that, after combining the prior with the data observed, 96% of the posterior belief about the pair of true rates falls in the region where B's rate exceeds A's. Two things follow. First, it is a direct statement about the parameters, not about the data or about how often an experiment procedure is right — so the leftover 4% is the posterior chance that A is at least as good as B, not a false-positive rate. Second, it is purely directional: it ranks the arms and carries no information about magnitude. A 96% probability of superiority can sit on top of a posterior lift centred on 0.05%, which is near-certain and still not worth acting on. That is why a probability of superiority is almost never the whole decision rule.

go deeper

for a junior

Be ready to state it in one sentence: the posterior probability that B's true rate exceeds A's, given the data and the prior. Then add that it says nothing about how big the gap is.

for a middle

Explain where the number comes from mechanically, and why the prior and the amount of data both move it. Show that it compresses the whole posterior for the difference into a single direction.

for a senior

Demonstrate that you never report it alone. Pair it with the posterior for the difference and a magnitude-aware statement, and be explicit that the complement is not a launch error rate.

for a principal

Own the reporting standard across teams: which quantities every experiment readout must carry, how priors are chosen and disclosed, and how you stop a single high-sounding percentage from driving launches.

## What the number is In a Bayesian analysis of an experiment you do not end with a single estimate of each arm's conversion rate; you end with a posterior distribution over each rate. The posterior expresses, after seeing the data, how plausible each possible value of the true rate is. Write the two unknown rates as `rate_A` and `rate_B`. The probability of superiority is then simply ``` P(rate_B > rate_A | data, prior) ``` the share of posterior belief that lands in the region where B's true rate is larger. When a tool reports "96% probability B beats A", that number is what it means. ## Why the conditioning matters Every part of the conditioning bar carries weight. - **Given the data**: the number is a summary of the evidence in hand right now. Collect more users and it moves. - **Given the prior**: the same data with a sceptical prior centred on "no difference" yields a lower probability of superiority than with a flat prior, especially early in a test when the data is thin. Two teams can honestly report different numbers from identical data if their priors differ, and a candidate who cannot say this has not internalised what a posterior is. - **About the parameters**: the randomness being described lives in the unknown rates, not in a hypothetical stream of repeated experiments. This is the structural difference from a p-value, which is a probability computed about data under an assumed no-difference world, not a probability attached to the hypothesis itself. ## The two standard misreadings **"So there is a 4% chance we are wrong."** Nearly, but be careful about what "wrong" means. The 4% is the posterior probability that A's rate is at least as high as B's. It is not the long-run error rate of your shipping process, and it is not a guarantee that four out of every hundred launches decided this way will regress. The operating characteristics of a shipping *procedure* — how often it ships a loser across many experiments — depend on the whole rule, including your threshold, your stopping behaviour and how often you test genuinely null ideas. **"96% is high, so the lift is big."** This is the more expensive mistake. Probability of superiority answers only "which arm is on top?" It compresses the entire posterior for the difference into a single direction bit weighted by belief. Precision drives it as much as effect size: with enough traffic, a true lift of 0.02% will eventually push the probability of superiority above 99%, because the posterior for the difference becomes narrow enough to sit almost entirely on the positive side of zero even though it is hugging zero. A tiny, certain win and a large, uncertain win can both report 96%. ## What to report alongside it Because of that second point, a probability of superiority should never travel alone. The same posterior gives you: - **The posterior for the difference itself**, `rate_B - rate_A`, with a credible interval — a range that holds a stated share of posterior belief, so a 95% credible interval is the range within which the difference lies with 95% posterior probability. - **A magnitude-aware probability**, such as the posterior probability that the lift exceeds some amount you actually care about rather than merely exceeds zero. - **The expected downside of choosing wrongly**, expressed in units of the metric. A candidate who reports "96% to beat control, posterior lift 1.2% with a 95% credible interval from 0.3% to 2.1%" has said something a decision-maker can act on. A candidate who reports only "96%" has said which arm won a coin-ranking and nothing about whether the win is worth having. ## Answering it out loud The compact interview answer is three beats: it is a posterior probability about the true rates, given data and prior; the complement is the posterior chance the control is at least as good, not an error rate; and it is directional only, so pair it with the size of the difference before deciding anything.

  • How does that differ from a p-value of 0.04 in the same comparison?
    A p-value is a probability about data: how likely a result at least this extreme would be if the two arms were truly identical. The probability of superiority is a probability about the parameters themselves, given the data you actually collected and your prior. One is a statement under an assumed no-difference world; the other attaches belief directly to the hypothesis that B is better.
  • Can the probability of superiority be 96% while the practical difference is negligible?
    Yes, and it happens constantly at high traffic. Probability of superiority rises with precision as well as with effect size, so a true lift of a few hundredths of a percent will eventually clear 96% once the posterior for the difference is narrow enough to sit mostly above zero. Always read it next to the posterior for the difference.
  • Two teams report different probabilities of superiority from the same counts. How?
    Different priors. A sceptical prior centred on no difference pulls early posteriors toward zero and reports a lower probability of superiority than a flat prior does; with enough data the two converge. It is worth stating the prior alongside the number, because the number is only defined relative to it.

It is like asking which of two runners is genuinely faster, given the races you have watched. Answering 'almost certainly the second one' is not the same as saying she wins by a wide margin.

saying these in an interview costs you the question

  • Calls it the probability that the experiment was run correctly
  • Reads the remaining 4% as a false-positive rate
  • Assumes a high probability implies a large lift
  • Describes it as the probability that the data would repeat
  • Reports it without ever mentioning the prior

context

open as a page

What is the difference between probability as long-run frequency and probability as degree of belief?

level: juniorimportance: must knowfreq 78%

basics

~20 s

The frequency reading defines probability as the proportion of times an outcome occurs in repeatable trials. The belief reading defines it as a numeric degree of confidence, so it can also score one-off events such as a single rocket launch.

open as a page

Starting from a Beta(1,1) prior, what posterior follows from 8 clicks in 100 impressions?

level: juniorimportance: must knowfreq 58%

basics

~10 s

Beta(1 + 8, 1 + 92), that is Beta(9, 93). The prior contributes one pseudo-success and one pseudo-failure, so the posterior mean is 9/102, about 0.088, slightly above the raw rate of 0.08.

open as a page

What does a 95% Bayesian credible interval of [1.9%, 2.5%] say about a conversion rate?

level: juniorimportance: must knowfreq 82%

basics

~20 s

It says 95% of the posterior probability for the conversion rate falls between 1.9% and 2.5%. Given the model and prior you used, there is a 95% probability the rate lies in that range. The probability attaches to the parameter.

open as a page

What is the difference between probability and likelihood after seeing 7 heads in 10 coin flips?

level: juniorimportance: must knowfreq 75%

basics

~20 s

Probability fixes the coin's bias and varies the data; likelihood fixes the observed data and varies the bias. After 7 heads in 10 flips the likelihood is a curve over p that is not a probability distribution.

open as a page

How does a MAP estimate differ from a maximum likelihood estimate of the same parameter?

level: juniorimportance: must knowfreq 68%

basics

~20 s

Maximum likelihood picks the parameter value that makes the observed data most probable. MAP picks the value that maximises the posterior, which is the likelihood multiplied by a prior. When the prior is flat, the two estimates coincide.

open as a page

What does MCMC give you when a posterior has no closed-form normalising constant?

level: juniorimportance: must knowfreq 72%

basics

~20 s

MCMC produces a stream of parameter draws that visit each region in proportion to posterior probability. You never compute the normalising integral: every posterior summary, such as a mean or a tail probability, becomes an average over those draws.

open as a page

What is a posterior predictive distribution in Bayesian inference?

level: juniorimportance: must knowfreq 72%

basics

~20 s

The posterior predictive distribution is the distribution of a future observation after averaging the likelihood over every parameter value the posterior still finds plausible. It carries sampling noise plus parameter uncertainty, instead of conditioning on a single estimate.

open as a page

What distinguishes an informative prior from a weakly informative or flat prior?

level: juniorimportance: must knowfreq 72%

basics

~20 s

An informative prior encodes real outside knowledge and moves the posterior by itself. A weakly informative prior only rules out implausible values, letting the data dominate. A flat prior spreads density evenly and is not genuinely neutral.

open as a page

How do you compute P(B > A) in a Bayesian A/B test from posterior draws?

level: middleimportance: must knowfreq 68%

basics

~20 s

Sample a large set of values from each arm's posterior, pair the draws index by index, and take the fraction of pairs in which B's draw exceeds A's. That fraction is a Monte Carlo estimate of the probability of superiority.

open as a page

In Bayesian versus frequentist inference, is the unknown parameter treated as random or as fixed?

level: middleimportance: must knowfreq 64%

basics

~20 s

Frequentists treat the unknown parameter as a fixed constant and put all the randomness in the data. Bayesians give the parameter a probability distribution that describes how uncertain they are about its value, and condition on the data actually observed.

open as a page

What makes a Beta prior conjugate to a binomial likelihood in Bayesian updating?

level: middleimportance: must knowfreq 70%

basics

~20 s

A Beta density and a binomial likelihood are both powers of theta and 1 minus theta, so their product is again a Beta: a Beta(a, b) prior with s successes in n trials gives Beta(a + s, b + n - s).

open as a page

Why does an MCMC run with 10,000 stored draws report an effective sample size of only 200?

level: middleimportance: must knowfreq 64%

basics

~20 s

Because MCMC draws are dependent. When each draw is nearly a repeat of the previous one, 10,000 correlated draws carry about as much information as 200 independent ones, and effective sample size reports that count rather than the stored count.

open as a page

What does the Gelman-Rubin R-hat statistic compare across MCMC chains?

level: middleimportance: must knowfreq 76%

basics

~20 s

R-hat compares the variation between several MCMC chains with the variation within each chain. If the chains have converged to the same distribution the two agree and R-hat sits near 1; values above roughly 1.01 mean the chains still disagree.

open as a page

How does an equal-tailed credible interval differ from a highest posterior density interval?

level: middleimportance: must knowfreq 61%

basics

~20 s

An equal-tailed interval cuts equal posterior probability off each tail: at 95% it runs from the 2.5th to the 97.5th percentile. A highest posterior density interval is the shortest range holding 95%, so on a skewed posterior it is shorter.

open as a page

In Bayes' rule for a parameter, what does 'posterior is proportional to prior times likelihood' mean?

level: middleimportance: must knowfreq 70%

basics

~20 s

It means the posterior density at each parameter value is prior times likelihood divided by a single constant, the evidence. That constant does not depend on the parameter, so the product alone determines the posterior's shape.

open as a page

In Metropolis-Hastings, why does the unknown normalising constant cancel out of the acceptance ratio?

level: middleimportance: must knowfreq 66%

basics

~20 s

The acceptance rule uses the target only as a ratio of its values at the proposed and current points. Both carry the same constant factor, so it cancels and the sampler needs nothing beyond likelihood times prior.

open as a page

Why is a posterior predictive interval for one new observation wider than the credible interval for the parameter?

level: middleimportance: must knowfreq 58%

basics

~20 s

A predictive interval must cover a single random observation, so it carries the model's sampling noise on top of the uncertainty about the parameter. A parameter interval carries only the second piece, and shrinks toward zero width as data accumulate.

open as a page

In Thompson sampling, how is the arm to serve chosen on a single request?

level: middleimportance: must knowfreq 72%

basics

~20 s

Thompson sampling draws one random value from each arm's posterior over its reward rate, then serves the arm whose draw came out largest. Every request repeats the draw, so allocation follows the posteriors on its own.

open as a page

What does maximising the ELBO achieve in variational inference?

level: middleimportance: must knowfreq 62%

basics

~20 s

Maximising the evidence lower bound (ELBO) picks the distribution in a chosen tractable family that is closest to the true posterior. The bound's gap to the log evidence is exactly a KL divergence, so raising the ELBO shrinks that gap.

open as a page

Why are the warm-up (burn-in) draws of an MCMC chain discarded before summarising the posterior?

level: juniorimportance: should knowfreq 58%

basics

~20 s

Early MCMC draws still reflect where the chain was started rather than the posterior, and in adaptive samplers the tuning is still changing. Those draws are dropped so the summary uses only draws from the stationary distribution.

open as a page

In Normal-Normal conjugate updating, how does the posterior mean combine prior and data?

level: middleimportance: should knowfreq 44%

basics

~20 s

As a precision-weighted average, where precision is one over variance. Precisions add, so the posterior is always more precise than the prior, and its mean lies between the prior mean and the sample mean, closer to whichever is more precise.

open as a page

When can you ignore the marginal likelihood in Bayes' rule, and when must you actually compute it?

level: middleimportance: should knowfreq 45%

basics

~20 s

Inside a single model the marginal likelihood is a constant that only rescales the posterior, so you can ignore it. You must compute it to compare models, since a Bayes factor is a ratio of two marginal likelihoods.

open as a page

Why does a zero-mean Gaussian prior on regression coefficients turn MAP into ridge regression?

level: middleimportance: should knowfreq 55%

basics

~20 s

Taking logs turns the posterior into log-likelihood plus log-prior. A zero-mean Gaussian prior contributes minus the sum of squared coefficients divided by twice the prior variance, so maximising it is least squares with an L2 penalty attached.

open as a page

Why does a Gibbs sampler need no accept-reject step when it draws from full conditionals?

level: middleimportance: should knowfreq 48%

basics

~20 s

A Gibbs sampler draws each block from its exact full conditional, its distribution given the current values of all other parameters and the data. Because the proposal is already exact, the Metropolis acceptance probability equals one.

open as a page

What does a prior predictive check reveal before any data have been observed?

level: middleimportance: should knowfreq 33%

basics

~20 s

A prior predictive check simulates fake datasets from the prior and the likelihood, then asks whether those datasets are physically sensible. It catches priors that imply absurd observables, such as four-metre human heights or negative revenue, before any fitting happens.

open as a page

Why is a flat prior on a probability not flat on the log-odds scale?

level: middleimportance: should knowfreq 46%

basics

~20 s

A density picks up a Jacobian factor under a change of variables. A uniform prior on p in [0,1] becomes a bell-shaped density on log-odds peaked at zero, favouring p near 0.5. No prior is uninformative on every scale.

open as a page

Why does Thompson sampling keep serving an arm with only 1 success in 2 trials?

level: middleimportance: should knowfreq 58%

basics

~20 s

Two trials leave a very wide posterior, so a draw from that arm lands above the leader's draw often enough to win some requests. Uncertainty, not a tuned exploration parameter, is what buys the arm its traffic.

open as a page

Why does a mean-field variational approximation understate posterior variance?

level: middleimportance: should knowfreq 52%

basics

~20 s

A mean-field family factorises the approximation into independent pieces, so it cannot represent correlation between parameters. On a correlated posterior it fits inside the narrow direction of the ridge rather than along it, producing spreads that are too small.

open as a page

In a Bayesian A/B test, what is the expected loss of shipping the variant?

level: seniorimportance: should knowfreq 50%

basics

~20 s

Expected loss of shipping B is the posterior average of how much worse B is than A, counting zero where B is better. It is in metric units, so you compare it to a tolerance.

open as a page

showing 1–30 of 60