skip to content

Bayesian Statistics

You will learn the Bayesian way of updating beliefs with data — priors, posteriors, MAP vs MLE, and credible intervals — and where it beats the frequentist toolkit. Interviewers use it to test whether you truly understand Bayes' theorem beyond plugging into the formula, and it underpins Bayesian A/B testing and probabilistic ML.

on this pageshow

explore

questions

page 2 of 2

How does a region of practical equivalence change a Bayesian ship decision?

level: seniorimportance: should knowfreq 42%

basics

~20 s

A region of practical equivalence is a band of differences you would call 'no real difference', fixed before analysis. If the posterior for the lift sits entirely inside it, the arms are equivalent and you choose on cost.

open as a page

A new feature shows 3 conversions in 40 sessions: why would a prior earn its keep here?

level: seniorimportance: should knowfreq 48%

basics

~20 s

At 40 sessions the raw 7.5% rate is mostly noise: one more conversion would read 10%. A prior built from comparable past features supplies the information the data lacks and pulls the estimate toward plausible values.

open as a page

Why does updating a Beta posterior event-by-event match one batch update of the same data?

level: seniorimportance: should knowfreq 46%

basics

~20 s

Because the posterior becomes the next prior and the update only adds sufficient statistics. For exchangeable data the likelihood factorises and multiplication commutes, so any batching or ordering accumulates the same success and failure counts and lands on identical parameters.

open as a page

A hierarchical model's sampler reports divergent transitions clustered where the group-level scale is near zero. What do you do?

level: seniorimportance: should knowfreq 43%

basics

~10 s

Treat the fit as biased, not noisy. Divergences near a near-zero scale parameter signal the funnel geometry a centred hierarchical parameterisation creates; the fix is a non-centred parameterisation, then a smaller step size.

open as a page

With 20 observations, a stronger prior halves your credible interval width - how do you report that?

level: seniorimportance: should knowfreq 44%

basics

~20 s

Report the prior explicitly and quantify its weight - a Beta prior contributes roughly its two parameters as pseudo-observations - then publish the interval under a ladder of priors. A narrow interval from a strong prior is an assumption, not evidence.

open as a page

Your posterior for a parameter is bimodal — what goes wrong if you report only the posterior mean?

level: seniorimportance: should knowfreq 48%

basics

~10 s

With two separated modes the posterior mean lands in the trough between them, a value the posterior itself calls unlikely. The single number also hides that two competing explanations are in play.

open as a page

When is the posterior mode a poor summary of a skewed posterior distribution?

level: seniorimportance: should knowfreq 45%

basics

~10 s

A skewed posterior has three different point summaries. Under right skew the mode is the smallest of the three, and it optimises an all-or-nothing loss that matches almost no real decision.

open as a page

Your random-walk Metropolis sampler accepts 2% of proposals — what is wrong and how do you fix it?

level: seniorimportance: should knowfreq 55%

basics

~20 s

The proposal steps are far too large, so almost every candidate lands in a low-density region and is rejected and the chain sits still for long stretches. Shrink the proposal scale until acceptance rises to roughly a quarter.

open as a page

In a posterior predictive check, replicated datasets show far fewer zero-count days than the observed data — what does that mean?

level: seniorimportance: should knowfreq 44%

basics

~20 s

The fitted model cannot generate the excess zeros the real data contain, so it is misspecified for that feature. Quantify the gap with the proportion of zeros as a test statistic, then extend the model to produce structural zeros.

open as a page

How does a hierarchical prior partially pool eight noisy per-group effect estimates?

level: seniorimportance: should knowfreq 52%

basics

~20 s

It treats the eight group effects as draws from a common distribution whose mean and spread are learned from the data. Each estimate is then pulled toward the overall mean, the noisiest groups moving furthest.

open as a page

How do you keep Thompson sampling responsive when an arm's true rate changes over time?

level: seniorimportance: should knowfreq 45%

basics

~20 s

A long-running Thompson sampler locks in because its posteriors become extremely tight and the changed arm barely gets served. Make it forget on purpose: discount old observations each round, or keep only a sliding window of recent ones.

open as a page

How does Thompson sampling's cumulative regret compare with a fixed 50/50 split?

level: seniorimportance: should knowfreq 50%

basics

~20 s

A fixed 50/50 split's cumulative regret grows linearly with traffic: 50,000 impressions parked on a 4% arm instead of a 5% arm costs about 500 conversions. Thompson sampling's grows sublinearly, roughly like the logarithm of the horizon.

open as a page

Why does variational inference minimise KL(q||p) rather than KL(p||q)?

level: seniorimportance: should knowfreq 38%

basics

~20 s

KL(q||p) takes expectations under the approximation you control, so it is computable up to a constant; the reverse direction needs expectations under the unknown posterior. The price is mode-seeking behaviour that ignores parts of a multimodal target.

open as a page

How do you set the ship threshold on probability of superiority for a Bayesian experiment program?

level: principalimportance: should knowfreq 37%

basics

~20 s

The threshold should encode the cost of being wrong, not a borrowed convention. Set it tighter for expensive, hard-to-reverse changes and looser for cheap reversible ones, pair it with a magnitude requirement, and fix it before the experiment runs.

open as a page

How would you decide whether a team reports Bayesian or frequentist results for its recurring decisions?

level: principalimportance: should knowfreq 37%

basics

~20 s

Decide by the shape of the decisions, not by taste. Bayesian reporting pays off when data per decision is thin, credible prior information exists, and stakeholders need a probability about the claim itself. Frequentist reporting suits high-volume standardised readouts.

open as a page

When is the convenience of a conjugate prior not worth the constraint it puts on your model?

level: principalimportance: should knowfreq 34%

basics

~20 s

When the family cannot express the belief or the structure the problem has. Closed form buys exact, constant-memory updates worth keeping at high throughput, but bending a bimodal belief or a covariate-driven model into a convenient family is a modelling error.

open as a page

How do you handle a Bayesian efficacy readout whose conclusion flips between a sceptical and an enthusiastic prior?

level: principalimportance: should knowfreq 38%

basics

~20 s

Report the flip as the finding: if the decision changes across priors reasonable people hold, the data are not decisive. Pre-specify the prior set, quantify the tipping point, and decide on the cost of being wrong.

open as a page

How do you decide between variational inference and MCMC for a production Bayesian model?

level: principalimportance: should knowfreq 35%

basics

~20 s

Decide by which error you can afford. Variational inference is fast and scales, but its approximation error is biased and unmeasured; sampling is asymptotically exact but slow. Match the choice to how sensitive the downstream decision is to uncertainty.

open as a page

Why do two analysts with different reasonable priors reach nearly the same conclusion as data grows?

level: middleimportance: nice to knowfreq 31%

basics

~20 s

Each observation multiplies more likelihood into the posterior, while the prior enters only once. With enough data the likelihood swamps it, so both posteriors concentrate on the same value and take the same approximately normal shape — the Bernstein-von Mises result.

open as a page

How does a Gamma prior on a support-ticket rate update after observing daily counts?

level: middleimportance: nice to knowfreq 30%

basics

~20 s

Events add to the shape, exposure to the rate parameter: in shape-and-rate form, Gamma(alpha, beta) with y tickets over n days becomes Gamma(alpha + y, beta + n). The prior reads as alpha pseudo-events over beta pseudo-days.

open as a page

Does thinning an MCMC chain to every 10th draw improve the posterior estimates?

level: middleimportance: nice to knowfreq 33%

basics

~20 s

No. For a fixed number of iterations, keeping every tenth draw throws away information, so estimates are no better and usually slightly noisier than using all draws. Thinning is justified by storage or downstream cost, not accuracy.

open as a page

Why is the Jeffreys prior for a binomial proportion Beta(1/2, 1/2) rather than uniform?

level: middleimportance: nice to knowfreq 26%

basics

~10 s

Jeffreys' rule sets the prior proportional to the square root of the Fisher information. For a binomial proportion that gives the Beta(1/2, 1/2) density. The motivation is invariance under reparameterisation, not neutrality.

open as a page

How does the Laplace approximation build a Gaussian approximation to a posterior?

level: middleimportance: nice to knowfreq 24%

basics

~20 s

It expands the log posterior to second order around its peak. The linear term vanishes there, leaving a quadratic, which is the log of a Gaussian centred at the peak whose covariance is the inverse of the negative second-derivative matrix.

open as a page

Does checking P(B > A) every morning inflate error the way repeated p-value checks do?

level: seniorimportance: nice to knowfreq 33%

basics

~20 s

The posterior itself needs no correction for how often you look: it summarises the data in hand at any moment. But a rule that stops the first time the probability crosses a bar still selects lucky data and overstates the lift.

open as a page

Two coin experiments with different stopping rules give proportional likelihoods — why is the posterior identical?

level: seniorimportance: nice to knowfreq 26%

basics

~20 s

Two likelihood functions that differ only by a factor free of the parameter give the same posterior, because that factor is absorbed by the normalising constant. With the same prior the two experiments' posteriors coincide exactly.

open as a page

Why is a MAP estimate not invariant under reparameterisation while the MLE is?

level: seniorimportance: nice to knowfreq 22%

basics

~20 s

The posterior is a probability density, so changing variables multiplies it by a Jacobian that reshapes the curve and can move its peak. The likelihood is not a density over the parameter, gets no Jacobian, so its maximiser transforms along.

open as a page

What does Hamiltonian Monte Carlo's momentum variable buy over random-walk Metropolis proposals?

level: seniorimportance: nice to knowfreq 33%

basics

~20 s

Momentum lets a proposal travel a long way while following the posterior's shape, so moves are coherent strides rather than a random walk. Distant proposals still get accepted, which is why exploration holds up as the number of parameters grows.

open as a page

Would you standardise on equal-tailed or highest-density credible intervals across your team's readouts?

level: principalimportance: nice to knowfreq 18%

basics

~20 s

Make equal-tailed the default, because it is reproducible, survives rescaling and is always one interval. Require the highest-density version, plus the posterior plot, in named exception cases: strongly skewed posteriors, posteriors piled against a boundary, and multimodal ones.

open as a page

How do you choose which test statistics to compare in a posterior predictive check?

level: principalimportance: nice to knowfreq 20%

basics

~20 s

Choose statistics that matter for the decision the model supports and that the fitting did not already force to match. A statistic the fit targets, such as the sample mean, passes regardless of how wrong the model is.

open as a page

How would you run Thompson sampling on a news feed where articles arrive and expire daily?

level: principalimportance: nice to knowfreq 28%

basics

~20 s

Thompson sampling handles arms coming and going without modification: each request draws over whatever arms are live, and a new arm's wide posterior earns it exploration. The real problems are cold-start priors, delayed clicks, and exploration swamping exploitation.

open as a page

showing 31–60 of 60