Bayesian Statistics
You will learn the Bayesian way of updating beliefs with data — priors, posteriors, MAP vs MLE, and credible intervals — and where it beats the frequentist toolkit. Interviewers use it to test whether you truly understand Bayes' theorem beyond plugging into the formula, and it underpins Bayesian A/B testing and probabilistic ML.
on this pageshowhide
explore
- Priors and Likelihoods21 questions
- Prior Choice5 questions
- Likelihood and Posterior5 questions
- Conjugate Updating6 questions
- Belief Versus Frequency5 questions
- Posterior Computation15 questions
- MCMC Samplers5 questions
- Convergence Diagnostics5 questions
- Variational Inference5 questions
- Posterior Summaries13 questions
- MAP Versus MLE4 questions
- Credible Intervals4 questions
- Predictive Distributions5 questions
- Posterior Decision Rules11 questions
- Probability to Beat Baseline6 questions
- Thompson Sampling5 questions
questions
page 2 of 2How does a region of practical equivalence change a Bayesian ship decision?
basics
~20 sA region of practical equivalence is a band of differences you would call 'no real difference', fixed before analysis. If the posterior for the lift sits entirely inside it, the arms are equivalent and you choose on cost.
A new feature shows 3 conversions in 40 sessions: why would a prior earn its keep here?
basics
~20 sAt 40 sessions the raw 7.5% rate is mostly noise: one more conversion would read 10%. A prior built from comparable past features supplies the information the data lacks and pulls the estimate toward plausible values.
Why does updating a Beta posterior event-by-event match one batch update of the same data?
basics
~20 sBecause the posterior becomes the next prior and the update only adds sufficient statistics. For exchangeable data the likelihood factorises and multiplication commutes, so any batching or ordering accumulates the same success and failure counts and lands on identical parameters.
A hierarchical model's sampler reports divergent transitions clustered where the group-level scale is near zero. What do you do?
basics
~10 sTreat the fit as biased, not noisy. Divergences near a near-zero scale parameter signal the funnel geometry a centred hierarchical parameterisation creates; the fix is a non-centred parameterisation, then a smaller step size.
With 20 observations, a stronger prior halves your credible interval width - how do you report that?
basics
~20 sReport the prior explicitly and quantify its weight - a Beta prior contributes roughly its two parameters as pseudo-observations - then publish the interval under a ladder of priors. A narrow interval from a strong prior is an assumption, not evidence.
Your posterior for a parameter is bimodal — what goes wrong if you report only the posterior mean?
basics
~10 sWith two separated modes the posterior mean lands in the trough between them, a value the posterior itself calls unlikely. The single number also hides that two competing explanations are in play.
When is the posterior mode a poor summary of a skewed posterior distribution?
basics
~10 sA skewed posterior has three different point summaries. Under right skew the mode is the smallest of the three, and it optimises an all-or-nothing loss that matches almost no real decision.
Your random-walk Metropolis sampler accepts 2% of proposals — what is wrong and how do you fix it?
basics
~20 sThe proposal steps are far too large, so almost every candidate lands in a low-density region and is rejected and the chain sits still for long stretches. Shrink the proposal scale until acceptance rises to roughly a quarter.
In a posterior predictive check, replicated datasets show far fewer zero-count days than the observed data — what does that mean?
basics
~20 sThe fitted model cannot generate the excess zeros the real data contain, so it is misspecified for that feature. Quantify the gap with the proportion of zeros as a test statistic, then extend the model to produce structural zeros.
How does a hierarchical prior partially pool eight noisy per-group effect estimates?
basics
~20 sIt treats the eight group effects as draws from a common distribution whose mean and spread are learned from the data. Each estimate is then pulled toward the overall mean, the noisiest groups moving furthest.
How do you keep Thompson sampling responsive when an arm's true rate changes over time?
basics
~20 sA long-running Thompson sampler locks in because its posteriors become extremely tight and the changed arm barely gets served. Make it forget on purpose: discount old observations each round, or keep only a sliding window of recent ones.
How does Thompson sampling's cumulative regret compare with a fixed 50/50 split?
basics
~20 sA fixed 50/50 split's cumulative regret grows linearly with traffic: 50,000 impressions parked on a 4% arm instead of a 5% arm costs about 500 conversions. Thompson sampling's grows sublinearly, roughly like the logarithm of the horizon.
Why does variational inference minimise KL(q||p) rather than KL(p||q)?
basics
~20 sKL(q||p) takes expectations under the approximation you control, so it is computable up to a constant; the reverse direction needs expectations under the unknown posterior. The price is mode-seeking behaviour that ignores parts of a multimodal target.
How do you set the ship threshold on probability of superiority for a Bayesian experiment program?
basics
~20 sThe threshold should encode the cost of being wrong, not a borrowed convention. Set it tighter for expensive, hard-to-reverse changes and looser for cheap reversible ones, pair it with a magnitude requirement, and fix it before the experiment runs.
How would you decide whether a team reports Bayesian or frequentist results for its recurring decisions?
basics
~20 sDecide by the shape of the decisions, not by taste. Bayesian reporting pays off when data per decision is thin, credible prior information exists, and stakeholders need a probability about the claim itself. Frequentist reporting suits high-volume standardised readouts.
When is the convenience of a conjugate prior not worth the constraint it puts on your model?
basics
~20 sWhen the family cannot express the belief or the structure the problem has. Closed form buys exact, constant-memory updates worth keeping at high throughput, but bending a bimodal belief or a covariate-driven model into a convenient family is a modelling error.
How do you handle a Bayesian efficacy readout whose conclusion flips between a sceptical and an enthusiastic prior?
basics
~20 sReport the flip as the finding: if the decision changes across priors reasonable people hold, the data are not decisive. Pre-specify the prior set, quantify the tipping point, and decide on the cost of being wrong.
How do you decide between variational inference and MCMC for a production Bayesian model?
basics
~20 sDecide by which error you can afford. Variational inference is fast and scales, but its approximation error is biased and unmeasured; sampling is asymptotically exact but slow. Match the choice to how sensitive the downstream decision is to uncertainty.
Why do two analysts with different reasonable priors reach nearly the same conclusion as data grows?
basics
~20 sEach observation multiplies more likelihood into the posterior, while the prior enters only once. With enough data the likelihood swamps it, so both posteriors concentrate on the same value and take the same approximately normal shape — the Bernstein-von Mises result.
How does a Gamma prior on a support-ticket rate update after observing daily counts?
basics
~20 sEvents add to the shape, exposure to the rate parameter: in shape-and-rate form, Gamma(alpha, beta) with y tickets over n days becomes Gamma(alpha + y, beta + n). The prior reads as alpha pseudo-events over beta pseudo-days.
Does thinning an MCMC chain to every 10th draw improve the posterior estimates?
basics
~20 sNo. For a fixed number of iterations, keeping every tenth draw throws away information, so estimates are no better and usually slightly noisier than using all draws. Thinning is justified by storage or downstream cost, not accuracy.
Why is the Jeffreys prior for a binomial proportion Beta(1/2, 1/2) rather than uniform?
basics
~10 sJeffreys' rule sets the prior proportional to the square root of the Fisher information. For a binomial proportion that gives the Beta(1/2, 1/2) density. The motivation is invariance under reparameterisation, not neutrality.
How does the Laplace approximation build a Gaussian approximation to a posterior?
basics
~20 sIt expands the log posterior to second order around its peak. The linear term vanishes there, leaving a quadratic, which is the log of a Gaussian centred at the peak whose covariance is the inverse of the negative second-derivative matrix.
Does checking P(B > A) every morning inflate error the way repeated p-value checks do?
basics
~20 sThe posterior itself needs no correction for how often you look: it summarises the data in hand at any moment. But a rule that stops the first time the probability crosses a bar still selects lucky data and overstates the lift.
Two coin experiments with different stopping rules give proportional likelihoods — why is the posterior identical?
basics
~20 sTwo likelihood functions that differ only by a factor free of the parameter give the same posterior, because that factor is absorbed by the normalising constant. With the same prior the two experiments' posteriors coincide exactly.
Why is a MAP estimate not invariant under reparameterisation while the MLE is?
basics
~20 sThe posterior is a probability density, so changing variables multiplies it by a Jacobian that reshapes the curve and can move its peak. The likelihood is not a density over the parameter, gets no Jacobian, so its maximiser transforms along.
What does Hamiltonian Monte Carlo's momentum variable buy over random-walk Metropolis proposals?
basics
~20 sMomentum lets a proposal travel a long way while following the posterior's shape, so moves are coherent strides rather than a random walk. Distant proposals still get accepted, which is why exploration holds up as the number of parameters grows.
Would you standardise on equal-tailed or highest-density credible intervals across your team's readouts?
basics
~20 sMake equal-tailed the default, because it is reproducible, survives rescaling and is always one interval. Require the highest-density version, plus the posterior plot, in named exception cases: strongly skewed posteriors, posteriors piled against a boundary, and multimodal ones.
How do you choose which test statistics to compare in a posterior predictive check?
basics
~20 sChoose statistics that matter for the decision the model supports and that the fitting did not already force to match. A statistic the fit targets, such as the sample mean, passes regardless of how wrong the model is.
How would you run Thompson sampling on a news feed where articles arrive and expire daily?
basics
~20 sThompson sampling handles arms coming and going without modification: each request draws over whatever arms are live, and a new arm's wide posterior earns it exploration. The real problems are cold-start priors, delayed clicks, and exploration swamping exploitation.
showing 31–60 of 60