How do you decide between variational inference and MCMC for a production Bayesian model?
answer
- which error can you afford
- bias you cannot measure versus variance you can
- means survive, tails do not
- fast path plus periodic reference run
basics
~20 sDecide by which error you can afford. Variational inference is fast and scales, but its approximation error is biased and unmeasured; sampling is asymptotically exact but slow. Match the choice to how sensitive the downstream decision is to uncertainty.
solid answer
~50 sFrame it as a decision about error, not about method preference. Sampling-based inference is asymptotically exact — run it long enough and the answer converges to the true posterior — but on a 50-million-row model a full run can take days, which rules it out of an hourly refit. Variational inference returns calibrated-enough posterior means in minutes and scales by subsampling, at the cost of a bias whose size you cannot measure, typically understated spread. So ask what the posterior actually feeds. If it drives point predictions, rankings or a recommendation ordered by posterior mean, the bias is largely harmless. If it drives a threshold, a risk statement or a decision priced off the tail, understated uncertainty converts directly into over-confident decisions. The mature answer is usually a hybrid: variational fits for routine refits, plus a periodic sampling-based audit on a subsample to bound how wrong the fast path is.
go deeper
Know the headline tradeoff: sampling is slow but converges to the true posterior, while variational inference is fast but leaves an approximation error that never disappears.
Be ready to explain why the variational error is bias rather than variance, and why more optimisation iterations cannot remove it when the posterior lies outside the chosen family.
Show that you validate the fast path against a reference fit on a subsample and that you know which downstream quantities — thresholds, tails, joint statements — the approximation damages most.
Own the error budget for the decision the model serves, define the hybrid arrangement and its monitoring, and state in advance the conditions that send the team back to the exact method.
## The two error budgets The honest framing of this choice is not "which is better" but **which error you can afford to carry**. Sampling-based posterior computation is *asymptotically exact*: as the chain runs longer, its draws characterise the true posterior arbitrarily well. Its error is therefore mostly **variance** — noise that shrinks with more compute — plus whatever bias remains from a chain that has not settled. You can spend money to reduce it, and you have established checks that tell you when you have not spent enough. Variational inference is *not* asymptotically exact. Running the optimiser longer converges to the best member of the chosen family, and if the true posterior is not in that family, the remaining error never goes away. Its error is therefore **bias**, and — this is the point interviewers push on — the size of that bias is the gap between the bound and the log evidence, which is exactly the quantity you could not compute. **You cannot spend compute to measure it.** ## The scale argument Consider a model over 50 million rows that has to be refit on a schedule. Sampling requires repeated passes over the likelihood and produces a stream of correlated draws; the wall-clock cost of a defensible run can reach days. If the refit cadence is hourly or nightly, that is not a slow option, it is a **non-option**. Variational inference optimises an objective that can be estimated from minibatches, so the cost per iteration is decoupled from the dataset size, and a fit that returns in minutes is achievable. Where the schedule and the sampling cost are incompatible, the decision is made for you; the remaining question is what you must do to make the fast path safe. ## What the posterior is used for This is the question that should drive the call, and it is the one weak candidates skip. - **Point predictions and rankings.** If a downstream system consumes posterior means — ordering items, producing a forecast, filling a dashboard — the characteristic variational bias hits the part of the answer you are not using. This is the strongest case for the fast path. - **Thresholded and risk-sensitive decisions.** If the decision is "act when the probability the effect exceeds zero passes 0.95", or a reserve priced off an upper quantile, understated spread pushes probabilities towards the extremes and systematically produces more confident decisions than the data support. Here the bias lands squarely on the quantity that matters. - **Joint and derived quantities.** Anything combining several parameters — a difference, a sum, a portfolio-style aggregate — depends on the dependence structure a simple approximating family may have discarded, and can be wrong in either direction. - **Model comparison.** Comparing fitted objective values across models compares bounds of unknown tightness, so it is a weak basis for a consequential model choice. ## The hybrid pattern The answer a lead is expected to give is rarely "pick one". A durable arrangement looks like: 1. **Fast path for production refits.** Variational fits on the full data at the required cadence. 2. **Periodic reference run.** A sampling-based fit on a subsample, an earlier snapshot, or a deliberately simplified model, run offline on whatever schedule the compute budget allows. 3. **A comparison you can act on.** Store both sets of posterior summaries and track the ratio of spreads and the shift in means. A stable, known shrinkage factor is something you can correct for or at least disclose; a factor that drifts is a signal the fast path has stopped being safe. 4. **Uncertainty inflation or a guard rail.** If the reference run shows the fast path's spread is consistently too small, either widen the reported uncertainty by the measured factor or move the decision threshold to compensate. 5. **An escape hatch.** Name in advance the conditions — a new model structure, a regime change, a decision whose stakes rise — under which the team goes back to the slow method before shipping. ## Other considerations that decide close calls - **Model structure.** Strong correlation between parameters, multimodality or weak identification are the conditions under which the approximation is most misleading, and they are also conditions under which sampling gets slow. Reparameterising the model often improves both. - **Reproducibility and audit.** A deterministic optimisation with fixed initialisation is easy to reproduce; whether that matters depends on your regulatory and review environment. - **Team capability.** Diagnosing a badly behaved sampler and diagnosing a badly behaved optimiser are different skills. A method nobody on the team can debug at 2am is a worse choice than a slightly less accurate one they can. - **Failure mode when things go wrong.** A sampler that has not converged usually announces itself through its standard checks. A variational fit that has converged to a poor optimum looks completely healthy — the objective is flat, the numbers are finite, the fit returns fast. **Silent failure is the real risk**, and it argues for building the comparison in step 3 rather than trusting the fast path unmonitored. ## The summary to say out loud Start from the decision the posterior serves, not the method. Establish what error that decision tolerates. Then choose the cheapest method whose error fits inside that tolerance, and — because the variational error is unmeasurable from inside — build a reference comparison that keeps the assumption honest over time.
- How would you quantify how wrong the fast variational path is?You cannot measure it from inside the fit, so you build an external reference: a sampling-based posterior on a subsample, an earlier snapshot, or a simplified model, and compare posterior means and spreads against it. A stable shrinkage factor can be corrected for or disclosed; a drifting one is a signal to stop trusting the fast path.
- Which failure mode worries you more in production, and why?A variational fit that converged to a poor optimum, because it looks entirely healthy — objective flat, numbers finite, runtime short. A poorly behaved sampler tends to advertise itself through its standard checks. Silent, confident wrongness is harder to catch than loud failure, so the monitoring effort belongs on the fast path.
- When would you refuse the fast path outright, regardless of runtime pressure?When the decision is priced off the tail or a probability threshold, when the posterior is plausibly multimodal or weakly identified, or when a regulator or reviewer needs an approximation whose error is bounded. In those cases understated or partial uncertainty translates directly into wrong decisions, and the runtime saving is not worth it.
saying these in an interview costs you the question
- Picks a method without asking what the posterior is used for
- Says variational inference just needs more iterations to be exact
- Treats a fast converged fit as validated
- Ignores that the approximation error cannot be measured internally
- Proposes sampling on a schedule the compute budget cannot meet