skip to content

Why do two analysts with different reasonable priors reach nearly the same conclusion as data grows?

level: middleimportance: nice to knowfreq 31%

answer

  1. count how many times each factor enters
  2. prior once, likelihood n times
  3. log posterior grows linearly in n
  4. approximately normal, spread like 1/sqrt(n)
  5. zero prior probability never recovers

basics

~20 s

Each observation multiplies more likelihood into the posterior, while the prior enters only once. With enough data the likelihood swamps it, so both posteriors concentrate on the same value and take the same approximately normal shape — the Bernstein-von Mises result.

solid answer

~50 s

The prior is applied once; the likelihood accumulates with every observation, so the ratio of evidence to prior information grows with the sample. Under standard regularity conditions — a correctly specified model with a fixed, finite number of parameters, and priors that assign positive density near the true value — the Bernstein-von Mises theorem says the posterior becomes approximately normal, centred near the frequentist estimate, with spread shrinking on the order of `1/sqrt(n)`. Two analysts starting from different but reasonable priors therefore end up in practically the same place, and in the same place a frequentist analysis lands. The important caveats are that a prior assigning zero probability to a region can never recover it no matter how much data arrives, and that the result assumes a well-behaved fixed-dimensional model, so it says nothing about nonparametric or badly misspecified settings.

go deeper

for a junior

Remember the shape of the argument: the prior counts once, the data counts many times, so with enough observations the data wins. Being able to say that in one sentence is enough here.

for a middle

Explain the mechanism and name Bernstein-von Mises with its conditions: a well-specified fixed-dimensional model and a prior that does not assign zero probability near the truth. Mention the 1/sqrt(n) rate for the posterior's spread.

for a senior

Use the result to decide where scrutiny belongs. Show you can spot the cases it excludes — rare events with small effective samples, thin subgroups, misspecified or growing-dimension models — rather than assuming any prior eventually stops mattering.

for a principal

Frame it for the organisation: on high-volume metrics the framework debate is not worth the cost, so standardise and move on, and reserve genuine review effort for thin-data launches where the prior visibly drives the decision.

## The counting argument The intuition needs no theorem. The posterior is proportional to prior times likelihood. The prior appears exactly once, whatever the sample size. The likelihood is a product over observations, so with `n` independent observations you multiply in `n` factors of evidence. As `n` grows, the prior becomes an ever-smaller share of the total information in that product. A prior expressing a moderate opinion is decisive when `n = 20`, mildly influential at `n = 500`, and invisible at `n = 500,000`. Another way to see it: take logs. The log posterior is the log prior plus the sum of `n` log-likelihood terms. The first is a fixed function; the second grows linearly in `n`. Whatever shape the prior contributes gets flattened by a term that keeps getting steeper. ## What the theorem actually says The formal statement is the **Bernstein-von Mises theorem**. In a well-behaved parametric model, for a prior with positive continuous density in a neighbourhood of the true parameter value, the posterior distribution of the parameter converges — in total variation distance — to a normal distribution centred near the efficient frequentist estimator, with covariance equal to the inverse Fisher information divided by `n`. Three things worth extracting from that: 1. **The prior disappears from the limit.** It appears in the conditions (positive density near the truth) but not in the limiting distribution. Any two priors satisfying the condition give asymptotically the same posterior. 2. **The Bayesian and frequentist answers coincide asymptotically.** The limiting spread is the same `1/sqrt(n)`-scale uncertainty a frequentist analysis reports. The two frameworks are not asymptotically in competition; they are asymptotically the same arithmetic with different narration. 3. **The rate is `1/sqrt(n)`.** Quadrupling the sample halves the spread. This is why the practical crossover from prior-dominated to data-dominated arrives quickly for common quantities, and slowly for rare events where the *effective* sample is the count of events, not the count of rows. ## Where the prior refuses to wash out The conditions are not decoration. Convergence fails when: - **The prior assigns zero probability to the truth.** Multiplying zero by any amount of likelihood leaves zero. If your prior rules out rates above 10% and the true rate is 15%, no volume of data rescues you. This is sometimes called Cromwell's rule: leave a little probability everywhere you are not certain to be impossible. - **The model is misspecified.** The theorem assumes the data really was generated by some member of your parametric family. Under misspecification the posterior still concentrates, but on the closest wrong answer, and two analysts may not agree on what "closest" means once their models differ. - **The parameter count grows with the sample.** Add a new parameter per group or per user and the effective data per parameter stops growing. Bernstein-von Mises is a fixed-dimension result and simply does not apply. Nonparametric and high-dimensional settings have their own, much weaker, guarantees. - **The parameter is not identified.** If two parameter values imply exactly the same distribution over data, no amount of data separates them, and the prior determines the answer forever. - **Data is thin where it matters.** A million rows overall is small comfort for a subgroup that contributes forty of them. ## Why this matters in an interview It reframes the framework argument. If someone insists the choice between Bayesian and frequentist analysis is a deep philosophical fork that changes every number, the honest response is that in large, well-specified problems it changes almost nothing — the answers converge, and the remaining difference is what you are permitted to *say* about them. The choice becomes practically consequential in exactly the opposite regime: small samples, rare events, cold starts, thin subgroups, and situations where genuine outside information exists and should be allowed in. So the convergence result is not a reason to stop caring about priors. It is a map of where caring is warranted. It tells you that arguing about a weakly informative prior on a high-traffic metric is wasted effort, and that scrutinising a prior on a low-traffic launch is exactly where your attention belongs. ## A crisp way to say it "The prior enters once and the likelihood enters `n` times, so evidence outgrows the prior. Bernstein-von Mises makes that precise: in a well-behaved fixed-dimensional model with a prior that does not rule out the truth, the posterior goes approximately normal around the frequentist estimate at rate `1/sqrt(n)`. The exceptions are priors that assign zero probability where they should not, misspecified models, and growing parameter counts." That answer states the intuition, names the result correctly, and shows the boundary — which is what separates recall from understanding.

  • Which kind of prior never gets washed out by data?
    One that assigns zero probability to the true value. Multiplying zero by any likelihood stays zero, so no sample size recovers a region the prior excluded — Cromwell's rule. Non-identified parameters behave similarly: if two values imply identical distributions over data, the evidence cannot separate them and the prior settles the answer permanently.
  • Does convergence mean the Bayesian versus frequentist choice never really matters?
    No, it maps where it matters. The choice is nearly irrelevant on large, well-specified problems where both land in the same place. It matters when data is thin relative to model complexity, in rare subgroups and cold starts, when genuine outside information exists, and whenever you need a probability statement about the parameter itself rather than about a procedure.
  • How does the convergence rate behave as the sample grows?
    The posterior's spread shrinks on the order of `1/sqrt(n)`, so quadrupling the sample halves it. The practical caveat is that the relevant `n` is the effective information, not the row count: for a rare outcome it is the number of events, so a million sessions with twenty conversions is a small-sample problem despite the headline size.

saying these in an interview costs you the question

  • Claims any prior washes out given enough data
  • Thinks the prior is re-multiplied for every observation
  • Says the two frameworks always disagree, regardless of sample size
  • Applies the convergence result to models whose parameter count grows with n
  • Treats large row counts as large samples for a rare outcome

context