skip to content

Why does updating a Beta posterior event-by-event match one batch update of the same data?

level: seniorimportance: should knowfreq 46%

answer

  1. yesterday's posterior is today's prior
  2. the likelihood factorises across observations
  3. multiplication does not care about order
  4. only running counts ever enter
  5. assumes the parameter never moves

basics

~20 s

Because the posterior becomes the next prior and the update only adds sufficient statistics. For exchangeable data the likelihood factorises and multiplication commutes, so any batching or ordering accumulates the same success and failure counts and lands on identical parameters.

solid answer

~50 s

Bayes' rule applied twice is the same as applying it once to the pooled data: `prior * L(batch1) * L(batch2)` is the same product however you group or order the factors. For a Beta-binomial pair the update is literally addition of counts, so 100 observations fed one at a time, in ten batches of ten, or shuffled all land on the same `Beta(a + s, b + f)`. That is why an online updater needs only two numbers of state rather than the dataset. The caveat is that order-invariance is a property of the *model*, not of reality: it holds because the model assumes the parameter is fixed and the observations exchangeable. If the true rate drifts, an old observation still carries full weight and the posterior lags badly — the fix is explicit down-weighting of old data, not a different update order.

go deeper

for a junior

Be able to state that the posterior from one batch becomes the prior for the next, so grouping the data differently does not change where you end up.

for a middle

Explain the mechanism: the likelihood factorises across independent observations and the update adds sufficient statistics, so multiplication order is irrelevant.

for a senior

Use it in production. Treat a mismatch between streaming and batch as a data bug, and recognise that a drifting rate needs explicit decay because the update itself is order-blind.

for a principal

Decide the memory policy. Choose how long evidence stays fully weighted, and justify the lag-versus-variance tradeoff that any decay or windowing scheme imposes.

## The algebra Split your data into two batches, `D1` and `D2`, and assume the observations are independent given the parameter theta. Update once with `D1`: `p(theta | D1)` is proportional to `p(theta) * L(theta; D1)` Now treat that as the prior and update with `D2`: `p(theta | D1, D2)` is proportional to `p(theta) * L(theta; D1) * L(theta; D2)` which is exactly what you get by updating the original prior with the pooled data in one step, since `L(theta; D1, D2) = L(theta; D1) * L(theta; D2)` under conditional independence. Multiplication is commutative and associative, so the grouping and the order of the factors are irrelevant. This is the formal content of "today's posterior is tomorrow's prior". ## What it looks like with counts For a Beta prior with binomial data the whole thing collapses to addition. Start at `Beta(1, 1)` and take 100 impressions containing 8 clicks: - All at once: `Beta(1 + 8, 1 + 92) = Beta(9, 93)`. - One at a time: each click does `a += 1`, each non-click does `b += 1`. After 100 events, `Beta(9, 93)`. - Ten batches of ten, in any order: each batch adds its own successes and failures. After all ten, `Beta(9, 93)`. - Shuffled: addition does not care about order. `Beta(9, 93)`. The only quantities that ever enter are the running totals — the sufficient statistics. Everything else about the data, including its arrival order, is discarded by the model because the model says it carries no information about theta. ## Why this is operationally valuable - **Constant state.** You store two numbers per item, not a growing event log, and memory does not scale with traffic. - **Restartability.** A crashed process resumes from the stored parameters; there is no replay to coordinate. - **Reconciliation.** A streaming updater and a nightly batch job over the same events must agree exactly. If they disagree, you have a data bug — duplicate events, dropped events, a window boundary counted twice — not a statistical subtlety. This makes order-invariance a genuinely useful test in production. - **Parallelism.** Two workers can accumulate counts on disjoint slices and merge by adding, because the merge is the same addition. The exactness matters here: this is not "approximately the same", it is the same integers. ## Where the property becomes a liability Order-invariance follows from the model's assumption that theta is fixed and the observations exchangeable. Reality frequently disagrees. Suppose a rate genuinely changed after a release. A posterior accumulated since launch weights the first month's events exactly as heavily as yesterday's, so it converges to a blend of the old and new regimes and moves toward the new rate only as fast as new data can outweigh the accumulated old counts. After a long history that is glacially slow. The posterior is not wrong given its assumptions; the assumptions are wrong. There is no ordering trick that fixes this, because the update is provably order-blind. The fixes are model changes: - **Down-weight the past.** Multiply the accumulated parameters by a factor slightly below one at each step, so old evidence decays and the effective window stays bounded. This deliberately breaks exchangeability and makes the result order-dependent — that is the point. - **Window the data.** Update only from a trailing window, accepting more variance for less lag. - **Model the change.** Fit a parameter that varies over time, or reset at a known intervention point such as a deploy. Each costs something. Decay and windows throw away real information and widen your intervals; a time-varying model gives up the closed-form conjugate update. ## A subtlety about "same" Exchangeability is an assumption about the *model*, not a claim that the order of your data is meaningless. If arrival order encodes something — a user's first session behaves differently from their tenth, a queue is served worst-first — then the observations are not exchangeable, and a model that ignores order is discarding information. Order-invariance of the update is then a symptom of a misspecified model rather than a feature. ## What to say in an interview Give the algebra in one line, show the counting version, and then volunteer the caveat before you are asked. "Identical parameters, because the posterior is the next prior and the update is addition of sufficient statistics — which also means a drifting rate will be tracked far too slowly unless I add explicit decay." That combination of the mechanism and its failure mode is what separates a memorised answer from an operated one.

  • Your streaming updater and a nightly batch job report different posteriors; what does that tell you?
    That you have a data bug, not a statistical one. Conjugate updating is exactly order- and batch-invariant, so identical event sets must give identical parameters. Look for duplicated events, dropped events, a window boundary counted in both jobs, or a late-arriving record excluded from one path. The invariance makes a good production assertion for precisely this reason.
  • The rate genuinely doubled last month but the posterior has barely moved; why?
    Because every historical observation still carries full weight. The model assumes one fixed rate, so months of old counts dominate the recent ones and the posterior converges to a blend that shifts only as new data accumulates. Fix it by down-weighting old evidence, updating from a trailing window, or resetting at the known change point — not by reordering the input.
  • Does order-invariance still hold if you use a decay factor on the accumulated counts?
    No, and deliberately so. Multiplying the parameters by a factor below one at each step makes recent observations count for more, which is exactly what you want under drift but which breaks exchangeability. The result then depends on order, and a batch job that applies decay per batch will not match a per-event streamer unless the decay schedule matches too.

It is a running tally on a scoreboard. Whether you record the goals one by one or write in the final count at the whistle, the number on the board is the same — and it is equally blind to whether a team improved during the match.

saying these in an interview costs you the question

  • Says batch updating is only approximately equal to streaming
  • Thinks feeding data one at a time overweights recent events
  • Claims reordering the input can fix a drifting rate
  • Cannot say which assumption makes the order irrelevant
  • Stores the full event log when two counters suffice

context