In Normal-Normal conjugate updating, how does the posterior mean combine prior and data?
answer
- work in one over variance
- each source weighted by its reliability
- the two precisions add up
- result sits between the two means
- posterior is narrower than both inputs
basics
~20 sAs a precision-weighted average, where precision is one over variance. Precisions add, so the posterior is always more precise than the prior, and its mean lies between the prior mean and the sample mean, closer to whichever is more precise.
solid answer
~50 sTake a `Normal(m0, t0^2)` prior on an unknown mean and `n` observations with known sampling variance `s2`. Work in precision, the reciprocal of variance. The prior carries precision `1 / t0^2` and the data carries precision `n / s2`. The posterior precision is simply their sum, `1 / t0^2 + n / s2`, and the posterior mean is the precision-weighted average `(m0 / t0^2 + n * ybar / s2) / (1 / t0^2 + n / s2)`. Two consequences fall out. First, the posterior is always at least as precise as the prior — information never subtracts, so the interval never widens on new data. Second, the estimate sits between the prior mean and the sample mean, dominated by whichever source has the smaller variance. As `n` grows, `n / s2` swamps the prior precision and the posterior mean converges on `ybar`.
go deeper
Be able to say the posterior mean lands between the prior mean and the sample mean, leaning toward whichever source is more trustworthy.
Write the formula. Prior precision one over prior variance, data precision n over sampling variance, add them, and weight each mean by its share.
Sanity-check the result: the posterior can never be wider than the prior, and correlated observations make the effective n smaller than the raw count.
Argue about where the prior variance comes from and who owns that number, since it silently decides how many observations of influence the prior gets.
## Setup You want to estimate an unknown mean, call it mu. You hold a prior `mu ~ Normal(m0, t0^2)`, and you take observations that are `Normal(mu, s2)` with the sampling variance `s2` treated as known. This last assumption is what keeps the problem conjugate in a single parameter; with the variance also unknown the conjugate story needs a joint prior on mean and variance, which is a heavier object. ## Precision is the natural currency Define **precision** as the reciprocal of variance. A tight, confident distribution has high precision; a vague one has low precision. Two facts make precision the right coordinate for this problem: - The prior contributes precision `1 / t0^2`. - A sample of `n` observations contributes precision `n / s2`, because the sample mean `ybar` has variance `s2 / n`. The conjugate update is then stated in one line each: `posterior precision = 1 / t0^2 + n / s2` `posterior mean = (m0 / t0^2 + n * ybar / s2) / (1 / t0^2 + n / s2)` The posterior is again Normal, which is what conjugacy means here. ## Reading the formulas **Precisions add.** Information accumulates. The posterior variance is `1 / (1 / t0^2 + n / s2)`, which is strictly smaller than both `t0^2` and `s2 / n`. Two mediocre measurements combine into something better than either. A posterior that came out wider than the prior would be a sign you had made an arithmetic error, not a discovery. **The mean is a weighted average.** Rewrite it as `w * m0 + (1 - w) * ybar` with `w = (1 / t0^2) / (1 / t0^2 + n / s2)`. The weight on the prior is its share of the total precision. So the posterior mean always lies between the prior mean and the sample mean, never outside them, and it leans toward whichever is more precise. **Data eventually wins.** The data precision grows linearly in `n` while the prior precision is fixed, so `w` tends to zero and the posterior mean tends to `ybar`. A fixed prior can bias a small-sample answer noticeably and a large-sample answer barely at all. ## A worked example Suppose you believe a package weighs about 500 grams, with a standard deviation of 10 grams on that belief: prior `Normal(500, 10^2)`, precision `1 / 100`. You put it on a scale known to have a measurement standard deviation of 10 grams and read 520 grams: data precision `1 / 100` as well. Equal precisions means equal weights, so the posterior mean is `(500 + 520) / 2 = 510` grams. The posterior precision is `1/100 + 1/100 = 2/100`, so the posterior variance is 50 and the posterior standard deviation is about 7.1 grams — narrower than either the 10-gram prior or the 10-gram reading. Now change one number. If the scale is far better, standard deviation 2 grams, its precision is `1 / 4`, twenty-five times the prior's. The posterior mean becomes `(500/100 + 520/4) / (1/100 + 1/4) = (5 + 130) / (0.26) = 519.2` grams — almost the reading, because the scale is the more precise source. If instead the scale is terrible, standard deviation 50 grams, its precision is `1 / 2500` and the posterior barely moves off 500. That is the whole intuition: **each source is heard in proportion to how precise it is.** ## The pseudo-count reading The Beta-binomial pair reads its prior as pseudo-successes and pseudo-failures. The Normal-Normal pair has the same interpretation in different units: the prior's precision `1 / t0^2` equals `k / s2` for `k = s2 / t0^2`, so the prior is worth `k` pseudo-observations at the sampling variance. In the first version of the example above, prior variance equals sampling variance, so the prior is worth exactly one observation — which is why a single reading moved the estimate half way. ## Where it breaks The update assumes the sampling variance is known and that the observations really are independent draws around a fixed mean. If `s2` is estimated from the same data, the posterior is overconfident, because you have ignored the uncertainty in the variance. If the observations are correlated — repeated readings from a scale with a systematic offset, say — the effective sample size is smaller than `n`, and using `n / s2` as the data precision overstates the information by a factor that can be large. If the true mean drifts over time, the model's assumption of a single fixed mu is wrong, and no amount of correct arithmetic rescues it. ## What to say in an interview Lead with "precision-weighted average, precisions add". Then show that the posterior variance is smaller than both inputs, note that the posterior mean is bracketed by the prior mean and the sample mean, and finish with the known-variance caveat. That sequence covers the mechanics, the sanity checks and the assumption in about a minute.
- Can the posterior variance ever exceed the prior variance in this model?No. The posterior precision is the sum of two non-negative precisions, so it is at least the prior precision, and the variance is its reciprocal. Data can move the mean anywhere between the prior mean and the sample mean, but it can never make you less certain. A widening interval means a modelling error or an arithmetic slip, not surprising data.
- How many observations is a Normal prior worth?Compare precisions: the prior's `1 / t0^2` equals `k / s2` when `k = s2 / t0^2`, so the prior is worth `k` observations at the sampling variance. A prior standard deviation equal to the measurement standard deviation is worth exactly one reading; a prior ten times tighter is worth a hundred.
- What breaks if the sampling variance is estimated rather than known?The posterior comes out overconfident, because it treats an estimated quantity as certain. The honest fix is to put a prior on the variance too and work with the joint posterior, which no longer reduces to a single Normal. With a decent sample size the difference is small; with a handful of observations it is not.
Two witnesses describe the same event. You do not average their accounts equally; you weight the careful one more heavily, and having both leaves you surer than either alone.
saying these in an interview costs you the question
- Averages the prior mean and sample mean with equal weights
- Weights each source by its variance instead of its precision
- Claims the posterior can be wider than the prior
- Forgets the data precision scales with n
- Ignores that the sampling variance is assumed known