In Bayes' rule for a parameter, what does 'posterior is proportional to prior times likelihood' mean?
answer
- one number is missing from the equation
- the denominator does not depend on theta
- shape from the product, scale from normalising
- posterior ratios cancel the denominator
- multiply on a grid, then rescale to area 1
basics
~20 sIt means the posterior density at each parameter value is prior times likelihood divided by a single constant, the evidence. That constant does not depend on the parameter, so the product alone determines the posterior's shape.
solid answer
~50 sBayes rule for a parameter reads `p(theta | data) = p(theta) * L(theta) / p(data)`, where `L(theta)` is the likelihood of the observed data and `p(data)` is the evidence, the integral of prior times likelihood over all parameter values. Because `p(data)` is a single number that does not vary with `theta`, it cannot change which parameter values look better or worse — it only rescales the whole curve so it integrates to 1. So the informative work is done by the product: the likelihood pulls mass toward values that explain the data, the prior downweights values you considered implausible beforehand, and the posterior ends up concentrated where both are non-negligible. Practically, this is why you can evaluate prior times likelihood on a grid of parameter values and normalise afterwards, and why posterior ratios between two values equal the ratio of their prior-times-likelihood products.
go deeper
Be able to write the full rule, name each of the four pieces, and say which one is missing from the proportional version. Knowing that the denominator is a constant with respect to the parameter is the core recall item.
Expect to explain why a constant denominator cannot change relative plausibility, and to show the cancellation in a ratio of posteriors at two parameter values. Be ready to sketch grid evaluation followed by normalisation.
Demonstrate what the product does in practice: regions killed by a zero prior, prior-data conflict showing up as a product that is small everywhere, and why a posterior can be tighter than both inputs. Interviewers want to hear diagnosis, not recitation.
Own the tradeoff between transparent numerical normalisation and heavier approximations as models grow, and set the team norm on when an unnormalised surface may be reported and when a fully normalised posterior is required for downstream decisions.
## The statement, written out For an unknown parameter `theta` and observed data, Bayes rule gives `p(theta | data) = p(theta) * p(data | theta) / p(data)` Read the pieces: - `p(theta)` is the prior density: what you believed about the parameter before this data arrived. - `p(data | theta)`, viewed as a function of `theta` with the data held fixed, is the likelihood `L(theta)`. - `p(data)` is the evidence, also called the marginal likelihood: `p(data) = integral of p(theta) * L(theta) over all theta`. - `p(theta | data)` is the posterior: the updated density over the parameter. The proportionality statement drops the denominator: `p(theta | data) is proportional to p(theta) * L(theta)` ### Why dropping the denominator is legitimate The evidence is computed by integrating `theta` out. Once that integral is done, `theta` no longer appears in it — it is a single number attached to the dataset, not a function of the parameter. A quantity that is the same at every parameter value cannot make one value more probable than another. All it does is scale: it is exactly the number that turns the unnormalised product into something that integrates to 1. A sharp way to see this is to take a ratio. For two candidate values `a` and `b`, `p(a | data) / p(b | data) = [p(a) * L(a)] / [p(b) * L(b)]` The evidence cancels completely. Every statement of the form *this value is k times more plausible than that one* can be made without ever computing it. ### Shape versus scale It helps to separate two things a function carries. Its **shape** says where the mass is — where the peak sits, how wide the bulk is, whether there are several humps. Its **scale** is one multiplicative number. Bayes rule says: the prior and the likelihood jointly determine shape, and normalisation supplies scale. Because the likelihood itself is only defined up to a constant, and the normaliser absorbs any such constant, the whole calculation is indifferent to constants on both sides. ### What the product actually does Multiplication is the right mental image. At each parameter value you multiply two numbers: - Where the likelihood is tiny because the data would be surprising under that value, the product is tiny regardless of what the prior said. - Where the prior is tiny because you regarded that region as implausible, the product is tiny regardless of how well the data fit. - The posterior lives where both are appreciable. Two consequences follow immediately. A prior that assigns exactly zero density to a region makes the posterior zero there forever, because zero times anything is zero — no amount of data can rescue a region you excluded outright. And when the prior and the likelihood are concentrated in strongly disagreeing places, the product is small everywhere and the posterior can end up spread or split; that is a signal of prior-data conflict worth investigating rather than reporting. ### Doing it numerically The proportional form is also how you compute a posterior when no tidy formula exists. Lay a fine grid over the plausible range of `theta`, evaluate prior times likelihood at each grid point, then divide every value by the sum of all values times the grid spacing. That divisor is a numerical approximation of the evidence integral, so you have not really escaped it — you have obtained it as a by-product of normalising. This grid recipe is transparent and fine in one or two parameters; its cost grows with the number of parameters, which is why higher-dimensional problems need different machinery. ### Common misstatements to avoid - Calling prior times likelihood a probability density before normalising. It is an unnormalised surface; its values are not probabilities and its area is arbitrary. - Saying the posterior is an average or a blend of prior and likelihood. It is a product, then a rescale; a product can be much more concentrated than either factor, which an average never is. - Believing the proportional form only holds in special models. It is the general statement of Bayes rule for a parameter; specific model families merely make the product easy to recognise. - Claiming the evidence is unimportant because it can be ignored. It can be ignored *inside one model*, where it is a constant. The moment you compare two models it stops being a shared constant and becomes the quantity of interest.
- How would you get a posterior numerically when the product has no recognisable closed form?Put a grid over the plausible range of the parameter, evaluate prior times likelihood at each point, and divide by the sum of those values times the grid spacing. That divisor is a numerical estimate of the evidence integral. The approach is transparent in one or two parameters and becomes impractical as the parameter count grows.
- What happens if the prior assigns exactly zero density to some region of the parameter space?That region has zero posterior density no matter what the data say, because the posterior is a product and zero times any likelihood is zero. Hard exclusions in a prior are permanent, which is why analysts prefer to give implausible regions small density rather than none unless the exclusion is a genuine constraint such as a variance being non-negative.
- Can the posterior be more concentrated than both the prior and the likelihood?Yes, and that is normal — multiplication is not averaging. Two moderately wide curves that overlap in the middle produce a product that is narrower than either, because values in the tails are penalised twice. This is the sense in which prior and data combine rather than compromise.
saying these in an interview costs you the question
- Calls prior times likelihood a probability density before normalising
- Describes the posterior as an average or blend of prior and likelihood
- Says the evidence changes which parameter values are favoured
- Believes the proportional form only applies to special model families
- Thinks the likelihood must be normalised before it is multiplied in