skip to content

When can you ignore the marginal likelihood in Bayes' rule, and when must you actually compute it?

level: middleimportance: should knowfreq 45%

answer

  1. constant with respect to what, exactly
  2. same inside a model, different across models
  3. prior predictive probability of the data
  4. a Bayes factor is a ratio of these
  5. averaging over the prior penalises spread

basics

~20 s

Inside a single model the marginal likelihood is a constant that only rescales the posterior, so you can ignore it. You must compute it to compare models, since a Bayes factor is a ratio of two marginal likelihoods.

solid answer

~50 s

The marginal likelihood, or evidence, is `p(data) = integral of p(theta) * L(theta) over the parameter space` — the probability the model assigns to the observed data, averaged over the prior. Within a single model it does not depend on the parameter, so anything driven by relative posterior values ignores it safely: the shape of the posterior, ratios between two parameter values, and any summary taken from the normalised posterior. It stops being ignorable the moment you compare two models, because then the two evidences differ and their ratio is the Bayes factor. That comparison is also where it bites: because the evidence averages the likelihood over the prior, spreading prior mass over badly-fitting parameter values lowers it, so a Bayes factor can be sensitive to a prior whose effect on the posterior is negligible.

go deeper

for a junior

Know the name, know it sits in the denominator of Bayes rule, and know that it does not depend on the parameter, which is why the proportional form works. Recognising the term marginal likelihood as the same thing as the evidence is enough here.

for a middle

Be ready to write the integral, explain why it is constant within a model, and name the case that breaks that convenience: a Bayes factor is a ratio of two evidences and nothing cancels.

for a senior

Show that you know the evidence averages over the prior rather than maximising, so it penalises models that spread prior mass over poorly-fitting values, and that a diffuse prior can wreck a Bayes factor while leaving the posterior intact.

for a principal

Own the governance question: whether the team reports model comparisons at all given their prior sensitivity, what sensitivity evidence must accompany a Bayes factor, and when predictive checks are the more defensible basis for choosing a model.

## What the quantity is Write Bayes rule for a parameter `theta` under a model `M`: `p(theta | data, M) = p(theta | M) * p(data | theta, M) / p(data | M)` The denominator `p(data | M)` is the marginal likelihood, usually called the evidence. It is defined by integrating the parameter out: `p(data | M) = integral of p(theta | M) * p(data | theta, M) over theta` Two readings are worth holding at once. Mechanically it is the normalising constant that makes the posterior integrate to 1. Interpretively it is the **prior predictive probability of the data**: before any data arrived, the model plus its prior implied a distribution over datasets, and this is the value that distribution places on the dataset you actually got. ## When it is safe to ignore Inside one model the evidence is a single number, identical at every parameter value. Everything that depends only on relative posterior values is untouched by it: - The **shape** of the posterior — where the peak is, how wide the bulk is, whether there are multiple modes. - **Ratios** of posterior density at two parameter values, where the constant cancels. - Any **point summary or interval** derived from the normalised posterior, since normalising restores the constant anyway. - **Numerical normalisation on a grid**: evaluate prior times likelihood at each grid point and divide by the total; that total, times the grid spacing, *is* an approximation of the evidence. You did not avoid it so much as obtain it as a by-product. This is why the working form `posterior is proportional to prior times likelihood` is enough for the overwhelming majority of single-model analysis. ## When you must compute it The evidence becomes the object of interest as soon as the comparison is between models rather than between parameter values. For two models `M1` and `M2`, the Bayes factor is `BF = p(data | M1) / p(data | M2)` Both terms are evidences. Nothing cancels: the number that was a shared constant inside one model is now the whole calculation. If you also carry prior probabilities over the models themselves, posterior model odds are prior odds times the Bayes factor, so again the evidences are required. ## Why the evidence penalises complexity Because the evidence averages the likelihood over the prior rather than maximising it, a model that spreads its prior over a large parameter space pays for it. Suppose a flexible model can fit the data superbly at one narrow region of parameter space but fits badly across the rest of the region its prior covers. The maximum likelihood is high, but the average is dragged down by all the parameter values that predicted poorly. A tighter model whose prior concentrates on decent parameter values can win the comparison even with a lower peak fit. This is often called an automatic Occam effect: it arrives from the arithmetic rather than from an added penalty term. ## The prior-sensitivity trap The same mechanism produces the classic gotcha. Take a prior and widen it substantially — spread the same total mass over a much wider range. Once you have plenty of data, the posterior barely notices, because the likelihood concentrates the mass anyway and the prior is nearly flat over the region that matters. The evidence, on the other hand, drops roughly in proportion to the widening, because most of the prior mass now sits over parameter values that explain the data poorly and drags the average down. So two analysts with essentially the same posterior can report materially different Bayes factors. The practical rule: a diffuse prior that is harmless for estimation is *not* harmless for model comparison, and any Bayes factor should be reported with the priors that produced it and a sensitivity check. ## Why it is hard The evidence is an integral over the whole parameter space, weighted by the prior. In one or two dimensions a grid handles it. As parameters multiply, the region where prior times likelihood is appreciable becomes a vanishing fraction of the space being integrated, so naive numerical integration wastes essentially all of its effort. The difficulty of this integral, rather than any conceptual subtlety, is the reason model comparison is treated as a harder task than estimation. ## What to say in an interview The crisp answer has three beats: it is a constant within a model so it drops out of everything relative; it is the entire quantity in a model comparison because a Bayes factor is a ratio of evidences; and it is an average over the prior, which makes it both a built-in complexity penalty and unusually prior-sensitive.

  • Why is the marginal likelihood called the prior predictive probability of the data?
    Because it averages the probability of the observed data over the prior: `integral of prior times likelihood`. Before any data arrive, model plus prior imply a distribution over possible datasets, and this quantity is the value that distribution assigns to the dataset you got. Higher means the model predicted your data better in advance.
  • How can a Bayes factor be very sensitive to a prior that barely affects the posterior?
    Widening a prior spreads mass over parameter values that fit the data poorly. The posterior renormalises and is almost unchanged once the likelihood dominates, but the evidence is an average over the prior, so those poorly-fitting regions drag it down. Two analysts with near-identical posteriors can therefore report very different Bayes factors.
  • Does normalising a posterior on a grid require the evidence?
    Yes, but you get it for free. Summing prior times likelihood across the grid and multiplying by the grid spacing approximates the evidence integral, and dividing every grid value by that total normalises the posterior. In one or two dimensions this is a perfectly serviceable estimate of the evidence itself.

saying these in an interview costs you the question

  • Says the evidence can be ignored when comparing two models
  • Thinks the marginal likelihood depends on the parameter value
  • Describes it as merely a constant with no interpretation
  • Assumes a diffuse prior is harmless for a Bayes factor
  • Confuses the average likelihood over the prior with the maximum likelihood

context