skip to content

What does a prior predictive check reveal before any data have been observed?

level: middleimportance: should knowfreq 33%

answer

  1. simulate before you fit
  2. same integral, prior instead of posterior
  3. judge on the observable scale
  4. four-metre humans, negative revenue

basics

~20 s

A prior predictive check simulates fake datasets from the prior and the likelihood, then asks whether those datasets are physically sensible. It catches priors that imply absurd observables, such as four-metre human heights or negative revenue, before any fitting happens.

solid answer

~50 s

The prior predictive distribution is the same integral as the posterior predictive, but taken against the prior: `p(y) = integral of p(y | theta) * p(theta) d(theta)`. You draw parameters from the prior, generate complete synthetic datasets from the likelihood, and look at them on the scale a domain expert understands. The value is that priors are specified on parameter scales nobody has intuition about — a coefficient, a log-scale intercept, a variance — while the simulated data land on a scale everyone can judge. If the simulated adult heights include four-metre people, or the simulated revenues are negative, the prior is on the wrong scale and you know it before touching the data. It is a cheap sanity check on the whole model, not just the prior: an impossible support or a broken link function shows up here too.

go deeper

for a junior

Know that you can simulate data from a model before fitting it, and that the point is to see whether those fake data are physically possible.

for a middle

Be able to write the prior predictive integral, describe the draw-then-simulate loop, and explain why a diffuse prior can still be a bad prior on the observable scale.

for a senior

Show that you use the check to debug the whole generative model — likelihood family, link, units — and that you summarise simulations with quantities a domain expert can actually judge.

for a principal

Argue for making prior predictive simulation a standard step before any expensive fit, and weigh its cost against the modelling errors it reliably catches early.

## The object Before any data arrive, a Bayesian model already makes a claim about what data are possible. That claim is the **prior predictive distribution**: ``` p(y) = INTEGRAL p(y | theta) * p(theta) d(theta) ``` It is structurally identical to the posterior predictive — average the likelihood over a distribution on the parameter — except the averaging distribution is the **prior** rather than the posterior. It answers: 'given everything I have assumed and nothing I have observed, what datasets does this model think it might see?' ## Why simulate instead of stare at the prior Priors are written on parameter scales that are hard to reason about. A standard deviation of 10 on a coefficient sounds modest until you realise the model applies it through an exponential link, at which point it implies effects spanning many orders of magnitude. Humans have no intuition about a prior on a log-scale intercept. They have excellent intuition about heights, revenues, wait times and counts. So the check pushes the prior forward onto the observable scale where intuition works: 1. Draw `theta` from the prior. 2. Simulate a full synthetic dataset from `p(y | theta)` — the same size, the same covariates, the same shape as the real study you plan to run. 3. Repeat, then plot or tabulate the simulated datasets and ask a domain expert whether they are remotely plausible. ## What it catches **Wrong scale.** A model of adult human height whose prior predictive routinely produces people four metres tall, or shorter than zero, is broadcasting that its prior is far too diffuse for the units in use. Nothing about that is visible from looking at the prior's parameters alone. **Impossible support.** Simulated negative revenue, negative durations or probabilities above one usually mean the likelihood family or the link function is wrong, not the prior. The check is a test of the whole generative model. **Vacuously flat priors.** A prior chosen to be 'uninformative' can be strongly informative on the observable scale, concentrating simulated data in the extreme tails. 'Flat is safe' is a misconception this check kills quickly. **Structural bugs.** If the simulation code cannot generate a dataset shaped like the real one, the model as written does not describe the data you have. Better to learn that in five minutes of simulation than after a long fit. ## Prior predictive versus posterior predictive Same integral, different weighting distribution, different job: - **Prior predictive**, taken against the prior, runs **before** the data are used. It asks whether the assumptions are sane and whether they are on the right scale. - **Posterior predictive**, taken against the posterior, runs **after** fitting. It asks whether the fitted model can reproduce the features of the data you actually observed. A model can pass one and fail the other. Sensible-looking simulated data before fitting says nothing about whether the fitted model reproduces the observed patterns; conversely, a wildly diffuse prior can still yield a well-fitting posterior when the data are plentiful — it just wasted the opportunity to encode what you already knew, and may have made fitting harder or slower. ## Practical notes Simulate at the **dataset** level, not just the single-observation level: many problems only show up in the shape of a whole simulated sample, such as an implied range across groups or an implausible spread between the largest and smallest values. Summarise the simulations with a handful of quantities a domain expert can judge — the maximum, the mean, the range, the fraction above a threshold — rather than expecting anyone to eyeball thousands of curves. And treat a failed check as information, not as a verdict on the prior alone: read whether it is the prior, the likelihood family, the link, or the units that produced the absurdity. ## What an interviewer is listening for That you can state the integral, that you understand why the observable scale is where judgment lives, that you know a diffuse prior is not automatically a harmless one, and that you keep the prior check and the posterior check in separate boxes with separate purposes.

  • A colleague argues a flat prior needs no prior predictive check because it assumes nothing. What do you say?
    Flatness on the parameter scale is not neutrality on the data scale. Pushed through a non-linear link or a multiplicative model, a flat prior can put most of its predictive mass on absurd observables. The check is exactly how you discover that, and diffuse priors are among the most common failures it catches.
  • How do you summarise thousands of simulated prior predictive datasets usefully?
    Reduce each simulated dataset to a few quantities a domain expert can judge — the maximum, the range, the mean, the share above a business threshold — and look at the distribution of those across simulations. Comparing summaries beats trying to eyeball raw simulated curves.
  • Can a model pass a prior predictive check and still fit the observed data badly?
    Easily. The prior check only asks whether the assumptions generate plausible data in the abstract; it never looks at the real dataset. Reproducing the actual observed patterns is what the posterior predictive check tests, and the two are independent hurdles.

It is a dress rehearsal with no audience: you run the whole production once to find out that a costume does not fit, rather than discovering it on opening night.

saying these in an interview costs you the question

  • Believes a flat prior is automatically harmless
  • Judges priors only on the parameter scale
  • Confuses the prior check with the posterior check
  • Simulates single values rather than whole datasets

context