skip to content

How do you choose which test statistics to compare in a posterior predictive check?

level: principalimportance: nice to knowfreq 20%

answer

  1. what the decision touches
  2. avoid statistics the fit already matched
  3. tails, dispersion, dependence
  4. data used twice, so conservative
  5. hold out or cross-validate for calibration

basics

~20 s

Choose statistics that matter for the decision the model supports and that the fitting did not already force to match. A statistic the fit targets, such as the sample mean, passes regardless of how wrong the model is.

solid answer

~50 s

Two criteria dominate. First, **decision relevance**: check the features whose failure would change an action — the maximum if you size capacity to peaks, the proportion of zeros if zeros trigger an alert, the number of runs if you care about temporal dependence. Second, **diagnostic power**: prefer statistics the fitting procedure did not already tune. A model fitted to match the mean will reproduce the observed mean almost by construction, so that check passes even for a badly wrong model; tail, dispersion and dependence statistics carry far more information. Report a small portfolio rather than one number, and remember that posterior predictive p-values are conservative — the observed data influenced the posterior that generated the replicates, so their distribution is not uniform but bunched toward 0.5. Where calibration matters, use held-out or cross-validated predictive checks instead.

go deeper

for a junior

Know that a posterior predictive check needs you to pick a summary quantity to compare, and that the choice decides what the check can detect.

for a middle

Be able to name concrete statistics for tails, dispersion, shape and dependence, and explain why a statistic the fit already matched is a weak check.

for a senior

Tie the statistics you check to the decisions the model drives, and be ready to justify why a documented misfit on one feature is or is not acceptable.

for a principal

Own the calibration caveat that posterior predictive p-values are conservative because the data are used twice, and set an organisation-wide standard set of checks with a held-out or cross-validated option where calibration matters.

## Why the choice is the whole game A posterior predictive check compares observed data with replicate datasets simulated from the fitted model. But raw datasets cannot be compared directly — you must reduce both to a **test statistic** `T`. That choice determines what the check can see. A badly chosen statistic makes a broken model look fine; a well-chosen one makes the specific failure obvious and points at the fix. Since the check is only as good as the statistic, the selection is a judgment call a lead is expected to own rather than a mechanical step. ## Criterion one: decision relevance Start from what the model is for. A model that is wrong in a direction nobody acts on is a different situation from one that is wrong exactly where a threshold sits. - Sizing infrastructure to peak load? Check the **maximum** and upper quantiles. - Triggering an alert when a day records nothing? Check the **proportion of zeros**. - Forecasting a sequence where streaks matter? Check the **number of runs** or the lag-1 autocorrelation, which expose dependence an independence assumption denies. - Reporting a per-segment breakdown? Check the **spread across segments**, since a model can match the pooled data while getting every group wrong. This framing also gives you a principled stopping rule: you check what the decision touches, not every statistic you can compute. ## Criterion two: diagnostic power Some statistics are structurally unable to fail. If the estimation procedure matches the sample mean, replicates will reproduce the observed mean whatever else is wrong, and a check on the mean is close to vacuous. The informative statistics are the ones the fit did **not** target: - **Tail behaviour** — the maximum, the minimum, the 99th percentile. - **Dispersion** — the variance or the ratio of variance to mean, when the assumed family ties them together. - **Shape** — skewness, the number of modes, the count of exact zeros. - **Dependence** — runs, autocorrelation, within-group similarity, all invisible to a model that assumes independence. The general principle: pick statistics that are **not** simple functions of the fitted quantities. That is what turns the check from a tautology into evidence. ## The conservatism of posterior predictive p-values The honest caveat, and the part that separates a principal-level answer. The replicates are generated from a posterior that was itself computed from the observed data — the data are used twice. The consequence is that a posterior predictive p-value does not behave like a frequentist p-value under a correct model: instead of being uniform on the unit interval, its distribution is concentrated toward 0.5. The check is therefore **conservative** — it under-detects misfit, and an unremarkable value is weaker evidence of a good model than it looks. This is a known property, not an implementation flaw. Two responses. Where you only need direction and diagnosis, accept it: an extreme value is still strong evidence of misfit precisely because the procedure is reluctant to produce one. Where calibration matters, move to a check that does not reuse the data — hold out part of the dataset and check the predictive distribution against it, or use a cross-validated predictive check that predicts each observation from a posterior that excluded it. ## Building a portfolio One statistic is a single question. A small, deliberately chosen set — say three to six, spanning tail, dispersion, shape and dependence — makes the check a profile. Reporting a table of observed values against replicate distributions is far more useful to a reviewer than a single number, and it makes the tradeoffs visible when fixing one failure worsens another. The counter-pressure is multiplicity: check enough statistics and one will look extreme by chance. This is why the checks should be **pre-declared** and tied to decisions, not mined after the fact. A statistic chosen because it happened to look bad is a story, not a diagnosis. ## Organisational angle At scale the durable move is a standard set of checks every model of a given kind ships with, plus a required, documented model-specific addition. That gives comparability across models and reviewers, prevents the check from degenerating into whatever the author found flattering, and makes 'this model has a known, accepted misfit in the tail' a statement a team can reason about rather than a surprise discovered in production.

  • Why are posterior predictive p-values described as conservative?
    Because the replicates come from a posterior fitted to the same data being checked, so the model has partly been tuned to the observation it is judged against. Under a correct model the p-value distribution is bunched toward 0.5 rather than uniform, meaning the check under-detects misfit and an extreme value is correspondingly strong evidence.
  • Why is checking the sample mean often a weak choice of test statistic?
    If the fitting procedure matched the mean, replicates reproduce it almost by construction, so the check passes for models that are badly wrong in the tails, in dispersion or in dependence. Informative statistics are the ones the fit did not target.
  • How do you avoid fooling yourself when checking many test statistics?
    Pre-declare the set and tie each to a decision the model supports, rather than scanning statistics and reporting the one that looks interesting. With enough statistics some will look extreme by chance, so a discovered anomaly should be treated as a hypothesis to investigate, not a finding.
  • What alternative gives a better-calibrated check than the standard posterior predictive one?
    A predictive check that does not reuse the data: hold out a portion of the dataset and compare the predictive distribution against it, or use cross-validated predictive checks where each observation is predicted from a posterior fitted without it. Both cost more compute and buy honest calibration.

saying these in an interview costs you the question

  • Checks only the mean and declares the model fine
  • Treats a posterior predictive p-value as calibrated and uniform
  • Picks the statistic after seeing which one looks bad
  • Believes more statistics always make a stronger check
  • Cannot name a statistic sensitive to dependence

context