How does multiple imputation produce standard errors that reflect missing-data uncertainty?
answer
- one dataset cannot express doubt
- draws, not conditional means
- average the estimates, then add their spread
- within-imputation plus between-imputation variance
basics
~20 sSeveral completed datasets are built by drawing missing values from a model, each analysed separately, then pooled. The variability of the estimates across datasets is added to the average within-dataset variance, so the standard error grows with how much was missing.
solid answer
~50 sYou build `m` completed datasets, each filling the gaps with a *random draw* from the predictive distribution of the missing values given the observed ones — draws, not conditional means, or the datasets would be near-identical and the whole point is lost. Run the same analysis on each, giving estimates `Q_1 ... Q_m` with variances `U_1 ... U_m`. Rubin's rules then pool them: the point estimate is the average `Qbar`; the within-imputation variance `Ubar` is the average of the `U_i`; the between-imputation variance `B` is the sample variance of the `Q_i`; and the total is `T = Ubar + (1 + 1/m) B`. That `B` term is exactly the uncertainty single imputation throws away, and the `(1 + 1/m)` factor corrects for using finitely many draws. The share of `T` coming from the between term estimates the fraction of missing information, which in turn tells you whether `m` was large enough.
go deeper
Be ready to describe the shape of it: fill the gaps several different ways, run the analysis on each version, then combine — and say that the disagreement between versions is what keeps the uncertainty honest.
You should be able to name the two variance components that get pooled, the within-imputation average and the between-imputation spread, and explain why filling with a predicted mean makes the second one vanish.
Show judgment about the imputation model itself — including the outcome and auxiliary predictors, checking imputed values against observed distributions, and choosing the number of imputations from the fraction of missing information.
Own the call on when this machinery is worth its complexity versus a documented complete-case analysis, and set the expectation that any headline number carries uncertainty inflated for what was missing.
## The three phases Multiple imputation replaces one guess per gap with a distribution of guesses, and then keeps track of how much the answer depends on which guess you drew. **1. Imputation.** Build a model for each incomplete variable given the others, and generate `m` completed datasets by drawing from it. The draw must be stochastic: a predicted mean plus a randomly drawn residual, or a sample from the predictive distribution, and ideally with the imputation model's own parameter uncertainty included. If you fill with the conditional mean instead, all `m` datasets come out nearly identical. **2. Analysis.** Run the identical analysis on each completed dataset, exactly as if it were the real data, producing `m` estimates `Q_i` of the quantity of interest and `m` variance estimates `U_i`. **3. Pooling.** Combine with Rubin's rules. ## Rubin's rules The pooled point estimate is the simple average: `Qbar = (1/m) * sum Q_i` The within-imputation variance is the average of the individual variances — the uncertainty you would have had with no missing data: `Ubar = (1/m) * sum U_i` The between-imputation variance is the sample variance of the estimates themselves — how much the answer moved depending on which plausible completion you drew: `B = sum (Q_i - Qbar)^2 / (m - 1)` And the total variance is `T = Ubar + (1 + 1/m) * B` with the standard error being `sqrt(T)`. The `(1 + 1/m)` inflation accounts for estimating `Qbar` from finitely many imputations rather than infinitely many; it is why small `m` is penalised rather than silently rewarded. Inference then uses a t-reference distribution with degrees of freedom derived from the ratio of the between and total terms, so heavy missingness both widens the interval and thickens its tails. ## Fraction of missing information The quantity `((1 + 1/m) * B) / T` estimates the fraction of missing information (FMI) — the share of the total uncertainty attributable to the gaps rather than to ordinary sampling noise. It is not the same as the fraction of cells missing: a well-predicted variable with 30% of its values absent can have a small FMI, while a poorly-predicted one with 10% missing can have a large one. FMI is the number to quote when someone asks how much the missing data cost you. ## Choosing m Five to ten completions was the classic advice and is adequate when little information is missing. Modern practice scales `m` with FMI — a common rule of thumb sets the number of imputations near a hundred times the fraction of missing information, so an FMI of 0.3 suggests around 30. More imputations cost only computation and make the pooled standard error more stable, so err upward. ## Building the imputation model The imputation model must be at least as rich as the analysis model. Concretely: - **Include the outcome.** Imputing predictors without the outcome draws values as if the two were unrelated, which attenuates the very association you are trying to estimate. - **Include auxiliary variables** that predict either the missing value or the fact that it is missing, even if the analysis will not use them. They make the MAR assumption more plausible and reduce FMI. - **Preserve structure.** Interactions, non-linear terms and skew present in the analysis model should be representable by the imputation model, or the completions will be systematically too tidy. - **Respect the variable type.** A binary field should get plausible binary values, not fractional ones. This requirement is called congeniality: an imputation model less rich than the analysis model biases results toward the simpler structure it assumed. ## Checking imputations Compare the distribution of imputed values against the observed values, per variable and per completion. They need not match — under MAR they legitimately differ, since the missing rows differ on the observed variables — but wildly implausible values, impossible categories or negative quantities point at a broken imputation model. Also confirm the `m` estimates genuinely differ; if `B` is near zero when a lot of data is missing, the draws are probably not stochastic. ## What it does not do Standard multiple imputation assumes missing at random: that the missing values behave like the observed ones once you condition on the variables in the imputation model. It is not a cure for MNAR. It also cannot invent information that was never collected — it is machinery for propagating honest uncertainty, not for manufacturing certainty. ## The common failure The most frequent mistake is to average the `m` completed datasets into one filled dataset and then analyse that. This collapses to a single, unusually smooth imputation, destroys the between-imputation variance and reproduces exactly the falsely narrow standard error that multiple imputation exists to prevent. You pool the *estimates*, never the datasets.
- Why must the imputations be random draws rather than predicted means?Predicted means give near-identical completed datasets, so the between-imputation variance collapses toward zero and the pooled standard error reverts to the falsely narrow single-imputation one. The added random component is what lets the completions disagree, and that disagreement is the honest measure of how much the missing values could have changed the answer.
- How many imputations are enough?Five to ten is the classic answer and is fine when little information is missing, but the requirement scales with the fraction of missing information; a common rule sets the number near a hundred times that fraction. Extra imputations cost only compute and stabilise the pooled standard error, so err upward.
- Does multiple imputation fix MNAR data?No. It assumes the missing values behave like the observed ones once you condition on the variables in the imputation model, which is the missing-at-random assumption. Under MNAR you need an explicit model of why values go missing, or a sensitivity analysis that shifts the imputed values and reports where the conclusion flips.
- Which variables belong in the imputation model?At minimum every variable in the analysis model, including the outcome, plus auxiliary variables that predict either the missing value or the fact of it. Leaving the outcome out draws imputations as though no relationship existed, which attenuates the very association you are trying to estimate.
Ask five careful colleagues to reconstruct the pages missing from a report, then have each of them draw a conclusion from their own reconstruction. How far their conclusions differ is your honest measure of how much those missing pages mattered.
saying these in an interview costs you the question
- Imputes with predicted means and calls it multiple imputation
- Pools estimates but ignores the between-imputation variance
- Omits the outcome from the imputation model
- Claims multiple imputation handles MNAR data
- Averages the completed datasets into one before analysing