skip to content

A key field may be MNAR and the data cannot prove otherwise — how do you proceed?

level: principalimportance: nice to knowfreq 28%

answer

  1. no test can settle it
  2. an assumption to declare, not to prove
  3. vary the assumption, report the tipping point
  4. make the hidden driver observable

basics

~20 s

Treat MNAR as an assumption to declare rather than a hypothesis to test. Report the missingness rate, run a sensitivity analysis that shifts the imputed values until the conclusion flips, and where the stakes justify it, go and observe the non-responders.

solid answer

~50 s

No test on the observed data can separate MAR from MNAR, because the evidence that would settle it is exactly what is missing. So the honest move is to make the assumption explicit and bound its consequences. A pattern-mixture sensitivity analysis does that: impute under missing-at-random, then shift the imputed values by a delta representing how different you believe non-responders to be, and re-run across a range of deltas. Report the tipping point — the delta at which the decision reverses — and let stakeholders judge whether that scenario is plausible. In parallel, attack the cause: chase a subsample of non-responders, link an administrative source, or change the instrument so the field is less avoidable, such as offering salary bands rather than an exact figure. Each of those turns an unobserved driver into a recorded one, which is what converts an MNAR problem into a tractable one.

go deeper

for a junior

Be ready to say that you cannot tell MNAR from MAR using the data alone, and that the right first move is to flag the missingness rate and the risk rather than quietly picking a filling rule.

for a middle

Explain that the usual imputation methods assume missing-at-random, so an MNAR worry is handled by varying that assumption and watching how far the answer moves, not by switching to a different default method.

for a senior

Show that you would rerun the analysis across a range of assumed departures, report where the conclusion flips, and pair that with a concrete plan to observe the driver — a follow-up of non-responders or a linked source.

for a principal

Own the decision rule and how it is communicated: what departure counts as plausible, who signs off when the tipping point is close, whether to fund recovering the data, and how the assumption appears in the written result.

## The identification problem MNAR means the probability that a value is missing depends on that value even after conditioning on everything observed. The uncomfortable consequence is that MNAR is not testable. Two models — one in which non-responders resemble comparable responders, and one in which they are systematically different — can produce identical distributions for the data you actually hold. They differ only in how they extrapolate into the region you never see. No amount of data, and no cleverer test, breaks that tie. This is an identification problem, not a sample-size problem, and treating it as the latter is the most common senior-level mistake. What follows is not despair but a change of job. The task stops being 'determine the mechanism' and becomes 'state the assumption, bound its consequences, and reduce the region you have to assume about'. ## Two ways to write the assumption down **Selection models** specify the probability of being missing as a function that includes the missing value itself, and fit that jointly with the outcome model. They are elegant and highly sensitive to the functional form you chose, which is rarely something you can defend. **Pattern-mixture models** instead specify how respondents and non-responders differ, then mix the two populations. They are usually the more practical framing in an interview and in practice, because the assumption is stated in the units of the thing itself — 'non-responders earn on average X more than comparable responders' — which non-statisticians can argue with productively. ## Delta adjustment and the tipping point The workhorse technique is a delta-adjusted analysis. Impute the missing values under the missing-at-random assumption, then add a shift `delta` to every imputed value, and re-run the whole analysis for a grid of deltas from zero out to implausibly large. Two outputs matter: - The **trajectory** of the estimate: how fast the conclusion degrades as the assumption is stretched. - The **tipping point**: the smallest `delta` at which the decision reverses. The tipping point is what you carry to a decision maker, because it converts an untestable statistical assumption into a domain judgment they are qualified to make. 'Non-responders would have to earn thirty percent more than otherwise-identical responders before this result reverses' is a sentence a business leader can evaluate. A single point estimate computed under an unstated assumption is not. ## Shrink the unobserved region Sensitivity analysis bounds the damage; it does not remove it. The higher-value moves make the driver observable: - **Follow up a subsample of non-responders.** A small, hard-won sample of the people who left the field blank gives you real values from the region you were assuming about, and lets you calibrate the delta rather than guess it. - **Link an external record.** An administrative or already-collected internal source that carries the same field for some rows converts unobserved values into observed ones for a subset. - **Collect auxiliary predictors.** Any variable that predicts *both* the value and the propensity to withhold it moves the situation closer to MAR, because conditioning on it absorbs part of the dependence. - **Redesign the instrument.** Bands instead of exact figures, an explicit 'prefer not to say' option that separates refusal from oversight, a less exposed placement of the question. Reducing the cost of answering directly reduces the informativeness of the blank. ## Always carry the fact of the blank Whether or not the value is recoverable, whether it was missing is itself data. Keep an indicator, look at whether it predicts your outcome, and report that relationship. Informative non-response that shows up as a strong indicator effect is one of the few visible signatures of an MNAR problem, and it is free to look for. ## The organisational call At a lead level the statistics is the easy half. The decisions you own are: - **What departure counts as plausible**, agreed with domain experts *before* the tipping point is computed, so the threshold is not chosen to suit the result. - **Whether a headline number ships at all.** If a mild and believable departure reverses the decision, publishing a single number implies a confidence the analysis does not have. Publish the range and the assumption, or delay. - **Whether to fund recovery.** Chasing non-responders costs money; compare it against the cost of the decision being wrong, and make that comparison explicit rather than defaulting to whichever is cheaper this quarter. - **How the assumption is recorded.** The missingness rate, the assumed mechanism and the sensitivity result belong in the write-up as first-class content, not a footnote, so the next person to reuse the number inherits the caveat with it. ## The one-line summary You cannot test your way out of MNAR. You can state the assumption, price it, and buy information that shrinks it.

  • What does a tipping-point analysis actually report to a decision maker?
    The size of departure from the missing-at-random assumption that would overturn the conclusion, expressed in the units of the field. Saying that non-responders would have to earn thirty percent more than comparable responders before the result reverses is something a business reader can judge, unlike a number computed under one silent assumption.
  • How does chasing a subsample of non-responders help?
    It produces observed values from the exact region you were assuming about, so the driver of the missingness becomes measurable rather than hypothetical. Even a small, expensive subsample lets you estimate how far non-responders really differ and calibrate the shift used in the sensitivity analysis instead of guessing it.
  • When would you refuse to publish a single headline number at all?
    When the tipping point sits well inside the plausible range — a mild, believable departure from the assumption reverses the decision. Then publish the range with the assumption attached, or delay until the missing information is recovered. A point estimate there implies confidence the analysis has not earned.

saying these in an interview costs you the question

  • Claims a statistical test can confirm MNAR
  • Reports one number without stating the missingness assumption
  • Treats multiple imputation as a cure for MNAR
  • Never considers observing the non-responders directly
  • Presents the missing-at-random result as conservative by default

context