What is the difference between MCAR, MAR and MNAR missing data?
answer
- three labels, not three severities
- ask what the missingness depends on
- nothing, observed columns, or the hidden value
- top earners blanking the salary field
basics
~20 sMCAR means missingness is unrelated to any variable. MAR means it depends only on variables you observed. MNAR means it depends on the missing value itself, as when the highest earners are the ones who leave salary blank.
solid answer
~50 sThe three labels describe *why* a value is absent, and they decide which repairs are honest. Missing Completely At Random (MCAR): the chance of being missing is unrelated to everything, observed or not — a wearable that drops readings whenever its battery dies, regardless of who is wearing it. Missing At Random (MAR): the chance depends only on data you have — the same wearable drops readings more often for older participants, and age is recorded, so missingness is random *given age*. Missing Not At Random (MNAR): the chance depends on the unseen value itself — a salary field left blank precisely by the top earners. MCAR is the only one under which dropping incomplete rows leaves an estimate unbiased. MAR is handled by methods that condition on the observed variables. MNAR cannot be repaired from the observed data alone.
go deeper
Be ready to give the three names in full and one concrete example of each, then say which mechanism makes dropping incomplete rows safe. An everyday example is accepted here in place of a formal definition.
An interviewer at this level expects each mechanism phrased as what the probability of being missing depends on, and a clear explanation of why MAR is only random once you condition on the recorded variables.
Show that you check the mechanism before choosing a treatment: compare complete against incomplete rows on the observed fields, say what that comparison can and cannot rule out, and name the assumption your chosen method rests on.
Own the framing that the mechanism is an assumption the team declares rather than a fact the data yields. Decide when an MNAR risk is large enough to block a decision or to justify paying to collect the missing information.
## Three mechanisms, not three severities Missing data is not one problem but three, and the label you attach decides which repairs are legitimate. The taxonomy comes from Rubin's 1976 formulation, which asks a single question: what does the *probability that a value is missing* depend on? Write `R` for the indicator that a cell was recorded, `X_obs` for the values you have, and `X_mis` for the values you do not. ## Missing Completely At Random (MCAR) A value is MCAR when `P(R)` depends on neither `X_obs` nor `X_mis` — not on the value itself, not on any other column, not on anything. A wearable that stops logging heart rate whenever its battery dies, with battery life unrelated to the wearer or to their heart rate, produces MCAR gaps. The consequence is the mildest one available: the recorded rows are a simple random subsample of all rows, so a mean, a proportion or a regression fitted on them targets exactly the quantity it would have targeted on the full data. You lose precision — standard errors grow because `n` is smaller — and you lose nothing else. MCAR is also the rarest of the three in real data, and assuming it because it is convenient is the most common error in this whole area. ## Missing At Random (MAR) MAR means `P(R)` may depend on `X_obs`, but given `X_obs` it does not depend on `X_mis`. Take the same wearable, but now older participants charge it less diligently, so their readings drop out more often. Age is recorded. Within an age band, whether a reading was captured tells you nothing further about what it would have been. The name is unfortunate and trips up most candidates. Missingness under MAR is plainly systematic in the raw table; it is random only *within the strata you can see*. Reading MAR as `missing at random given what we recorded` avoids the confusion. Because the driver is observable, methods that condition on the observed variables can recover unbiased estimates: multiple imputation, likelihood-based analyses, or weighting complete rows by the inverse of their estimated probability of being complete. A plain mean over the complete rows generally cannot, because it silently over-weights the groups that responded. ## Missing Not At Random (MNAR) MNAR means `P(R)` still depends on `X_mis` even after conditioning on everything observed. The canonical case is a salary field left blank precisely by the highest earners: the probability of a blank rises with the very number the blank conceals. Contrast that with the same field being blank because a form section rendered badly for a random slice of users, which would be MCAR. Under MNAR the observed values are a distorted view of the whole by construction. The observed mean salary is too low; imputing from the observed values reproduces the distortion; deleting the blanks entrenches it. No method that uses only the recorded data can fix this, because the information required is exactly what is absent. It is handled by assumption — an explicit model of why values go missing, or a sensitivity analysis over how different the hidden values might be — or by going out and observing the missing information. ## How far the data can take you You can compare complete and incomplete rows on every field you *do* observe. If people missing a salary skew young, or come disproportionately from one region, MCAR is implausible. Little's MCAR test formalises this by checking whether the observed means differ across missingness patterns. Rejection is evidence against MCAR; failing to reject is weak comfort rather than proof, and neither outcome speaks to MNAR at all. That last point is structural, not a limitation of a particular test: the observed data are equally compatible with a MAR story and with an MNAR story, because the values that would separate them were never recorded. ## Why interviewers ask this The taxonomy is the vocabulary for every later decision. Complete-case deletion is defensible under MCAR and dangerous otherwise. Multiple imputation assumes MAR. A missingness indicator is a hedge against informative non-response. A candidate who cannot name the three mechanisms will reach for whichever cleaning step is nearest to hand and will not be able to say what it assumes. ## The practical stance State the mechanism you are assuming, in writing, next to the result. Report the missingness rate per column and per row. Keep a flag marking which values were absent, so the fact of the blank stays available as data. And treat the percentage missing and the mechanism as two separate facts: the first bounds how much precision you lost, the second bounds how wrong you might be.
- Can you test from the data whether missingness is MCAR rather than MAR?Partly. Compare the rows that are complete against the rows missing that field on every variable you do observe; systematic differences make MCAR implausible. That evidence can refute MCAR but can never confirm it, and it says nothing at all about MNAR, because the values that would settle that question were never recorded.
- Why is MAR called 'at random' when the missingness clearly depends on something?It means conditionally random. Given the observed variables, whether a value is missing carries no extra information about what the value would have been. The pattern is systematic in the raw table and random within the strata you can see. Many people read it as 'missing at random given what we recorded', which removes the confusion.
- Does a higher percentage of missing values mean a worse bias?No. The percentage governs how much precision you lose; the mechanism governs bias. Five percent missing under a strong MNAR mechanism can move an estimate further than forty percent missing under MCAR, which only widens the interval. Report both facts, and never treat a small missingness rate as self-clearing.
Think of a clipboard survey losing pages. MCAR is pages blowing off at random; MAR is pages from the windy outdoor booth blowing off more often, and you noted which booth each respondent used; MNAR is respondents tearing out the page whose answer embarrasses them.
saying these in an interview costs you the question
- Treats MCAR, MAR and MNAR as severity levels
- Says MAR means missingness follows no pattern at all
- Assumes a small missing percentage means no bias
- Claims a test on the observed data can confirm MNAR
- Deletes incomplete rows without naming a mechanism