skip to content

Complete-case analysis drops 40% of your rows — when does that bias the estimate?

level: seniorimportance: should knowfreq 50%

answer

  1. count is precision, mechanism is bias
  2. who survived, not how many
  3. a random subsample keeps the target
  4. predict completeness from what you observed

basics

~20 s

The share dropped controls precision; the mechanism controls bias. If values are missing completely at random, the surviving 60% are a random subsample and the estimate stays unbiased, only noisier. Under MAR or MNAR it can be badly off.

solid answer

~50 s

Deleting incomplete rows is unbiased only when the survivors are a random subsample — that is, under MCAR. You pay for it in precision, with the standard error inflated by roughly `sqrt(n / m)` where `m` rows survive out of `n`, but you are still estimating the right quantity. Under MNAR the survivors differ systematically: if the top earners are the ones leaving salary blank, the complete-case mean salary is too low no matter how many rows remain. Under MAR a plain marginal mean is generally biased too, though there is an important exception in regression — if the probability that a row is complete depends only on the covariates in the model and not on the outcome given them, the fitted coefficients stay consistent. So the first question is never how many rows you lost, but what predicts being complete. Compare kept rows against dropped rows on every observed field before deciding.

go deeper

for a junior

Be ready to say that deleting incomplete rows is only safe when values are missing completely at random, and that a large drop is a warning worth raising rather than a routine cleanup step.

for a middle

Explain the two separate effects: fewer rows widen the standard error by roughly the square root of the ratio of full to complete counts, while a non-random pattern of survival shifts the estimate itself.

for a senior

Show the diagnostic habit — flag complete against incomplete rows, compare them on every observed field, and state which mechanism your chosen treatment assumes. Interviewers here want the investigation, not a rule of thumb.

for a principal

Own when the loss is unacceptable regardless of the statistics: set the threshold at which an analysis must be rerun with imputation, escalated, or blocked, and decide whether to fund recovering the missing information at source.

## What complete-case analysis is Complete-case analysis, also called listwise deletion, keeps only rows with no missing value in any variable the analysis touches, and discards the rest. It is the default in most tooling and the most common thing done to missing data without anyone deciding to do it. ## Losing 40% is easier than it sounds The loss compounds across columns, because a row survives only if it clears every one. Ten independent columns each 5% missing leave about `0.95^10 ≈ 0.60` of rows complete. No single column looked alarming; 40% of the data is gone. This is why the first diagnostic is a missingness rate per column *and* the complete-row rate, not one or the other. ## The two effects, kept separate **Precision.** Going from `n` rows to `m` complete ones inflates the standard error of a mean by about `sqrt(n / m)`. With `n = 100` and `m = 60` that is roughly 29% wider intervals. This effect is real, predictable, and honest — the interval widens to reflect that you have less data. **Bias.** Whether the estimate itself moves depends entirely on the mechanism. - **MCAR.** The complete rows are a simple random subsample of all rows. Every estimate keeps its target; only precision suffers. Deletion is defensible, and the only thing to report is the loss of power. - **MAR.** Missingness depends on observed variables, so the complete rows over-represent some groups. A marginal mean computed on them is generally biased: if under-35s answer a spending question far more often, the complete-case mean spend is really a mean weighted toward the young. - **MNAR.** The complete rows differ on the very quantity being estimated. Salaries left blank by high earners produce a complete-case mean that is too low, and no amount of remaining data corrects it. ## The regression exception worth knowing Candidates who only know the rule 'deletion is safe under MCAR' miss a result that comes up constantly in practice. In a regression of an outcome `Y` on covariates `X`, complete-case analysis remains consistent whenever the probability that a row is complete depends only on `X` and not on `Y` given `X` — even if that dependence on `X` is strong, and even if the covariate missingness would be called MNAR in the covariate. The reason is that conditioning on `X` is exactly what the regression already does, so restricting to a subset selected on `X` leaves the conditional distribution of `Y` given `X` intact. The converse is the danger: selection driven by the outcome, or by unmeasured causes of the outcome, bends the fit. This is why 'what predicts being complete' is the question to ask, and why the answer must be checked against the model you intend to fit rather than against a generic rule. ## Diagnostics before deletion 1. Build a complete-versus-incomplete flag for the rows. 2. Compare the two groups on every observed variable: means, distributions, category shares. 3. Model that flag from the observed fields and see which ones predict completeness, and how strongly. 4. Ask whether any predictor of completeness is also a driver of your outcome. If yes, deletion is not neutral. Systematic differences refute MCAR. Finding *no* differences is weak comfort, since the variable that decides completeness may itself be one you never recorded. ## Available-case analysis is not a free upgrade Pairwise or available-case deletion — computing each statistic from whichever rows have the variables it needs — retains more data but introduces its own problems: different cells of a covariance matrix rest on different row sets and different sample sizes, and the assembled matrix can fail to be a valid covariance matrix at all. It also makes the effective `n` of a reported result ambiguous. ## Alternatives when deletion is not defensible Keep the rows and carry the uncertainty: add an indicator marking that the value was absent, so informative non-response can be detected rather than discarded, and build several stochastic completions to be analysed and pooled so the standard errors reflect what was missing. Neither recovers information that was never recorded; both stop you from throwing away the information you do have and from overstating your confidence in what remains. ## Reporting Whatever you choose, report three things: the fraction of rows dropped, how kept and dropped rows differ on observed variables, and the mechanism you are assuming. An analysis that silently loses 40% of its rows is not a cleaned analysis; it is an unstated subgroup analysis.

  • How would you check, from the data you have, whether the dropped rows differ?
    Create a complete-versus-incomplete flag and compare the two groups on every observed variable — means, distributions, category shares. You can also model that flag from the observed fields and see what predicts it. Systematic differences refute MCAR; finding none is weak comfort, since the deciding variable may itself be unobserved.
  • Is a missingness indicator plus a filled value a legitimate alternative to deleting the row?
    It keeps the sample size and lets the analysis learn whether the blank itself carries signal, which matters when non-response is informative. It does not recover the hidden value, and under MNAR the estimate can still be biased. Treat the indicator as a diagnostic and a hedge, never as a repair.
  • Why can dropping rows with any missing field lose far more than any single column's rate?
    Losses compound, because a row survives only if it clears every column. Ten independent columns each 5% missing leave about 0.95 to the tenth power, roughly 60%, of rows complete. That is how a dataset with no alarming per-column rate ends up discarding 40% of its rows.

saying these in an interview costs you the question

  • Judges bias by the percentage of rows dropped
  • Assumes deletion is safe because the sample is large
  • Never compares dropped rows against kept rows
  • Calls complete-case analysis unbiased under any mechanism
  • Uses pairwise deletion without noticing inconsistent sample sizes

context