skip to content

Imputation Strategies

Drop rows or columns, or fill with the mean, median, mode, a kNN estimate or iterative regression, keeping a missing-indicator flag. Interviewers check that the imputer is fitted inside the fold.

on this pageshow

questions

4

20% of survey respondents skipped the income question — do you drop those rows, drop the column, or impute?

level: juniorimportance: must knowfreq 82%

answer

  1. think about what happens at scoring time
  2. you cannot delete a live request
  3. skewed money column: mean or median?
  4. categoricals get a level, not the mode
  5. freeze the fill value from training

basics

~20 s

Impute in most cases. Deleting the rows discards a fifth of the data and has no equivalent at scoring time, where every request must return a prediction. Drop the column only if it is nearly empty or adds no lift.

solid answer

~50 s

I start from what happens in production: a live request with no income still needs a score, so row deletion is a training-time convenience, not a strategy the pipeline can carry to serving. At 20% missing I would keep the column and fill it. For a skewed money quantity the median is the safer constant than the mean, because a handful of very high earners pull the mean to a value almost no respondent actually sits near. If the field were categorical I would add an explicit `Missing` level rather than fill with the most common answer, since the mode invents a response the person never gave. I would only drop the column if it is nearly empty or shows no measurable lift on a validation split. Whichever fill I use, I keep a binary flag recording that the value was absent.

go deeper

for a junior

Know the three options by name — drop rows, drop the column, fill — and be ready to say why filling is the usual answer. Remember median over mean for skewed numbers and an explicit Missing level over the mode for categories.

for a middle

Explain why the fill statistic must be computed on the training data and then frozen, why the median resists skew where the mean does not, and what a constant fill does to the column's spread and to its relationship with the target.

for a senior

Show that you reason from the serving path backwards: row deletion has no inference-time equivalent, so training on complete cases alone leaves the model untested on exactly the rows it will meet. Justify keep-or-drop with a validation comparison, not instinct.

for a principal

Own the standard: which columns are allowed a fill at all, when the honest answer is to fix the collection instrument rather than paper over a fifth of the responses, and what a team is permitted to ship without a documented missingness policy.

## The three options, stated precisely When a feature column has gaps, there are exactly three things you can do with them, and they operate on different objects: 1. **Complete-case deletion (drop rows)** — remove every row that has a gap in any column you intend to use. You lose rows. 2. **Drop the column** — remove the feature entirely for every row. You lose a feature. 3. **Impute (fill)** — substitute a value for each gap and keep both the rows and the column. You lose the distinction between an observed value and a manufactured one, unless you record it separately. At 20% missing in a survey income field, all three are on the table, and the decision is not a matter of taste. ## The constraint that settles it: scoring time A predictive model exists to score rows it has never seen. When a request arrives with no income, you cannot answer "no prediction, that row was deleted" — the caller needs a number. So **row deletion is not a transformation your pipeline can apply at inference**; it only ever applies to the training table. That asymmetry has a practical consequence. If you train only on complete cases, you never taught the model what to do with an incomplete one, yet at serving you will hand it incomplete ones anyway, filled by whatever your feature code happens to do. Training on the same representation you will serve — filled value plus a missingness flag — keeps the two paths consistent. Deleting rows is defensible in a narrow case: the missing fraction is genuinely small (a fraction of a percent), you have plenty of data, and you still define a serving-time fill for the rows that will inevitably show up incomplete later. Twenty percent is not that case. ## Choosing the fill value **Numeric columns.** The two standard constants are the mean and the median of the observed values. The median is more robust to skew and to extreme values, which is why it is the default choice for money, counts, durations, and anything with a long right tail. Income is the textbook skewed quantity: the arithmetic mean of a population's incomes sits well above the typical income, so mean-filling places every non-respondent at a level most respondents never reach. Compute whichever statistic you choose on the training data only, and keep the number — it becomes a fixed parameter of the model that gets applied unchanged to future rows. **Categorical columns.** The reflex is to fill with the mode, the most frequent level. That is usually the wrong instinct: it silently asserts that every non-responder gave the most common answer, which inflates that level and blurs whatever distinction the column encoded. The better default is to make absence its own level — a category literally named `Missing`. This is honest, it costs nothing, and it lets any model that handles categories learn whatever pattern non-response carries. **Never fill with zero by default.** Zero is a real, meaningful value in most numeric columns; using it as a placeholder mixes "no income reported" with "income of nothing", and shifts every summary of the column. ## When to drop the column instead Dropping the column is the right call when the feature is so sparse that whatever you fill dominates it — a column that is 95% empty is, after filling, 95% one constant, which is nearly a constant column and carries almost no usable signal in its values. It is also right when the feature simply does not earn its place: fill it, train, compare validation score with and without, and if the difference is inside the noise, the simpler pipeline with one fewer column to produce, monitor and serve wins. The important caveat is that a mostly-empty column may still be worth keeping in a reduced form: throw away the values and keep only a binary present/absent flag. That costs one column, requires no fill value, and captures whatever the act of answering encoded. ## What imputation costs you Filling is not free. Any single constant pretends you know something you do not: after the fill, the model cannot tell an imputed entry from a genuine one, and 20% of the column now sits stacked on one value, which weakens whatever relationship that feature had with the target. You mitigate this by keeping a binary indicator of which entries were filled, so the model can treat the two groups differently rather than being lied to. ## The decision in order Check how much is missing and whether the feature matters at all; decide whether the column survives; pick a fill that respects the column's type and shape (median for skewed numerics, an explicit level for categoricals); compute the fill on training data and freeze it; add a missingness flag; and confirm with a validation comparison rather than an opinion.

  • When is deleting the incomplete rows genuinely the right call?
    When the missing fraction is tiny — well under a percent — you have abundant data, and dropping them cannot plausibly reshape who is left in the training set. Even then it is only a training-table decision: you still have to define what the serving path fills in, because incomplete requests will arrive regardless.
  • The income column turns out to be 95% empty — is anything worth keeping?
    Usually yes, but not the values. After filling, 95% of the column is a single constant, so its numeric content is near-worthless. Keep it instead as a single binary present/absent flag: that needs no fill value, costs one column, and preserves whatever the act of answering signals about the respondent.
  • Why not just fill the numeric gaps with zero and move on?
    Because zero is a legitimate value in most numeric columns, so the fill becomes indistinguishable from a real reading of nothing. It also drags the column's centre downward and can invert the sign of a learned relationship. If you want a value that means absent, encode absence explicitly with a flag rather than smuggling it into the number.

Deleting incomplete rows is like a shop that only serves customers who filled in every field on the form. It works until the customers are standing at the counter, and then you still have to serve them.

saying these in an interview costs you the question

  • Says always fill numeric columns with the mean
  • Deletes a fifth of the rows without asking who remains
  • Assumes a row can simply be skipped at scoring time
  • Fills a categorical with the mode and calls it neutral
  • Uses zero as a universal placeholder for missing numbers
  • Drops any column containing missing values on sight

context

open as a page

Why add a binary missing-indicator column next to a feature you have imputed?

level: middleimportance: should knowfreq 58%

basics

~20 s

Because filling destroys the fact that the value was absent, and absence is often predictive in its own right. A binary flag restores that information, and it lets the model treat filled rows differently from genuinely observed ones.

open as a page

A scoring request arrives with a null device-age field — what value do you fill and where does it come from?

level: seniorimportance: should knowfreq 46%

basics

~20 s

The fill comes from a statistic computed once on the training data and shipped with the model as a fixed parameter. Never recompute it from live traffic: the same request would then score differently depending on its neighbours.

open as a page

When does kNN or iterative imputation beat filling a column with its median?

level: middleimportance: nice to knowfreq 38%

basics

~20 s

When the incomplete column is strongly predictable from the other columns. Both borrow that structure — kNN from similar rows, iterative imputation from a regression on the other features — where a median gives every gap the same answer.

open as a page