When does kNN or iterative imputation beat filling a column with its median?
answer
- one number ignores the rest of the row
- borrow from similar rows instead
- the incomplete column becomes the target
- cycle column by column until stable
- compare downstream score, not intuition
basics
~20 sWhen the incomplete column is strongly predictable from the other columns. Both borrow that structure — kNN from similar rows, iterative imputation from a regression on the other features — where a median gives every gap the same answer.
solid answer
~50 sA median fill uses one number for every gap, so it throws away everything the rest of the row says. Take property listings with a missing floor area: bedroom count, price band and postcode predict floor area well, and a market-wide median will be badly wrong for both a studio and a six-bedroom house. kNN imputation fills the gap from the k most similar listings measured over the observed features, typically their average. Iterative imputation instead treats the incomplete column as a target, regresses it on the other columns, predicts the missing entries, and cycles through the incomplete columns repeatedly until the filled values stop changing. Both pay off when the correlation is strong and much is missing. They cost more: kNN needs reference rows wherever you score, iterative needs a fitted chain of models, and the gain over median plus a missing-indicator is often smaller than expected — so measure it first.
go deeper
Know that beyond mean and median fills there are methods that use the rest of the row: kNN borrows from similar records, and iterative imputation predicts the gap from the other columns. Be able to say why that can beat one constant.
Describe both procedures step by step, including kNN's scaling sensitivity and its k tradeoff, and iterative imputation's round-robin passes until values stabilise. Explain why the target must never be a predictor in the imputation model.
Show that you weigh the cost, not just the sophistication: reference rows or a model chain to deploy, version and monitor, against a gain that is often inside the noise of a median-plus-indicator baseline. Insist on a masked-value reconstruction test and a downstream comparison.
Own whether the platform supports conditional imputation at all. A kNN fill in the serving path means training records inside the deployed artifact, and a model chain means more things that can drift — decide when that complexity is worth its permanent maintenance cost.
## The limitation the fancy methods address A constant fill answers every gap identically. If floor area is missing for a listing, a median fill says "assume the typical property" whether the row describes a one-room studio in a dense city centre or a six-bedroom house with four bathrooms. Yet the row is telling you plenty: bedrooms, bathrooms, price band, postcode, property type. All of that is discarded. Both kNN and iterative imputation exist to use it. They are called **multivariate** or **conditional** imputation, because the value they insert depends on the rest of the row rather than only on the column. ## kNN imputation For a row with a gap in one column, find the k rows most similar to it, measured over the features that *are* observed in both. Fill the gap with an aggregate of those neighbours' values in that column — usually the mean for a numeric column, or the most common level for a categorical one, sometimes weighted by distance so nearer neighbours count for more. Properties worth knowing: - It is **non-parametric**: it assumes no functional form, so it can capture whatever local structure exists. - It is **sensitive to feature scaling**, because similarity is a distance and a feature with a large numeric range will dominate it unless the features are put on comparable scales first. - It has a real **k tradeoff**: small k tracks local structure but is noisy; large k pulls every fill toward the global average, which in the limit collapses back to the constant fill you were trying to beat. - It is **memory-resident by nature**: to fill a gap you need the reference rows. That is fine offline and expensive online. ## Iterative imputation Also described as chained-regression imputation. The procedure: 1. Initialise every gap with a simple fill — the column median, say — so that a complete table exists. 2. Pick one incomplete column. Treat it as the target and the remaining columns as predictors. Fit a regressor (for a numeric column) or a classifier (for a categorical one) on the rows where that column was actually observed. 3. Use that fitted model to predict the entries that were originally missing in that column, and overwrite their current fills. 4. Move to the next incomplete column and repeat, cycling round-robin through all of them. 5. Run several passes until the filled values stop changing meaningfully, then stop and keep the single filled table. Properties worth knowing: - It models each column **conditionally on all the others**, so it captures relationships a per-column constant cannot. - It handles **several incomplete columns at once**, which kNN also does but less explicitly. - It is **iterative and can fail to settle**: with strongly collinear columns the values can oscillate, and you cap the number of passes. - The choice of the inner model matters. A linear regressor imposes linear relationships on the filled values; a tree-based one captures interactions but can overfit small observed subsets. ## When the extra machinery is worth it Three conditions, roughly, and you want all three: 1. **The incomplete column is genuinely predictable from the others.** You can check this directly: on the rows where the value *is* observed, fit the imputation model and see how well it reconstructs held-out true values. Weak reconstruction means there is nothing to borrow, and the sophistication buys nothing. 2. **A meaningful share of the column is missing.** Filling 0.5% of rows more cleverly cannot move a validation score. 3. **The downstream model cannot route around missingness itself** and the column matters to it. ## When to stay with the simple fill **Serving cost.** kNN imputation means shipping reference rows with the model and searching them per request; iterative imputation means deploying and versioning a chain of extra models. Both turn a constant lookup into infrastructure. **The gain is often inside the noise.** In practice, a median fill plus a binary missing-indicator is a strong baseline, and the improvement from conditional imputation is frequently small enough that a careful validation comparison cannot distinguish them. Run that comparison before you commit. **Overconfidence.** Both methods produce a single confident-looking number where the truth is unknown, and the downstream model then treats it as observed. The filled values are smoother and less variable than real data, so they can make relationships look cleaner than they are. ## One rule that is not optional Do not use the target variable as a predictor inside the imputation model. It is tempting — the target is usually the most informative column you have — and it produces filled features that encode the answer, which flatters validation and collapses in production where the target is unknown by definition. The imputation model gets the features, not the label. ## How to decide, concretely Mask a random sample of *observed* values, impute them with each candidate strategy, and measure the reconstruction error against the values you hid. That ranks the imputers on their own terms. Then, separately, train the downstream model on each imputed dataset and compare validation scores — because reconstruction accuracy is a means, and the model's performance is the thing you actually care about. The two rankings do not always agree, and when they disagree the downstream score wins.
- Why is feature scaling important for kNN imputation?Because neighbours are chosen by distance, and a feature measured in large units contributes far more to that distance than one measured in small units. Without putting the features on comparable scales, the neighbour set is effectively decided by whichever column happens to have the widest numeric range, and the borrowed value reflects similarity on that column alone.
- What happens if you set k too large in a kNN fill?Each gap is filled from an ever-wider set of rows, so the filled value drifts toward the column's global average. At the extreme it is the constant fill you were trying to improve on, with all the computational cost and none of the benefit. Small k is noisier but local; the value of k is a tradeoff you tune, not a default you accept.
- Can you include the target variable among the predictors when imputing a feature?No. It would produce filled features that carry information about the label, which inflates validation performance and cannot be reproduced at prediction time, when the label is unknown by definition. The imputation model sees the other features only.
- How do you actually measure whether a fancier imputer helped?Two ways, used together. Mask a sample of values you can see, impute them, and score the reconstruction against the truth you hid. Then train the downstream model on each imputed dataset and compare validation performance. Reconstruction quality is diagnostic; the downstream score is what decides, and when they disagree the downstream score wins.
Valuing a house with no floor area recorded: a median fill quotes the national average size, while kNN asks what the nearest comparable listings measured and iterative imputation predicts it from bedrooms, price and postcode.
saying these in an interview costs you the question
- Assumes model-based imputation always beats a median fill
- Runs kNN imputation without putting features on comparable scales
- Includes the target among the imputation model's predictors
- Ignores that kNN needs reference rows wherever you score
- Treats filled values as if they were observed measurements