How do scikit-learn's SimpleImputer, KNNImputer and IterativeImputer differ?
answer
- a fill value is learned, not computed inline
- one number per column, or a model
- cost grows from lookup to search to fit
- the fancy ones need the whole row
- one of the three is still experimental
basics
~20 sSimpleImputer learns one constant per column (mean, median, most frequent or a fixed value) and stores it in statistics_. KNNImputer fills from similar training rows, and IterativeImputer models each column from the others — both far more expensive.
solid answer
~40 sAll three are transformers that learn their fill strategy from training data only. `SimpleImputer(strategy=...)` computes one statistic per column — `'mean'`, `'median'`, `'most_frequent'` or a literal `'constant'` with `fill_value` — and stores it in `statistics_`; it is cheap, works on strings with `'most_frequent'`, and is the default choice. `KNNImputer` keeps the training rows and fills a missing entry from its `n_neighbors` nearest complete-enough rows using the `nan_euclidean` metric, so transform cost grows with training size. `IterativeImputer` treats each feature with missing values as a regression target on the remaining features and cycles round-robin for `max_iter` passes; as of scikit-learn 1.9 it is still experimental, so you must run `from sklearn.experimental import enable_iterative_imputer` before importing it. Any of them can add `add_indicator=True` to append binary columns marking where values were missing.
code
python · 7 linesimport numpy as np
from sklearn.impute import SimpleImputer
X = np.array([[1.0, np.nan], [3.0, 2.0], [np.nan, 4.0]])
imp = SimpleImputer(strategy="median", add_indicator=True).fit(X)
print(imp.statistics_) # [2. 3.]
print(imp.transform(X)) # 2 filled columns + 2 missingness flagsgo deeper
Know that scikit-learn imputers are fitted objects: the fill value is learned from training data and then applied, rather than recomputed on whatever frame you hand them.
Be able to name statistics_, the four SimpleImputer strategies, and describe how KNNImputer and IterativeImputer differ in what they store and what they cost.
Show judgment about cost and serving: KNNImputer ships the training data in the artifact, IterativeImputer runs models per prediction, and add_indicator usually beats a fancier fill. Justify choices with a measured comparison on fixed splits.
Own the question of whether the missingness is random at all. If it is systematic, no imputer fixes it and the answer is upstream: instrumentation, collection, or a model that treats missingness as a first-class feature.
## The shared contract All three live in `sklearn.impute` and are ordinary transformers: `fit()` learns whatever it needs from training data, `transform()` fills. That matters because the naive alternative — filling with a mean computed over the whole frame before splitting — derives the fill value from rows that include the test set, so information crosses the split. Imputers exist so the fill value becomes part of the fitted model. All of them accept `missing_values` (default `np.nan`, but you can point it at a sentinel like `-999`), `add_indicator`, and `keep_empty_features`, which controls whether a column that was entirely missing during fit survives as zeros rather than being dropped. ## SimpleImputer One statistic per column, held in `statistics_`: - `'mean'` and `'median'` — numeric only; median is the robust choice when the column is skewed or has outliers. - `'most_frequent'` — works on numeric and string/object columns, and is what you use ahead of an encoder. - `'constant'` with `fill_value` — fills with a literal, including a string like `'missing'` that then becomes its own category downstream. It is fast, its fitted state is tiny, and it is what you should reach for first. Its weakness is that it flattens variance: filling everything with one number pulls the column towards its centre and understates uncertainty, which biases downstream variance estimates. ## KNNImputer For each row with a missing value, it finds the `n_neighbors` closest training rows using the `nan_euclidean` distance — a Euclidean distance computed over the coordinates both rows have present, then rescaled for the number of missing coordinates — and fills with their average, optionally distance-weighted via `weights='distance'`. It preserves relationships between correlated columns far better than a column mean. The costs are real: the fitted object retains the training matrix, transform is a neighbour search per row rather than a lookup, and because distance is Euclidean the features should be on comparable scales or the largest-unit column dominates the neighbourhood. ## IterativeImputer The most expressive of the three. It initialises the missing entries (`initial_strategy`, default `'mean'`), then treats each column with missing values in turn as the regression target and the others as predictors, fitting the supplied `estimator` (a `BayesianRidge` by default) and replacing that column's missing entries with predictions. It repeats the round-robin for `max_iter` passes or until the change falls below `tol`. `imputation_order` controls the sweep order, and `sample_posterior=True` draws from the posterior instead of taking the point estimate, which is how you build multiple-imputation-style variability. Two API facts to state: as of scikit-learn 1.9 it is still gated behind the experimental flag, so `from sklearn.experimental import enable_iterative_imputer` must execute before `from sklearn.impute import IterativeImputer` or the import fails; and it fits one model per column per iteration, so on a wide matrix it is dramatically slower than the alternatives, at both fit and transform time. ## The missingness indicator `add_indicator=True` appends binary columns — the same thing `MissingIndicator` produces standalone — flagging which entries were missing. This is often the highest-value line in the whole preprocessing stack, because missingness is frequently informative: a blank income field may say more about the applicant than any imputed value would. Filling without an indicator destroys that signal; filling with one lets the model decide. ## Choosing Start with `SimpleImputer(strategy='median')` for numeric columns and `'most_frequent'` or `'constant'` for categorical, both with `add_indicator=True`, and treat that as the baseline. Move to `KNNImputer` or `IterativeImputer` only when a measured comparison — same splits, same model — shows the extra cost buys accuracy. The sophisticated imputers also complicate serving: `KNNImputer` ships the training data inside the artifact, and `IterativeImputer` runs a set of models on every prediction. ## How to answer Name what each one learns (`statistics_`, the training rows, per-column estimators), name the cost of each, mention the experimental import for `IterativeImputer`, and finish with `add_indicator`. The interviewer is checking whether you know these are fitted objects rather than one-line frame operations.
- Why is add_indicator often more valuable than the choice of fill strategy?Because missingness is frequently informative — a blank field may signal something the value itself never would. Imputing without an indicator overwrites that signal permanently; adding binary flags keeps it and lets the model weigh it. Comparing strategies rarely moves the score as much as recovering that column does.
- What must you do before importing IterativeImputer in scikit-learn 1.9?Execute `from sklearn.experimental import enable_iterative_imputer` first. The estimator is still flagged experimental, and that import is what registers it in `sklearn.impute`; without it the import raises. It also signals that the API may change between releases, so pin your version.
- Why should features be on comparable scales before KNNImputer?It fills from neighbours found with the `nan_euclidean` metric, so a column measured in tens of thousands dominates the distance and the 'nearest' rows are chosen almost entirely by that one feature. Put the columns on comparable footing first, keeping in mind that the scaler itself must tolerate NaN if it runs before the imputer.
saying these in an interview costs you the question
- Computing the fill value over the whole dataset before the split
- Using mean imputation on a heavily skewed column
- Assuming IterativeImputer imports without the experimental flag
- Thinking imputation adds no bias because nothing errors
- Dropping every row with a missing value by default