How does scikit-learn's OneHotEncoder handle categories unseen during fit()?
answer
- the vocabulary is frozen at fit
- transform meets a value fit never saw
- the default is loud, not lenient
- handle_unknown chooses raise vs zeros
- all-zero row means none of the known
basics
~10 sBy default OneHotEncoder raises a ValueError at transform time for any category it did not see in fit(). Setting handle_unknown='ignore' encodes the unknown value as an all-zero row instead, so scoring continues.
solid answer
~40 s`OneHotEncoder` learns the category list per column during `fit()` and stores it in `categories_`. At `transform()` time the default `handle_unknown='error'` raises a `ValueError` when a value is not in that list — safe, but it means one new city name in production takes the service down. `handle_unknown='ignore'` instead emits a row of all zeros for that column's block, so the model still scores, but it cannot distinguish 'unknown' from 'none of the known categories'. The third option, `handle_unknown='infrequent_if_exist'`, routes unknowns into the rare-category bucket when you have enabled grouping via `min_frequency` or `max_categories`, and falls back to all zeros when you have not. Note that combining `drop='first'` with `handle_unknown='ignore'` is ambiguous: the dropped reference category is also encoded as all zeros, so unknown and reference become indistinguishable.
code
python · 8 linesfrom sklearn.preprocessing import OneHotEncoder
enc = OneHotEncoder(handle_unknown="ignore", sparse_output=False)
enc.fit([["cat"], ["dog"]])
print(enc.categories_) # [array(['cat', 'dog'], dtype=object)]
print(enc.get_feature_names_out()) # ['x0_cat' 'x0_dog']
print(enc.transform([["dog"], ["fox"]])) # [[0. 1.]
# [0. 0.]] <- unknowngo deeper
Know that one-hot encoding turns each category into its own 0/1 column, and that a value the encoder never saw during fit is a problem you have to configure for.
Be able to name handle_unknown and its modes, describe the all-zero encoding, and explain that categories_ is the frozen vocabulary that fixes the output width.
Show production judgment: choose 'ignore' or 'infrequent_if_exist' for a live service, monitor how often the unknown path fires, and use min_frequency/max_categories to keep the encoding width bounded as cardinality grows.
Own the policy for how new categories enter a served model at all: whether the encoder vocabulary is frozen per release, how drift in the unknown rate triggers a refit, and who is accountable when a new value silently degrades predictions.
## What fit() learns `OneHotEncoder` is a transformer whose learned state is the category vocabulary. With the default `categories='auto'`, `fit(X)` collects the sorted unique values of each input column and stores them in the list `categories_`, one array per column. `transform()` then produces one binary column per known category, in that stored order, and `get_feature_names_out()` names them `<column>_<category>`. Because the vocabulary is frozen at fit time, the output width is fixed — which is exactly what a downstream estimator needs, and exactly why unseen values are a problem. ## The three handle_unknown modes **`'error'` (default).** A value absent from `categories_` raises `ValueError` during `transform()`. This is the right default: it is loud, it surfaces schema drift immediately, and it prevents you from silently scoring garbage. It is also the wrong setting for a long-lived service where a new category is expected and must not cause an outage. **`'ignore'`.** The unknown value produces zeros across that column's whole block. Nothing raises. The cost is representational: the model sees a vector that says 'none of the categories I know', which is a real signal but not a distinguishable one — every different unknown value maps to the same all-zero encoding. **`'infrequent_if_exist'`.** When you have enabled rare-category grouping, unknown values are mapped into the infrequent bucket, so an unknown is treated like a rare-but-seen value rather than like nothing at all. Where no grouping is active for that column, it degrades to the all-zero behaviour of `'ignore'`. ## Rare-category grouping High-cardinality columns are the practical reason one-hot encoding goes wrong: a `user_agent` column with 40,000 distinct values produces 40,000 columns, most of which are almost always zero. `min_frequency` (an integer count or a float fraction) and `max_categories` (a cap on the number of output columns) collapse the tail into a single infrequent bucket; the groupings are readable afterwards through `infrequent_categories_`. This is a scikit-learn feature rather than a general technique, and it is worth naming — the alternative many candidates reach for is hand-rolled tail lumping in pandas before the split, which leaks and is harder to serve. ## drop and its interaction with unknowns `drop='first'` removes one column per feature to eliminate the perfect collinearity that upsets unregularised linear models; `drop='if_binary'` drops the second column only for two-category features, which is the usually-sensible middle ground. Combining a `drop` with `handle_unknown='ignore'` creates the ambiguity noted above: the dropped category is represented by all zeros, and so is an unknown, so the two collide. Tree ensembles do not care about collinearity, so dropping is mostly a linear-model concern. ## Output type `sparse_output` (the parameter was named `sparse` before scikit-learn 1.2) defaults to `True`, so the result is a SciPy CSR matrix. That is memory-sensible for wide encodings, and most scikit-learn estimators accept sparse input, but code that expects a NumPy array — or a `DataFrame` — needs `sparse_output=False`, or `set_output(transform='pandas')` if you want named columns back. Inside a `ColumnTransformer`, whether the final block is sparse or dense is decided by that container's `sparse_threshold`, not by the encoder alone. ## Missing values `np.nan` and `None` are treated as their own category by `OneHotEncoder` rather than raising, so a column with missing values quietly gains an extra indicator column. If you would rather impute first, put a `SimpleImputer(strategy='most_frequent')` ahead of the encoder in the pipeline; if missingness is informative, leaving it as a category is a defensible choice — just make it deliberate rather than accidental. ## How to answer State the default is an error, name `handle_unknown='ignore'` and its all-zero encoding, and then say what you would actually do in a production service: `'ignore'` or `'infrequent_if_exist'` with monitoring on how often the unknown path fires, because an encoder that silently zeros out a growing share of your traffic is a model quietly degrading.
- You have a categorical column with 40,000 distinct values. What do you do before one-hot encoding it?Use the encoder's own tail-lumping rather than encoding all of it: `min_frequency` (count or fraction) and `max_categories` collapse rare values into a single infrequent bucket, inspectable via `infrequent_categories_`. Doing the lumping inside the encoder means the threshold is learned from training data only and ships with the fitted object.
- Why is drop='first' together with handle_unknown='ignore' a problem?Both are encoded as all zeros. The dropped reference category is represented by the absence of any one-hot column, and an unknown category under 'ignore' produces exactly the same zero vector, so the model cannot distinguish them. Prefer `drop='if_binary'`, or keep all columns and rely on regularisation.
- What does OneHotEncoder do with NaN values in a column?It treats missing as just another category and gives it its own output column, rather than raising. That is fine when missingness is informative, but it should be a decision: if you want it filled instead, put a `SimpleImputer` ahead of the encoder in the pipeline so the fill value is learned from training data.
saying these in an interview costs you the question
- Assuming unknown categories are silently ignored by default
- Fitting the encoder on train and test concatenated
- Thinking all-zero encoding distinguishes different unknown values
- Believing OneHotEncoder returns a dense array by default
- One-hot encoding a 40,000-value column without any grouping