When is scikit-learn's OrdinalEncoder a safe choice instead of OneHotEncoder?
answer
- one column of numbers, not many
- the numbers look like magnitudes
- alphabetical order is still an order
- trees split, linear models multiply
- explicit categories for real orderings
basics
~20 sOrdinalEncoder maps each category to an integer in one column, which implies an ordering. That is safe for genuinely ordered features and for tree-based models that split on thresholds, but misleading for linear models, SVMs and distance-based estimators.
solid answer
~40 s`OrdinalEncoder` replaces each category with an integer code, keeping one output column per input column; `OneHotEncoder` expands to one binary column per category. The integer codes carry an implied ordering and implied distances, so any estimator that reads a feature as a magnitude — `LogisticRegression`, `Ridge`, `SVC`, `KNeighborsClassifier` — will conclude that `red=0` is closer to `green=1` than to `blue=2`, which is nonsense for an unordered column. Tree-based estimators only ask 'is this value above a threshold?', so integer codes are usually acceptable there and keep the feature matrix narrow. The other safe case is a genuinely ordinal column, where you pass the order explicitly as `categories=[['small', 'medium', 'large']]` rather than relying on the alphabetical default. For unseen values, set `handle_unknown='use_encoded_value'` together with `unknown_value=-1`.
code
python · 10 linesfrom sklearn.preprocessing import OrdinalEncoder
enc = OrdinalEncoder(
categories=[["small", "medium", "large"]],
handle_unknown="use_encoded_value",
unknown_value=-1,
)
enc.fit([["small"], ["large"]])
print(enc.transform([["medium"], ["huge"]])) # [[ 1.]
# [-1.]]go deeper
Be able to say that OrdinalEncoder gives one column of integer codes while OneHotEncoder gives one binary column per category, and that the integers imply an order that may not exist.
Explain why the implied ordering harms linear, SVM and distance-based models but not tree splits, and name the explicit categories argument for real ordinal features.
Bring the cardinality tradeoff: one-hot widens the matrix and slows tree training, so ordinal codes or native categorical support win on wide columns. Handle unseen values with use_encoded_value rather than letting production raise.
Own the encoding policy across a feature platform: which columns are genuinely ordinal, how new categories are admitted over time, and whether encoders are shared artifacts or refit per model.
## The two encoders side by side Both learn a per-column vocabulary during `fit()` and expose it as `categories_`. They differ in what `transform()` emits. `OrdinalEncoder` produces one float column per input column, holding the index of each value within `categories_`. Three categories become the codes 0.0, 1.0, 2.0. Width is unchanged; nothing is sparse. `OneHotEncoder` produces one binary column per category. Three categories become three columns. Width grows with cardinality, and the output is a sparse matrix by default. ## Why the integer codes are dangerous With the default `categories='auto'` the codes are assigned in sorted order, which for strings means alphabetical. A `city` column encoded this way tells a linear model that Amsterdam < Berlin < Cairo and, worse, that the gap between Amsterdam and Berlin equals the gap between Berlin and Cairo. Every estimator that multiplies a feature by a coefficient, or measures a distance in feature space, will act on that fiction: - linear and logistic regression fit one coefficient for the whole column, forcing a monotone effect across an arbitrary alphabetical order; - SVMs with an RBF kernel and k-nearest-neighbours compute distances in which the numeric gaps are literal; - anything preceded by a scaler will happily standardise the fake magnitudes. Nothing raises. The model trains, produces a plausible score, and is systematically wrong about that feature. This is the classic silent-wrong-answer of scikit-learn preprocessing. ## Where it is fine **Genuine ordinal features.** Sizes, education levels, satisfaction ratings and severity grades have a real order, and collapsing them to one column preserves information a one-hot encoding throws away. Pass the order explicitly — `OrdinalEncoder(categories=[['low', 'medium', 'high']])` — because the alphabetical default would give you high < low < medium. **Tree-based models.** Decision trees, random forests and gradient boosting split on `feature <= threshold`. A tree can isolate any single category by stacking two splits, so an arbitrary integer ordering costs depth rather than correctness. In exchange you keep the matrix narrow, which matters a great deal for high-cardinality columns where one-hot encoding would produce thousands of nearly-empty columns and slow every split search down. scikit-learn's histogram gradient-boosting estimators also accept a `categorical_features` argument, which handles the column natively rather than through either encoder. ## The parameters that come up - `categories` — pass a list of lists to fix the order per column instead of inferring it. - `handle_unknown='use_encoded_value'` with `unknown_value` — required together; unseen values get that code (commonly `-1` or `np.nan`) rather than raising. The default is `'error'`. - `encoded_missing_value` — the code used for `NaN` inputs, which by default are encoded as `np.nan` rather than treated as a category. - `min_frequency` and `max_categories` — the same rare-category grouping the one-hot encoder offers, so a long tail collapses into one code. ## The related mistake: LabelEncoder on features A very common error is reaching for `LabelEncoder` to encode a feature column. `LabelEncoder` is documented for the target `y` only: it accepts a 1-D array, has no notion of multiple columns, and does not fit the transformer contract that `Pipeline` and `ColumnTransformer` require. `OrdinalEncoder` is the 2-D feature-space equivalent and is the one to name in an interview. ## How to answer Lead with the implied ordering — that is the whole point of the question. Then split the world in two: estimators that read magnitudes (use one-hot, or a native categorical handler) and estimators that split on thresholds (ordinal codes are fine, and cheaper). Close with the explicit-`categories` trick for real ordinal columns and the `use_encoded_value`/`unknown_value` pair for production robustness.
- Why not just use LabelEncoder on the feature columns?`LabelEncoder` is intended for the target `y`: it takes a 1-D array, cannot handle a 2-D feature matrix, and does not implement the transformer contract that `Pipeline` and `ColumnTransformer` rely on. `OrdinalEncoder` is the feature-space equivalent and works on multiple columns at once.
- What does OrdinalEncoder do with an unseen category by default, and how do you change it?By default `handle_unknown='error'` raises a `ValueError` at transform time. Set `handle_unknown='use_encoded_value'` and supply `unknown_value` — commonly `-1` or `np.nan` — and the two must be given together. Choose a code that cannot collide with a real category index.
- You have a 5,000-category column and you are training gradient boosting. One-hot or ordinal?Ordinal, or the estimator's native categorical support. One-hot would add 5,000 near-empty columns and slow every split search; a tree can still isolate categories through repeated splits on the integer codes. Combine with `max_categories` or `min_frequency` to collapse the rare tail.
Numbering the cities on a map 0, 1, 2 does not make city 0 nearer to city 1 than to city 2 — but a model that only sees the numbers will believe it does.
saying these in an interview costs you the question
- Claiming integer codes are always fine because models are numeric
- Using LabelEncoder on feature columns
- Relying on alphabetical order for a genuinely ordinal column
- One-hot encoding every categorical column regardless of cardinality
- Thinking scaling the integer codes removes the implied ordering