Why does a Spark MLlib StringIndexer fail on a category it never saw during fit()?
answer
- the vocabulary is frozen when you fit
- three strategies, one of them is the default
- the default is the noisy one
- unseen labels land at index numLabels
basics
~20 sStringIndexer learns its label vocabulary during fit(), and its handleInvalid parameter defaults to error, so any category absent from the training data throws. Set it to keep, which assigns the extra index numLabels, or skip, which drops the row.
solid answer
~40 s`StringIndexer` is an Estimator: `fit()` learns the set of label strings in the training column and produces a `StringIndexerModel` holding that vocabulary. At transform time anything outside it is an *unseen label*, and the `handleInvalid` parameter decides what happens. There are three strategies, and the default is to throw an exception. `setHandleInvalid("skip")` drops the offending rows entirely; `setHandleInvalid("keep")` puts every unseen label into one extra bucket at index `numLabels`. In production, `error` turns a brand-new category into an outage, and `skip` silently deletes rows from a scoring batch so the output has fewer rows than the input — usually the worse failure, because nothing alerts. `keep` is normally the right choice, but remember to set `handleInvalid` consistently on the downstream `OneHotEncoder` too, since the extra index widens the encoded vector.
code
python · 13 linesfrom pyspark.ml.feature import StringIndexer, OneHotEncoder
indexer = StringIndexer(
inputCol="category",
outputCol="categoryIndex",
handleInvalid="keep", # default is "error"
stringOrderType="frequencyDesc",
)
encoder = OneHotEncoder(
inputCol="categoryIndex",
outputCol="categoryVec",
handleInvalid="keep", # must match, or the extra index errors
)go deeper
Know that a StringIndexer learns its label vocabulary during fit() and that a category it never saw causes an error by default, rather than quietly becoming zero.
Explain all three handleInvalid options, say which index a kept label receives, and be able to work out the indices the default frequencyDesc ordering assigns to a small example.
Discuss what each option costs in production: skip silently deletes rows from a scoring batch, keep widens downstream vectors and hides new categories in one bucket, and error turns a new value into an outage.
Own the policy across the platform: which categorical columns are allowed to grow, how new values are detected and reviewed, what fraction landing in the unseen bucket triggers a re-fit, and who owns that decision.
## What StringIndexer learns `StringIndexer` encodes a string column of labels into a column of numeric indices in the range `[0, numLabels)`. It is an *Estimator*, not a plain transformer, because it must see the data first: `fit()` scans the input column, collects the distinct label strings, and returns a `StringIndexerModel` that carries that vocabulary. Everything about later behaviour follows from the fact that this vocabulary is frozen at fit time. Four ordering options control which label gets which index, via `stringOrderType`: `frequencyDesc` (the default — the most frequent label gets 0), `frequencyAsc`, `alphabetDesc` and `alphabetAsc`. Under the frequency orderings, labels with equal frequency are further sorted alphabetically. If the input column is numeric it is cast to string and the string values are indexed. The documented example makes the default concrete. Fitting on a `category` column holding `a, b, c, a, a, c` gives `a` index 0.0 (three occurrences), `c` index 1.0 (two), and `b` index 2.0 (one) — frequency order, not alphabetical order. ## The three unseen-label strategies When a fitted `StringIndexerModel` meets a label that was not in the training column, `handleInvalid` chooses among three documented strategies: - **`error`** — throw an exception. **This is the default.** - **`skip`** — drop the row containing the unseen label entirely. - **`keep`** — put unseen labels in a special additional bucket, at index `numLabels`. So with the example above, fitting on `{a, b, c}` and then transforming rows containing `d` and `e`: the default throws; `skip` returns a DataFrame with those two rows missing; `keep` assigns both of them index `3.0`, since `numLabels` is 3. ## Why this bites in production Every real categorical column grows. A new merchant, a new device model, a new country code appears in yesterday's data and the pipeline meets it today. Each strategy fails in a characteristic way: - `error` is loud. The scoring job dies with an exception naming the column. Unpleasant, but you find out. - `skip` is silent and is usually the worst option at scoring time. Rows vanish from the output. If downstream consumers join predictions back onto the input by key, they get nulls; if they count rows, the count is quietly short. Nothing raises an alarm. `skip` is defensible during *training*, where dropping a handful of malformed rows is harmless, and dangerous during *inference*, where every input row is expected to produce a prediction. - `keep` is the usual production answer, but it is not free: every unseen category collapses into a single "other" bucket, so the model treats a brand-new high-value merchant identically to a typo. That is a modelling compromise you should make deliberately, and it is a signal worth monitoring — a rising share of rows landing at index `numLabels` means the vocabulary is stale. ## The downstream consequence `StringIndexer` rarely appears alone. Its output normally feeds a `OneHotEncoder`, which itself supports `handleInvalid` with the options `keep` (invalid inputs get an extra categorical index) and `error`. If the indexer keeps unseen labels but the encoder is left at `error`, the pipeline still fails — the encoder now receives an index it was not fitted for. Set the policy consistently across the stages that share it, and be aware that keeping the extra category changes the width of the encoded vector, which in turn changes the assembled `features` vector the model expects. ## The re-fitting trap A subtler failure has nothing to do with unseen labels. Because the default ordering is `frequencyDesc`, index 0 means "the most frequent label *in the data this indexer was fitted on*". Re-fit the indexer on a newer month where a different category became most frequent, and the meaning of every index shifts underneath a model whose coefficients were learned against the old assignment. The fix is not clever: persist the fitted `StringIndexerModel` (or the whole `PipelineModel`) and reuse it, rather than re-fitting a fresh indexer at scoring time. If you truly must re-fit independently, pin `stringOrderType` to `alphabetAsc` or `alphabetDesc` so the mapping depends only on the label set, not on its frequency distribution — and even then, adding a label reshuffles alphabetical order. ## How to answer Give the default first (`error`), then the two alternatives and the index that `keep` assigns. Then show judgment: name `skip` as the dangerous one at inference because it removes rows without telling anyone, name `keep` as the usual choice, and mention that the policy must be mirrored on the downstream encoder. Finish with the operational point — track how many rows land in the unseen bucket, because that number is your signal that the pipeline needs re-fitting.
- What does handleInvalid="skip" cost you at scoring time?Rows disappear from the output with no error and no log line, so the scored batch has fewer rows than the input. Downstream joins produce nulls, counts come up short, and nothing alerts. It is reasonable while training, where dropping a few malformed rows is harmless, and a poor default for inference where every input row must yield a prediction.
- Why can re-fitting a StringIndexer on newer data change the meaning of index 0?The default `stringOrderType` is `frequencyDesc`, so index 0 is whichever label was most frequent in the data used for fitting. If frequencies shift between months, the mapping shifts too, and a model trained against the old assignment now sees scrambled indices. Persist and reuse the fitted `StringIndexerModel` instead, or pin the ordering to `alphabetAsc`.
- If unseen labels are kept, what changes downstream in OneHotEncoder?The indexer can now emit `numLabels` as a valid index, one beyond what the encoder was fitted for. The encoder must also use `handleInvalid="keep"`, which assigns that input its own extra categorical index — and that widens the encoded vector, and therefore the assembled features vector the model consumes.
saying these in an interview costs you the question
- Says unseen categories map to zero by default
- Proposes re-fitting the indexer on the scoring data
- Uses skip in production and never checks the dropped rows
- Thinks StringIndexer assigns indices alphabetically by default
- Assumes a label's index is stable across separate fits