skip to content

Categorical and Text Encoding

Turning labels and free text into numbers: one-hot columns, out-of-fold target encoding, the hashing trick and TF-IDF vectors for a tabular model. Interviewers probe high cardinality and leakage.

on this pageshow

explore

questions

17

Why is coding education level as 0-3 acceptable but coding job family the same way risky?

level: juniorimportance: must knowfreq 78%

answer

  1. does the order exist in the world?
  2. ordered versus merely labelled
  3. one coefficient forces a monotone effect
  4. codes also claim equal gaps
  5. one indicator column per level when unordered

basics

~20 s

Education level has a real order - high-school below bachelor below master below PhD - so integer codes carry that information. Job family has no order, so codes 0-3 invent a ranking the model takes literally. Encode unordered categories as one-hot indicators instead.

solid answer

~50 s

An ordinal (integer) code says two things to the model: these levels have an order, and the gap between consecutive codes is the same size. For education that first claim is true, so a single numeric column is defensible and cheap. For job family it is false: coding sales=0, legal=1, engineering=2, support=3 tells a linear model that legal sits between sales and engineering and that support is three units away from sales, which is meaningless. A linear model then fits one coefficient that forces a monotone, evenly spaced effect across an arbitrary alphabetical order, and a distance-based model computes distances off the same fiction. One-hot encoding gives each level its own indicator column and its own weight, so the model can learn any pattern across levels with no ordering assumed. The cost is one column per level instead of one column total.

go deeper

for a junior

Be ready to state the difference in one breath: ordinal codes assume an order, one-hot assumes none. Then name a column of each kind and say which encoding you would use.

for a middle

Explain the mechanism, not just the rule - one coefficient times the code forces a monotone, equally spaced effect, and distances are computed on the same invented scale.

for a senior

An interviewer expects you to say what actually goes wrong downstream: silent bias in a linear model, distorted neighbourhoods, and extra depth burned by trees recovering a grouping the ordering scattered.

for a principal

Own the tradeoff at pipeline scale: a blanket one-hot rule is safe but expensive across dozens of columns, so decide per column and write the ordering assumption down where the next person will find it.

## The two encodings A categorical column holds labels, and a model needs numbers. There are two elementary ways to get them. **Ordinal (integer) encoding** replaces each level with one integer: `high-school -> 0, bachelor -> 1, master -> 2, PhD -> 3`. The column count does not change - one text column becomes one numeric column. **One-hot encoding** creates one binary indicator per level. A row whose education is `master` gets `[0, 0, 1, 0]`. Exactly one entry is 1 and the rest are 0, which is where the name comes from. A column with `k` levels becomes `k` columns (or `k-1` if a reference level is dropped). ## What an integer code actually asserts Writing `bachelor = 1` and `master = 2` is not a neutral relabelling. It asserts two things: 1. **Order.** Level 2 is greater than level 1 in whatever sense the model uses numbers. 2. **Equal spacing.** The distance from 0 to 1 equals the distance from 1 to 2. For education, claim 1 is true in the world. Claim 2 is arguable - the jump from bachelor to master may not be worth the same as the jump from master to PhD - but it is a modelling simplification, not a fabrication. For job family, claim 1 is already false. `sales`, `legal`, `engineering`, `support` have no natural order at all; whatever integers you assign come from alphabetical order or the order the levels happened to appear in the data. The model has no way to know the ordering is arbitrary. It will use it. ## How each model family is hurt **Linear models.** A single coefficient `w` multiplies the code, so the fitted effect is `w * code`. That forces the effect to move monotonically and in equal steps as the code increases. If the true pattern is that `legal` and `support` behave alike while `engineering` differs, no single `w` can express it - the model is structurally unable to fit the truth, and it will spend that inability as bias. **Distance-based models** (nearest-neighbour style methods, clustering). Distance is computed on the numbers, so `sales` (0) and `support` (3) are treated as far apart while `sales` and `legal` are treated as near neighbours. The neighbourhoods are built on an ordering that does not exist. **Trees and tree ensembles** are the mildest case, and this is the reason integer codes survive as long as they do in practice. A tree splits on a threshold such as `code <= 1.5`, which carves the levels into two groups. Because the threshold is a cut on the number line, the groups it can form are always **contiguous ranges of the arbitrary ordering**. To isolate `sales` and `support` together against the others, the tree needs several splits stacked on the same column, which costs depth and data. So a tree is not immune - it is merely capable of recovering, given enough splits, from an ordering it was handed for no reason. ## The practical rule Ask one question about the column itself, not about the model: **does the order exist in the world?** - Yes, and equal spacing is a tolerable simplification -> an integer code is fine and cheap. Education level, T-shirt sizes, a rating band. - Yes, but you doubt equal spacing and have enough rows -> one-hot lets each level carry its own weight and the model discovers the spacing, including a non-monotonic pattern. - No -> one-hot. Job family, browser, colour, country. ## The equal-spacing caveat, concretely A 1-5 satisfaction response is genuinely ordered, but there is no reason the step from 4 to 5 is worth the same as the step from 1 to 2 - respondents cluster at the top, and the difference between very satisfied and satisfied may drive far more behaviour than the difference between two flavours of unhappy. Coding it 1-5 bakes equal gaps in. One-hot removes the assumption at the cost of four extra columns and the loss of the ordering information the column did carry. ## What one-hot costs One-hot is the safe default for unordered columns, not a free one. It multiplies the column count by the number of levels, produces a matrix that is mostly zeros, and gives every level its own parameter - which means a level with a handful of rows gets a parameter estimated from a handful of rows. That is a real tradeoff, not a reason to fall back on inventing an order.

  • Even when the order is genuinely real, what does an integer code still assume?
    That the gaps are equal. A single coefficient multiplies the code, so the fitted effect moves in equal steps from level to level. If the jump from master to PhD matters more than the jump from high-school to bachelor, the code cannot say so. One-hot drops the assumption and lets each level carry its own weight, at the cost of extra columns and thinner data per weight.
  • Does a decision tree care whether the integer codes are in a meaningful order?
    Less than a linear model, but it is not indifferent. A tree splits at a threshold on the code, so any group it forms is a contiguous run of the imposed ordering. If the levels that behave alike are scattered across that ordering, the tree needs several stacked splits to reunite them, spending depth and data it would not spend on one-hot indicators.
  • Are the gaps between the levels of a 1-5 satisfaction rating equal?
    There is no reason to assume so. Respondents bunch at the top, and the behavioural difference between 4 and 5 is often larger than between 1 and 2. Feeding the raw 1-5 code to a linear model asserts equal gaps; one-hot indicators let the model learn the real spacing, including a non-monotonic one, if you have the rows to support five separate weights.

T-shirt sizes S, M, L can be numbered 1, 2, 3 because they really do line up. Numbering shirt colours red, green, blue the same way tells the model that green sits between red and blue.

saying these in an interview costs you the question

  • Integer-codes every categorical column because it saves memory
  • Claims tree models make the encoding choice irrelevant
  • Treats one-hot and ordinal encoding as two names for the same thing
  • Assumes alphabetical code order encodes something meaningful
  • Says a nominal integer code is fine because the model will figure it out

context

open as a page

In a bag-of-words matrix of movie reviews, what does TF-IDF weighting do to words like 'the' and 'and'?

level: juniorimportance: must knowfreq 78%

basics

~20 s

TF-IDF multiplies each word's count by an inverse-document-frequency factor that shrinks as the word appears in more documents. Words like 'the' and 'and' sit in nearly every review, so their columns collapse to near-zero while rarer words keep large weights.

open as a page

What does an information value of 0.015 tell you about a scorecard feature?

level: juniorimportance: must knowfreq 58%

basics

~20 s

An information value of 0.015 sits below the conventional 0.02 floor, so on its own the feature barely separates defaulters from non-defaulters. Standard practice is to drop it from the scorecard unless policy requires it or it earns its place alongside other characteristics.

open as a page

Why does target-encoding a 40,000-level seller ID leak, and how does out-of-fold encoding fix it?

level: middleimportance: must knowfreq 72%

basics

~20 s

Computing each seller's target mean over all training rows puts a row's own label inside its own feature, so the model reads the answer back. Out-of-fold encoding builds each fold's means from the other folds only.

open as a page

Why must a TF-IDF vocabulary and its idf values be fitted on the training fold only?

level: middleimportance: must knowfreq 55%

basics

~20 s

The vocabulary and the idf values are learned parameters. Building them over the whole corpus before splitting lets held-out documents shape the features that describe them, so validation scores come out optimistically biased and overstate what production will do.

open as a page

Why do credit scorecards replace binned features with their weight of evidence?

level: middleimportance: must knowfreq 68%

basics

~20 s

Weight of evidence replaces a bin with ln(share of non-defaults / share of defaults), which is the bin's log-odds offset from the population. That makes every feature monotone, on one common scale, linear in log-odds, auditable, and lets missing values form their own bin.

open as a page

What does count encoding do to a 30,000-level ZIP code column, and when does the count itself carry signal?

level: juniorimportance: should knowfreq 41%

basics

~20 s

Count encoding replaces each ZIP code with how many rows carry it, turning 30,000 levels into one numeric column. It helps when frequency proxies something real, like population density. It uses no labels, so it cannot leak.

open as a page

What does one-hot encoding a 50-level US state column do to a linear model's fit?

level: middleimportance: should knowfreq 56%

basics

~20 s

It adds fifty binary columns that are 98 percent zeros and gives every state its own weight. Each weight is driven only by that state's rows, so small states get unstable estimates and total variance rises. Penalising the weights becomes close to mandatory.

open as a page

In target encoding, a seller with three historical rows all converted — smooth toward the prior or bucket it as rare?

level: seniorimportance: should knowfreq 46%

basics

~10 s

Three rows cannot support a rate of 1.0. Smoothing shrinks the estimate toward the global rate, weighted by the row count, so some seller signal survives; a rare bucket discards it entirely. Prefer smoothing.

open as a page

What does a one-hot encoder emit for a browser string never seen during training?

level: seniorimportance: should knowfreq 50%

basics

~20 s

Either it raises an error or it emits all zeros across that column's indicators. All zeros is the dangerous case: if a reference level was dropped, that pattern already means the reference level, so the unknown browser is silently scored as it.

open as a page

Your TF-IDF matrix has 50,000 columns for 20,000 documents — how do you decide what to prune?

level: seniorimportance: should knowfreq 41%

basics

~20 s

Check the arithmetic before pruning: stored sparsely, 20,000 documents touching a hundred terms each is a few million values, not a billion. Cut with document-frequency thresholds at both ends, and let held-out folds decide how far to go.

open as a page

A scorecard's age bands break monotonic WoE in one bin — how do you fix the binning?

level: seniorimportance: should knowfreq 52%

basics

~20 s

First decide whether the dip is real or sampling noise in a thin band. If it is noise, merge that band with the neighbour it sits closest to in weight of evidence and re-cut under a monotone constraint, accepting a small loss of information value.

open as a page

How do you choose an encoding for a 40,000-level seller ID that must be refreshed and served daily?

level: principalimportance: should knowfreq 37%

basics

~20 s

Decide on four axes: how much signal the identity carries, what the model family can consume, what state serving can refresh, and who must explain the feature. Start cheap and escalate only when a holdout says it pays.

open as a page

How does the hashing trick encode millions of ad publisher domains, and what do collisions cost?

level: middleimportance: nice to knowfreq 34%

basics

~20 s

A hash function maps each domain to a bucket index modulo a fixed bucket count, say 2^20, and that bucket is the feature slot. Nothing is stored, so new domains need no special case. Colliding domains share one blended weight.

open as a page

When do character 3-5-grams beat word unigrams as features for matching messy product titles?

level: middleimportance: nice to knowfreq 34%

basics

~20 s

Character n-grams win when the tokens themselves are unreliable: misspelled or run-together brand names, inconsistent punctuation, model codes. A typo changes only a few of a word's character n-grams, whereas it destroys the word unigram entirely, so overlap survives.

open as a page

Should a one-hot column drop a reference level before an L2-penalised linear fit?

level: seniorimportance: nice to knowfreq 30%

basics

~20 s

Usually no. Dropping a level exists to remove the exact redundancy between the full set of indicators and the intercept, which an L2 penalty already resolves on its own. Keeping every level and leaving the intercept unpenalised treats all levels symmetrically.

open as a page

How do you turn a scorecard's fitted log-odds into points at base 600 with a PDO of 20?

level: seniorimportance: nice to knowfreq 26%

basics

~20 s

Apply score = offset + factor times ln(odds), where factor = PDO / ln 2, so 20 / 0.693 is about 28.85, and offset = 600 minus factor times the log of the anchor odds. Every doubling of the odds then adds exactly 20 points.

open as a page