Why is coding education level as 0-3 acceptable but coding job family the same way risky?
answer
- does the order exist in the world?
- ordered versus merely labelled
- one coefficient forces a monotone effect
- codes also claim equal gaps
- one indicator column per level when unordered
basics
~20 sEducation level has a real order - high-school below bachelor below master below PhD - so integer codes carry that information. Job family has no order, so codes 0-3 invent a ranking the model takes literally. Encode unordered categories as one-hot indicators instead.
solid answer
~50 sAn ordinal (integer) code says two things to the model: these levels have an order, and the gap between consecutive codes is the same size. For education that first claim is true, so a single numeric column is defensible and cheap. For job family it is false: coding sales=0, legal=1, engineering=2, support=3 tells a linear model that legal sits between sales and engineering and that support is three units away from sales, which is meaningless. A linear model then fits one coefficient that forces a monotone, evenly spaced effect across an arbitrary alphabetical order, and a distance-based model computes distances off the same fiction. One-hot encoding gives each level its own indicator column and its own weight, so the model can learn any pattern across levels with no ordering assumed. The cost is one column per level instead of one column total.
go deeper
Be ready to state the difference in one breath: ordinal codes assume an order, one-hot assumes none. Then name a column of each kind and say which encoding you would use.
Explain the mechanism, not just the rule - one coefficient times the code forces a monotone, equally spaced effect, and distances are computed on the same invented scale.
An interviewer expects you to say what actually goes wrong downstream: silent bias in a linear model, distorted neighbourhoods, and extra depth burned by trees recovering a grouping the ordering scattered.
Own the tradeoff at pipeline scale: a blanket one-hot rule is safe but expensive across dozens of columns, so decide per column and write the ordering assumption down where the next person will find it.
## The two encodings A categorical column holds labels, and a model needs numbers. There are two elementary ways to get them. **Ordinal (integer) encoding** replaces each level with one integer: `high-school -> 0, bachelor -> 1, master -> 2, PhD -> 3`. The column count does not change - one text column becomes one numeric column. **One-hot encoding** creates one binary indicator per level. A row whose education is `master` gets `[0, 0, 1, 0]`. Exactly one entry is 1 and the rest are 0, which is where the name comes from. A column with `k` levels becomes `k` columns (or `k-1` if a reference level is dropped). ## What an integer code actually asserts Writing `bachelor = 1` and `master = 2` is not a neutral relabelling. It asserts two things: 1. **Order.** Level 2 is greater than level 1 in whatever sense the model uses numbers. 2. **Equal spacing.** The distance from 0 to 1 equals the distance from 1 to 2. For education, claim 1 is true in the world. Claim 2 is arguable - the jump from bachelor to master may not be worth the same as the jump from master to PhD - but it is a modelling simplification, not a fabrication. For job family, claim 1 is already false. `sales`, `legal`, `engineering`, `support` have no natural order at all; whatever integers you assign come from alphabetical order or the order the levels happened to appear in the data. The model has no way to know the ordering is arbitrary. It will use it. ## How each model family is hurt **Linear models.** A single coefficient `w` multiplies the code, so the fitted effect is `w * code`. That forces the effect to move monotonically and in equal steps as the code increases. If the true pattern is that `legal` and `support` behave alike while `engineering` differs, no single `w` can express it - the model is structurally unable to fit the truth, and it will spend that inability as bias. **Distance-based models** (nearest-neighbour style methods, clustering). Distance is computed on the numbers, so `sales` (0) and `support` (3) are treated as far apart while `sales` and `legal` are treated as near neighbours. The neighbourhoods are built on an ordering that does not exist. **Trees and tree ensembles** are the mildest case, and this is the reason integer codes survive as long as they do in practice. A tree splits on a threshold such as `code <= 1.5`, which carves the levels into two groups. Because the threshold is a cut on the number line, the groups it can form are always **contiguous ranges of the arbitrary ordering**. To isolate `sales` and `support` together against the others, the tree needs several splits stacked on the same column, which costs depth and data. So a tree is not immune - it is merely capable of recovering, given enough splits, from an ordering it was handed for no reason. ## The practical rule Ask one question about the column itself, not about the model: **does the order exist in the world?** - Yes, and equal spacing is a tolerable simplification -> an integer code is fine and cheap. Education level, T-shirt sizes, a rating band. - Yes, but you doubt equal spacing and have enough rows -> one-hot lets each level carry its own weight and the model discovers the spacing, including a non-monotonic pattern. - No -> one-hot. Job family, browser, colour, country. ## The equal-spacing caveat, concretely A 1-5 satisfaction response is genuinely ordered, but there is no reason the step from 4 to 5 is worth the same as the step from 1 to 2 - respondents cluster at the top, and the difference between very satisfied and satisfied may drive far more behaviour than the difference between two flavours of unhappy. Coding it 1-5 bakes equal gaps in. One-hot removes the assumption at the cost of four extra columns and the loss of the ordering information the column did carry. ## What one-hot costs One-hot is the safe default for unordered columns, not a free one. It multiplies the column count by the number of levels, produces a matrix that is mostly zeros, and gives every level its own parameter - which means a level with a handful of rows gets a parameter estimated from a handful of rows. That is a real tradeoff, not a reason to fall back on inventing an order.
- Even when the order is genuinely real, what does an integer code still assume?That the gaps are equal. A single coefficient multiplies the code, so the fitted effect moves in equal steps from level to level. If the jump from master to PhD matters more than the jump from high-school to bachelor, the code cannot say so. One-hot drops the assumption and lets each level carry its own weight, at the cost of extra columns and thinner data per weight.
- Does a decision tree care whether the integer codes are in a meaningful order?Less than a linear model, but it is not indifferent. A tree splits at a threshold on the code, so any group it forms is a contiguous run of the imposed ordering. If the levels that behave alike are scattered across that ordering, the tree needs several stacked splits to reunite them, spending depth and data it would not spend on one-hot indicators.
- Are the gaps between the levels of a 1-5 satisfaction rating equal?There is no reason to assume so. Respondents bunch at the top, and the behavioural difference between 4 and 5 is often larger than between 1 and 2. Feeding the raw 1-5 code to a linear model asserts equal gaps; one-hot indicators let the model learn the real spacing, including a non-monotonic one, if you have the rows to support five separate weights.
T-shirt sizes S, M, L can be numbered 1, 2, 3 because they really do line up. Numbering shirt colours red, green, blue the same way tells the model that green sits between red and blue.
saying these in an interview costs you the question
- Integer-codes every categorical column because it saves memory
- Claims tree models make the encoding choice irrelevant
- Treats one-hot and ordinal encoding as two names for the same thing
- Assumes alphabetical code order encodes something meaningful
- Says a nominal integer code is fine because the model will figure it out