skip to content

A deployed 20-way product classifier needs a 21st category. What happens to the softmax head?

level: seniorimportance: should knowfreq 58%

answer

  1. head shape follows the label set
  2. one new row, one new bias
  3. everything renormalizes across 21
  4. old labels may hide the new class
  5. old thresholds and metrics no longer apply

basics

~20 s

The head gains one weight row and one bias for the new class; the trunk keeps its shape. All 21 classes then share one probability budget, so old outputs shift and old thresholds no longer transfer.

solid answer

~40 s

The head is a `20 x d` weight matrix plus 20 biases; a 21st class means one new row of `d` weights and one new bias. Everything below the head keeps its shape, and the 20 existing rows are a legitimate warm start. But you cannot ship it unretrained: the new row has learned nothing, and because all 21 outputs share a single budget of one, every prediction renormalizes. The real work is data. You need examples of the new category, and — the step people forget — you must find examples currently labelled among the old 20 that actually belong to category 21, or you train the old rows to keep claiming them. Then decide whether head-only fine-tuning suffices, and re-derive thresholds and per-class metrics afterwards.

go deeper

for a junior

Know that the head's output count equals the class count, so a new category means the head must grow and the model must be retrained. Recognize that the old model cannot predict a class it has no unit for.

for a middle

Explain the shapes: one new weight row of length d plus one bias, trunk unchanged, old rows reusable as a warm start. Say why the shared probability budget means every prediction shifts once the class is added.

for a senior

Show the operational sequence: relabel historical data that belongs to the new class, choose head-only versus full fine-tuning based on whether the trunk already separates it, then re-derive thresholds and watch per-class recall on the old classes for cannibalization.

for a principal

Own the cost curve. If the taxonomy grows on a schedule, argue explicitly about whether every addition should trigger a retrain-and-relabel cycle, what that costs per quarter, and at what point a design where classes are added as examples rather than parameters is worth its own migration.

## What changes structurally The classification head over a `d`-dimensional final representation is a weight matrix of shape `20 x d` plus a bias vector of length 20 — one weight vector and one bias per class. Adding a 21st category means the matrix becomes `21 x d` and the bias vector length 21. That is `d + 1` new parameters, and nothing else in the network changes shape: the trunk, all its layers and the representation it produces are untouched. That structural minimalism is why people underestimate the change. The parameter delta is tiny; the behavioural delta is not. ## Why the head cannot be reused unchanged With 20 outputs, category 21 is not merely unlikely — it is **unrepresentable**. Every example of the new category must be assigned to one of the 20 existing classes, and the model will do so confidently, because it has no vocabulary for "none of these". A shipped 20-way head therefore does not degrade gracefully as the new category's traffic grows; it silently pollutes the metrics of whichever old classes are nearest in feature space. Once the 21st row exists, the softmax's fixed budget bites. All 21 outputs must sum to one, so probability mass that used to be split among 20 classes is now split among 21. Even for an input that has nothing to do with the new category, the old probabilities generally drop a little. Any rule of the form "act when the top class exceeds `t`" was tuned against the old normalization and now means something slightly different. ## What survives, and how to initialize - **The trunk survives.** If the new category is visually or semantically similar to what the model already sees, its features are probably already discriminative enough. - **The 20 old rows survive as a warm start.** Keeping them is far better than random re-initialization, but expect them to move: every training example of category 21 now pushes the old rows' scores down on those inputs. - **The new row starts untrained.** Initializing it to zeros gives it a score of exactly zero for every input, so before any training it takes a non-trivial slice of mass on inputs where the old scores happen to be small — which is not harmless, just predictable. Small random values behave similarly. Either way the row is meaningless until it sees data. Freezing the old 20 rows and training only the new one is a tempting shortcut. It is usually wrong: the frozen rows cannot learn to *stop* claiming examples that now belong to category 21, so the new class has to out-shout entrenched competitors that never yield. ## The data work is the real work Two distinct data tasks: 1. **Collect labelled examples of category 21.** Obvious, and usually the smaller half. 2. **Relabel the historical data.** Some examples currently labelled with one of the old 20 genuinely belong to the new category — they had nowhere else to go. If you leave them, you are training the old rows and the new row on contradictory targets for near-identical inputs, and the new class will underperform for reasons that look like a modelling problem but are a labelling problem. Auditing the classes nearest to the newcomer, and the low-confidence tail of the old model's predictions, is a productive way to find them. Class balance matters too: a brand-new category typically has far fewer examples than the mature 20, so the head sees it rarely and its row stays weak. ## Retraining choices - **Head-only fine-tuning.** Cheapest and fastest. Appropriate when the trunk's representation already separates the new category from its neighbours — you are only learning a new direction in an existing feature space. - **Fine-tuning the whole network.** Needed when the new category requires distinctions the trunk never learned to make, because no linear row over the current features can isolate it. - **Full retraining from scratch.** Rarely justified by one added class alone; it is a reasonable choice when the taxonomy is being restructured rather than extended. ## Evaluation is not a drop-in comparison You cannot compare the old model's 20-way accuracy against the new model's 21-way accuracy — the label space is different and the second task is harder by construction. Build a common, relabelled test set and report: - per-class precision and recall for the original 20, specifically to detect **cannibalization**, where the new row steals examples that legitimately belonged to an old class; - the new class's own precision and recall, which will lag until it has comparable data; - re-derived operating thresholds, since the old ones were calibrated against a 20-way normalization. ## The design question underneath If the taxonomy is going to keep growing — new product categories every quarter — a fixed-width head means a retraining event every time, plus a relabelling pass and a threshold re-derivation. At that point the honest senior answer is to question the design: an approach where a class is represented by a prototype or reference embedding rather than by a parameter row lets a new category be added by supplying examples rather than by retraining a head. That is a bigger change with its own costs, and the right time to argue for it is when the second or third category request arrives, not the tenth.

  • Can you initialize the new row to zeros and freeze the other 20?
    You can, but it usually underperforms. The frozen rows cannot learn to stop claiming examples that now belong to the new category, so the new row must beat entrenched competitors that never yield ground. The frozen classes' probabilities also drop anyway, because all 21 outputs share one budget. Fine-tuning all rows of the head is the cheap default.
  • What if the 21st category is a catch-all `other` rather than a real product type?
    The mechanics are identical — one row, one bias — but the modelling is harder. A catch-all is heterogeneous, so a single weight vector has to cover a scattered region of feature space, and its examples are defined by exclusion rather than by shared structure. Expect weak recall, and expect the boundary to drift every time the real taxonomy changes.
  • How would you evaluate the 21-way model against the 20-way one it replaces?
    Not by comparing headline accuracy — the label spaces differ and the new task is harder. Use one relabelled test set, report per-class precision and recall for the original 20 to expose classes cannibalized by the newcomer, report the new class separately, and re-derive any operating thresholds against the new normalization before comparing decisions.

saying these in an interview costs you the question

  • Says you can add a class without touching the head
  • Assumes the original 20 probabilities are unaffected
  • Claims head-only fine-tuning always suffices
  • Forgets to relabel old data that is really the new class
  • Compares 20-way and 21-way accuracy directly

context