Why must a softmax classification head have exactly one output unit per class?
answer
- width is a modelling decision, not tuned
- one score per candidate answer
- normalized over the units present
- a missing unit is an impossible prediction
- K rows of weights, K biases
basics
~20 sA softmax head emits one score per class and normalizes those scores into probabilities that sum to one. K classes need K units: with fewer, some class can never be predicted; extra units create classes no label ever selects.
solid answer
~50 sThe head is a linear layer mapping the final hidden vector `h` (dimension `d`) to `K` scores, so its weights are `K x d` plus `K` biases: class `k` owns its own weight vector, `s_k = w_k . h + b_k`. Softmax turns those into `p_k = exp(s_k) / sum_j exp(s_j)`, a distribution over exactly the units present. The width is therefore fixed by the label space, not tuned: a 40-way language identifier has 40 units, a 1,000-way object head has 1,000. Drop a unit and that class is unrepresentable — the model must misfile it. Add spare units and they still consume probability mass and can win an argmax on odd inputs. Because the outputs sum to one, the classes compete: one probability rises only if others fall, which is exactly the mutual-exclusivity assumption the head encodes.
go deeper
Be ready to state the head's width from the task alone: one unit per class, no exceptions. Know that the outputs sum to one and that the prediction is the largest one.
Explain the mechanics: a K x d weight matrix plus K biases, one weight vector per class, and why encoding the class as a single number or as bits breaks. Say why one-hot rows and integer ids are the same target.
Show you treat the head width as a contract with the label space. Interviewers want to hear what happens operationally when a class is missing from the head — the traffic does not vanish, it gets misfiled into neighbours, and your per-class recall hides it.
Own the tradeoff when the taxonomy is unstable. A fixed-width head ties the model's parameters to today's label set, so argue when to accept that versus a design where a new class is a new prototype rather than a new parameter row, and who pays the retraining cost each time the taxonomy moves.
## What the head actually is A classification network ends with a **head**: a linear layer that takes the last hidden representation `h` — a vector of length `d` produced by whatever trunk sits underneath — and produces one raw score per class. Concretely the head holds a weight matrix of shape `K x d` and a bias vector of length `K`, where `K` is the number of classes. Class `k` owns row `k`: ``` s_k = w_k . h + b_k for k = 1..K ``` Those `K` numbers are the **logits**. Softmax then maps them to a probability distribution: ``` p_k = exp(s_k) / (exp(s_1) + ... + exp(s_K)) ``` Every `p_k` is strictly positive and the `K` of them sum to `1`. ## Why one unit per class, and not fewer Each class needs its **own** weight vector — its own direction in feature space that says "this input looks like class `k`". That is the whole content of the head. If a class has no row, no input can ever score highest for it, so the class is not merely unlikely, it is **unrepresentable**: every example of it must be filed under one of the classes you did keep. This is the single most important practical consequence of head width. Two tempting shortcuts both fail: - **One unit predicting the class index.** Regressing to the integer id imposes an ordering and a metric that the label set does not have: it makes class 3 sit "between" class 2 and class 4, so an error toward a neighbouring id looks cheaper than an error toward a distant one, and there is no way to express "probably 7, possibly 2". - **Binary-encoding `K` classes into about `log2(K)` units.** The head would emit a code rather than a distribution: nothing constrains it to a valid code, and flipping one output bit lands on an arbitrary unrelated class rather than a near miss. The `K`-unit design avoids both. It is redundant in a precise sense — the distribution is unchanged if you add the same constant to every logit — but that redundancy costs one extra weight vector and buys a clean, symmetric parameterization. ## Why not extra units Spare units are not free. They are never the true class, so training pushes their logits down, but never to minus infinity; they keep taking a slice of the probability budget, and on an unusual input a stale, undertrained row can score highest and win the argmax. They also cost `d` weights each. Width is a modelling decision determined by the taxonomy, not a hyperparameter to sweep. Note that head width has nothing to do with the width of the layer feeding it. A 7-class head fed by a 512-unit hidden layer is `7 x 512` weights plus 7 biases: `d` sets how many weights each class gets, `K` sets how many classes there are. ## Targets: one-hot rows and integer ids are the same thing The head's output is a distribution over `K` classes, so the target is naturally a distribution too: a **one-hot** row of length `K` with a `1` at the true class and `0` everywhere else. In practice labels are usually stored as a single **integer class id** instead. These are not two different setups — the id is simply the index of the `1` in the one-hot row. The training objective reads off the predicted probability of the true class either way, so the model, the head width and the gradients are identical; ids are just far cheaper to store and pass around when `K` is large. If an interviewer asks "are your labels one-hot or integers?", the correct answer is that it is a storage question, not a modelling one. ## The fixed budget: classes compete Because the `K` outputs must sum to `1`, they share a fixed budget. A class's probability can only rise if others fall. That coupling is the point: it encodes the assumption that **exactly one** of the `K` classes is correct for any input. A head over a label space where two labels can be simultaneously true is the wrong design, because you cannot raise both without robbing the rest. The budget also explains a behaviour beginners misread as a bug. Feed a 40-way spoken-language head a two-second clip that genuinely could be Portuguese or Spanish and you may see roughly `0.45` and `0.44` on those two units with the remaining `0.11` spread thinly over 38 others. Nothing has failed. The head is reporting that it cannot separate the two candidates, and the only way it could report `0.9` for one of them is to take that mass from the other. ## At inference Prediction is the `argmax` over the `K` units — the index of the largest score. The probabilities are what you look at when you need to know *how* strongly the head prefers that class, or how the runner-up compares.
- Your labels are stored as integer class ids rather than one-hot rows. Does the head change?No. The integer id is just the index of the `1` in the one-hot row, so it names the same target. Head width is still the class count, the weights are the same shape, and training sees the same signal. Ids are only a cheaper encoding, which matters when the class count is in the thousands.
- What actually goes wrong if you give a 10-class problem a 12-unit head?The two spare units are never the true class, so training drives their scores down but not to negative infinity. They keep consuming part of the probability budget, they waste one weight vector each, and on an unusual input an undertrained spare row can score highest and be returned as the prediction. Nothing crashes; the failure is silent and occasional.
- Does head width depend on how many hidden units feed it?No. Width is the class count. The incoming dimension `d` only sets how many weights each class gets: a `K`-class head over a `d`-dimensional feature vector is `K x d` weights plus `K` biases. Widening the trunk changes the parameter count, never the number of outputs.
It is a show of hands where every candidate must have a hand to raise. Leave one candidate out of the room and nobody can vote for them, however obviously they should have won.
saying these in an interview costs you the question
- Calls the output width a hyperparameter to tune
- Suggests one output unit regressing to the class index
- Says the softmax outputs are independent per-class confidences
- Confuses head width with the hidden layer's width
- Thinks a missing class merely gets low probability