What does the softmax function do to the class scores of a multiclass linear model?
answer
- turns K scores into one distribution
- everything positive, total fixed at one
- the denominator is shared by all classes
- exponentiate, then divide by the sum
- ranking unchanged; only confidence changes
basics
~20 sSoftmax exponentiates each of the K class scores and divides by their total, turning them into K positive numbers that add up to one. The largest score keeps the largest probability, so the predicted class is unchanged.
solid answer
~50 sA multiclass linear model keeps one weight vector per class, so an input produces K raw scores (logits), `z_k = w_k . x + b_k`. Softmax maps them to probabilities: `p_k = exp(z_k) / sum_j exp(z_j)`. Exponentiating makes every value positive, and dividing by the shared sum forces the K values to add to exactly one, which is what lets you read them as a distribution over mutually exclusive classes. Routing a support ticket to one of 8 queues, softmax gives 8 numbers like 0.61, 0.22, 0.09 and so on, summing to 1 — so you can act on the top class or hold the ticket for review when the top probability is low. Softmax is monotone in the scores, so it never changes which class wins; it changes only how confident the model looks. Training uses multiclass cross-entropy, not squared error.
go deeper
Be ready to write the formula on a whiteboard and say in one sentence what the exponential and the denominator each do. Know that the outputs sum to one and that the top class does not change.
Explain that only the differences between logits matter, that the exponential exaggerates score gaps into confidence gaps, and that the two-class case collapses to the sigmoid of the score difference.
Show you know what the sum-to-one constraint costs you in production: the model cannot say "none of these", so out-of-scope traffic arrives looking confident and needs its own detection path.
Own the decision of whether a mutually-exclusive distribution is the right output contract at all, and what downstream consumers are allowed to assume from a probability that is only ever as good as the calibration you can demonstrate.
## The shape of a multiclass linear model A binary linear classifier holds one weight vector: a single score `z = w . x + b` decides the answer. A multiclass model with K classes holds **K weight vectors**, one per class. For an input `x` it produces K raw scores, usually called **logits**: ``` z_k = w_k . x + b_k for k = 1..K ``` These logits are unbounded real numbers. They may be negative, they may all be large, and they carry no meaning as probabilities. Softmax is the function that turns them into one. ## What softmax computes ``` p_k = exp(z_k) / (exp(z_1) + exp(z_2) + ... + exp(z_K)) ``` Two things happen. **Exponentiating** each logit makes it strictly positive, so no class can end up with a negative probability. **Dividing by the sum of all the exponentials** — the shared normaliser — forces the K outputs to add up to exactly one. That second step is the important one. It couples the classes: raising one logit lowers every other class's probability, because the denominator grows. Softmax does not score classes independently and then tidy up; it produces a single distribution in which the classes compete for a fixed budget of probability mass. Worked example. Suppose an 8-queue support-ticket router produces logits `3.0, 1.5, 1.0` for the top three queues and clearly lower values for the rest. Exponentiating gives `20.1, 4.48, 2.72`. If the remaining five exponentials total `3.0`, the grand total is about `30.3`, so the probabilities are roughly `0.66, 0.15, 0.09`, with the rest sharing the last `0.10`. They sum to one by construction. ## Properties worth being able to state **It is monotone.** The class with the largest logit always gets the largest probability, because `exp` is increasing and every class is divided by the same number. Softmax therefore never changes the argmax; it only converts a ranking into calibrated-looking magnitudes. If all you need is the predicted label, you can skip softmax entirely and take the largest logit. **Differences, not levels, matter.** Because everything is divided by a common sum, only the gaps between logits determine the probabilities. Logits `1, 0, 0` and `101, 100, 100` give exactly the same distribution. **It exaggerates gaps.** Because the transform is exponential, a logit gap of 4 turns into a probability ratio of `exp(4)` ≈ 55 to 1. Small changes in score become large changes in stated confidence, which is why a softmax model can look very confident on inputs it has no business being confident about. **Binary logistic regression is the K = 2 case.** With two classes, `p_1 = exp(z_1) / (exp(z_1) + exp(z_2)) = 1 / (1 + exp(-(z_1 - z_2)))`, which is the sigmoid applied to the score difference. Softmax is the K-class generalisation of the sigmoid, and multinomial logistic regression is the K-class generalisation of logistic regression. ## Why probabilities that sum to one matter operationally Because the outputs form a distribution over classes assumed **mutually exclusive**, you can do arithmetic with them. You can set an abstain rule ("if the top probability is below 0.5, send the ticket to a human queue"), aggregate over classes ("probability the ticket is any billing-related queue = sum of those queues' probabilities"), or feed the probabilities into an expected-cost calculation where different misroutings cost different amounts. None of that is safe with scores that are merely ordered. The flip side: because the K probabilities are forced to sum to one, softmax can never say "none of these". An input from an unseen class still produces a confident-looking distribution. Out-of-scope detection has to come from somewhere else — a low top probability, a high entropy across the K outputs, or a separate model. ## Training The model is fit by minimising **multiclass cross-entropy**: for each row, the negative log of the probability assigned to the true class. Under that loss, the gradient with respect to class k's weights is proportional to `(p_k - y_k) * x`, where `y_k` is 1 for the true class and 0 otherwise — the model is pushed to raise the true class's probability and lower every other class's. Squared error is not used here: it is not the negative log-likelihood of the model, and it produces vanishing gradients exactly where the model is confidently wrong. ## What interviewers are checking Mostly that you can write the formula, say why the exponential and the normaliser are there, and state that the outputs are a distribution over mutually exclusive classes rather than K independent confidences. The common failure is describing softmax as "dividing the scores by their sum", which would break on negative scores and is not what the function does.
- If softmax never changes which class has the highest score, why compute it at all?For the label alone you do not need it — the largest logit already wins. You compute softmax when you need calibrated magnitudes: an abstain threshold, a cost-weighted decision, aggregating several classes into a group probability, or a proper scoring rule such as log loss during training and evaluation.
- What happens if you feed softmax an input from a class the model was never trained on?It still returns K probabilities that sum to one, often with one of them large. Softmax has no way to express "none of these" — the normaliser guarantees the mass is spent somewhere. You detect out-of-scope inputs separately, for example by flagging a low top probability or a high-entropy output distribution.
- How does softmax relate to the sigmoid used in binary logistic regression?It is the K-class generalisation. With two classes, `exp(z_1) / (exp(z_1) + exp(z_2))` simplifies to `1 / (1 + exp(-(z_1 - z_2)))`, the sigmoid of the score difference. That is why binary logistic regression needs only one weight vector: only the difference between the two class scores is meaningful.
It is like converting eight raw vote counts into vote shares, except the counts are exponentiated first, so a modest lead in raw score becomes a lopsided share.
saying these in an interview costs you the question
- Says softmax divides the raw scores by their sum
- Claims softmax can change which class is predicted
- Treats the K outputs as independent per-class confidences
- Thinks a high softmax probability proves the input is in-scope
- Says the model is trained with squared error on the probabilities