Why does adding 5 to every logit of a softmax classifier leave its probabilities unchanged?
answer
- the constant factors out of the ratio
- only score differences are identified
- K vectors, K-1 degrees of freedom
- the likelihood has a flat ridge
- regularisation picks one representative
basics
~20 sSoftmax divides by the sum of the exponentials, so a constant added to every score contributes the same factor to numerator and denominator and cancels. Only differences between class scores are identified, which leaves softmax over-parameterised.
solid answer
~50 sWrite it out: `exp(z_k + c) / sum_j exp(z_j + c) = (exp(c) * exp(z_k)) / (exp(c) * sum_j exp(z_j))`, and the `exp(c)` cancels. Softmax is shift-invariant, so with K classes the model has K weight vectors but only K-1 identifiable degrees of freedom — you can add any constant vector to all K weight vectors and get a model that predicts identically. Three consequences matter. The unregularised likelihood has a flat ridge, so individual weights are not interpretable in isolation and will not be stable across refits; an L2 penalty removes the ambiguity by selecting the minimum-norm solution among the equivalent ones. Some formulations instead pin one class as a reference with weights fixed at zero, which is why the two-class case needs only one weight vector. And the identity is exploited numerically: subtract the largest logit before exponentiating to avoid overflow, with provably identical output.
code
python · 14 linesimport math
def softmax(scores):
m = max(scores) # shift-invariance lets us do this
exps = [math.exp(s - m) for s in scores]
total = sum(exps)
return [e / total for e in exps]
base = [2.0, 1.0, 0.1]
shifted = [s + 5 for s in base]
print([round(p, 6) for p in softmax(base)])
print([round(p, 6) for p in softmax(shifted)])
print(round(sum(softmax(base)), 12))go deeper
Be able to do the one-line cancellation on paper: the same exponential factor appears above and below the fraction line, so it divides out.
Explain that the cancellation is a statement about the parameterisation, not just one evaluation: K weight vectors carry only K-1 worth of information, so the fitted weights are not unique.
Show you have hit this in practice — coefficients that move between refits with unchanged predictions — and know the fixes: regularise, pin a reference class, and report contrasts rather than raw per-class weights.
Own what the team is allowed to claim from model coefficients at all, and set the convention (reference class or centred weights) so that interpretation, monitoring and version-to-version comparison stay meaningful.
## The algebra Softmax over K logits is `p_k = exp(z_k) / sum_j exp(z_j)`. Add the same constant `c` to every logit: ``` exp(z_k + c) / sum_j exp(z_j + c) = (exp(c) * exp(z_k)) / (exp(c) * sum_j exp(z_j)) = exp(z_k) / sum_j exp(z_j) ``` The `exp(c)` factors out of the numerator and out of every term of the denominator, so it cancels exactly. Adding 5 to every logit changes nothing. Only the **differences** `z_k - z_j` determine the output. ## Why that means the model is over-parameterised The logits come from `z_k = w_k . x + b_k`. Pick any fixed vector `v` and scalar `d`, and replace every class's parameters with `w_k + v` and `b_k + d`. Every logit shifts by the same amount `v . x + d` for that input, so by the identity above the probabilities are unchanged — for every input, not just one. The model has K weight vectors and K intercepts, but only **K - 1** of them are identified: there is an entire family of parameter settings that produce byte-identical predictions. Statisticians call this non-identifiability. The log-likelihood surface is not a bowl with a single lowest point; it has a **flat ridge** along the direction of the common shift. Optimisation still converges — it just lands somewhere on that ridge, and where it lands depends on the initialisation, the solver and the stopping rule. ## Consequence 1: the individual weights are not interpretable If you fit the same 8-queue ticket router twice from different starting points, class 3's weight on a given feature can come out as `0.4` in one run and `-1.1` in the other, with identical predictions from both. Only contrasts are meaningful: `w_k - w_j` tells you how much a feature pushes class k relative to class j. A statement like "the weight for the billing queue on this feature is positive, so the feature indicates billing" is not well-founded on its own; the comparable statement is "the feature raises billing relative to shipping". This also means you cannot compare a coefficient across two refits, or across two model versions, without pinning down the parameterisation first. ## Consequence 2: regularisation makes the solution unique Add an L2 penalty `lambda * sum_k ||w_k||^2` to the loss. All the settings on the ridge give the same likelihood, so the penalty picks out the one with the smallest norm. For each feature, shifting all K weights by a common amount is a one-dimensional problem, and the norm is minimised when the shift makes them **sum to zero** across the classes. So the L2-regularised fit has, per feature, class weights centred on zero — the ridge collapses to a single point and refits become reproducible. This is one of the underrated benefits of regularisation here: it is not only shrinking, it is choosing a representative from an equivalence class. An alternative fix is a **reference class**: fix `w_K = 0` and `b_K = 0` and fit only the other K-1. That is the classical statistical parameterisation, and it is exactly why binary logistic regression carries a single weight vector rather than two — the second class is the pinned reference, and `sigmoid(z)` is `softmax` with one logit held at zero. Both fixes remove the same redundancy; they just choose different representatives, so their coefficients are not comparable to each other. ## Consequence 3: numerical stability Logits are unbounded. If a well-separated model produces a logit of 1000, `exp(1000)` overflows to infinity and the ratio becomes meaningless. Because a common shift changes nothing, implementations subtract the largest logit first: ``` m = max(z) p_k = exp(z_k - m) / sum_j exp(z_j - m) ``` Now the largest exponent is `exp(0) = 1`, everything else is between 0 and 1, and nothing overflows. Very negative shifted logits underflow harmlessly to 0. The result is mathematically identical to the naive formula — shift invariance is what licenses the trick. The same reasoning underlies computing the loss in log space instead of exponentiating and then taking a log. ## What an interviewer is really probing This question separates people who memorised the formula from people who have thought about it. The strong answer is three sentences of algebra followed by the three consequences: coefficients are only meaningful as contrasts, regularisation or a reference class picks a unique representative, and the max-subtraction trick is a free consequence rather than an approximation. A weak answer stops at "it cancels" and misses that the cancellation is a statement about the model's parameterisation, not just about one evaluation of the function.
- Given that redundancy, why do implementations still keep K weight vectors instead of K-1?Symmetry and convenience. Keeping all K vectors makes the gradient uniform across classes, makes the code identical for every class, and pairs naturally with an L2 penalty that resolves the ambiguity anyway. Pinning a reference class saves parameters but makes every coefficient a contrast against whichever class you pinned, which is awkward when the class set changes.
- How would you explain to a stakeholder why a coefficient changed sign between two refits with identical accuracy?Because softmax identifies only the differences between class scores, two fits can sit at different points on a flat ridge of equally good solutions and still predict identically. The fix is to report contrasts between classes rather than raw per-class coefficients, and to regularise so the fit lands at the same representative every time.
- Does shift invariance also mean scaling all logits by a constant is harmless?No, and that is the key contrast. Multiplying every logit by a constant changes the ratios of the exponentials, so it sharpens or flattens the distribution while leaving the argmax alone. Only additive shifts cancel; multiplicative rescaling is a real change to the predicted probabilities.
It is like measuring altitude: only differences between two points are meaningful, so moving the sea-level reference up by five metres changes every reading and no conclusion.
saying these in an interview costs you the question
- Says the shift cancels only when the constant is small
- Claims scaling all logits is equally harmless
- Reads a single class weight as a standalone effect
- Thinks subtracting the max is an approximation
- Cannot say why binary logistic regression needs one weight vector