skip to content

Why can a trained classifier output a 99.9% softmax score on pure noise?

level: middleimportance: must knowfreq 62%

answer

  1. closed-world question, always answered
  2. the outputs must sum to one
  3. only logit differences survive normalization
  4. the loss keeps pushing gaps wider
  5. piecewise-linear extrapolation grows the gap

basics

~20 s

Softmax normalizes scores across the trained classes, so its output ranks them and must sum to one — never evidence that the input belongs to any. Nothing in training penalised confident nonsense, and far from the data the gaps can grow.

solid answer

~50 s

A softmax answers a closed-world question: of these K classes, which fits best? It exponentiates the K output scores and normalizes them, so one unit of probability is always handed out no matter what the input was — noise included. Cross-entropy training then rewards ever-larger gaps between the correct score and the rest, since the loss keeps falling as the target class approaches one, and the training set never contains an example labelled `none of these`. Worse, for networks built from piecewise-linear units, moving far from the training data along a direction tends to scale the scores rather than flatten them, so confidence can *rise* off-distribution. The consequence is practical: a high softmax value is a rank among the known classes, not a probability of being right, and it must never be used on its own as an out-of-distribution detector.

go deeper

for a junior

Recall that softmax always sums to one across the trained classes, so a high score means best of these options, not this input is one of these options. Being able to say that clearly is enough at this level.

for a middle

Explain the mechanics: shift invariance, the exponential saturation of the score in the logit gap, and why cross-entropy keeps widening that gap because its output gradient never reaches zero for the correct class.

for a senior

Show the operational habit. Say that you never gate on the top score alone, that you pair it with a disagreement or novelty signal, and describe how you would demonstrate the failure to a sceptical stakeholder with a rotated or noise input.

for a principal

Own the risk framing. Decide when a closed-world classifier is acceptable at all, when the product needs an explicit reject path or an open-set formulation, and what the organisation commits to when a confident wrong answer reaches a user.

## What softmax actually computes Given K real-valued outputs (logits) `z_1 ... z_K`, softmax returns `p_i = exp(z_i) / sum_j exp(z_j)`. Three properties matter here. 1. **It is shift-invariant.** Adding the same constant to every logit leaves the output unchanged. Only *differences* between logits carry information, so the absolute magnitude of the network's evidence is thrown away. 2. **It is normalized by construction.** The outputs sum to one for every possible input, including inputs that resemble no class at all. There is no spare mass for *none of the above*; whatever mass exists must be distributed among the K trained classes. 3. **It saturates.** A logit gap of about 7 already gives a top score above 0.999. Confidence is an exponential function of a gap that the network is free to grow without bound. Put together: softmax reports a *relative ranking* under the assumption that the input is one of the K classes. That assumption is baked in and never tested. ## Why training makes it worse, not better Cross-entropy with a one-hot target has gradient `dL/dz = p - y` at the logits. For the correct class this is `p_c - 1`, which is negative until `p_c` reaches exactly one — it shrinks as the model gets confident, but it never becomes zero. So on every example the model is still nudged, however slightly, to push the correct logit further above the rest. Anything that lets it do so cheaply — weights growing, a duplicated easy example, a long training run — increases the typical gap. The result is a network whose gaps are systematically larger than the true evidence warrants, which is the ordinary overconfidence seen even *inside* the training distribution. Meanwhile, the training set contains no examples of nonsense. The objective never once asked the model what to do with noise, and a loss can only shape behaviour on inputs it has seen. So the behaviour off-distribution is whatever the architecture happens to extrapolate. ## Why the extrapolation goes the wrong way A network of piecewise-linear units (a rectifier network) partitions input space into regions, and inside each region it is an affine function. Far away from the data you are inside some outer region, and the function there is affine and unbounded: scale the input far enough along a direction and the logits scale with it. The largest logit typically grows fastest, so the *gap* grows and the softmax score approaches one. This is the formal version of the demo everyone reproduces: feed pure noise, or an image rotated far past anything in training, and get 99.9% for a specific class. There is nothing pathological about the trained network — the geometry says confidence does not decay with distance from the data, it often grows. ## What the number is and is not - **Is:** a monotone score usable for ranking within the known classes, and a usable input to a decision threshold *if* it is checked against held-out data from the same distribution. - **Is not:** a probability that the prediction is correct; a measure of how familiar the input was; an out-of-distribution detector; a decomposition into noise versus ignorance. A fourth trap: because softmax is shift-invariant, two inputs with wildly different amounts of evidence — one strongly matching a class, one matching nothing — can produce identical outputs if their logit *differences* happen to match. The evidence magnitude that would have distinguished them is discarded by normalization. ## What to do about it The honest fixes all add information the single forward pass does not have. - **Model disagreement.** Sample several plausible models — independently trained networks, or stochastic forward passes — and read their spread. Where the data was thin, they disagree, and that disagreement is the missing epistemic signal. - **An explicit novelty score.** Score how typical the input's internal representation is relative to the training data — a distance in feature space, a density estimate, or the score of a model trained to recognise the training distribution itself — and treat it as a separate gate from the class prediction. - **A reject option.** Let the system decline to answer and route the case to a human or a fallback, based on that gate rather than on the top score alone. - **Training-time changes.** Label smoothing (targets of `1 - e` and `e/(K-1)` instead of one-hot) removes the incentive to grow gaps without bound; exposing the model to auxiliary out-of-distribution examples with a flat target teaches it that some inputs deserve a flat answer. ## The interview point Candidates who have only used models say "softmax gives probabilities". Candidates who have shipped them say "softmax gives a normalized ranking over the classes I chose to train on, whose magnitude is an artefact of how long I trained, and I never let it decide whether an input is in-distribution". The second answer is what the question is testing for.

  • Two inputs give identical softmax vectors but one clearly matched a class strongly. What did the softmax discard?
    The magnitude of the evidence. Softmax is shift-invariant, so adding a constant to every logit changes nothing; only the differences survive. An input producing logits of 12 and 5 and one producing 2 and -5 give the same output. Any notion of `how much did anything fire at all` is normalized away, which is exactly the notion novelty detection needs.
  • Does label smoothing fix overconfidence on out-of-distribution input?
    Only partly. Smoothing replaces the one-hot target with a slightly flattened one, so the loss stops rewarding unbounded logit gaps and in-distribution scores become less extreme. But it says nothing about inputs the model never saw: the extrapolation geometry off-distribution is unchanged, and a smoothed model can still be very confident on noise. It moderates the symptom inside the data, not the blind spot outside it.
  • Why is thresholding the top softmax score still a common production baseline despite all this?
    Because it is free, monotone, and often good enough when inputs really are in-distribution — it does separate easy cases from ambiguous ones within the trained classes. The failure is specific: it cannot detect that an input is outside the world it was trained on. Teams use it as a first-pass filter and pair it with a separate novelty or disagreement signal for that job.

A multiple-choice exam with no none of the above option forces a letter on every question, including one printed in a language the student cannot read.

saying these in an interview costs you the question

  • States that softmax outputs are true posterior probabilities
  • Uses the top score as an out-of-distribution detector
  • Thinks a bigger training set alone fixes off-distribution confidence
  • Believes softmax flattens automatically for unfamiliar input
  • Confuses the size of a logit with the size of the softmax output

context