skip to content

How does label smoothing change the target that a classifier's cross-entropy loss fits?

level: middleimportance: must knowfreq 58%

answer

  1. the target stops being one-hot
  2. a little mass on every class
  3. blend of one-hot and uniform
  4. the optimal logit gap becomes finite
  5. ten classes at 0.1 gives 0.91

basics

~10 s

Label smoothing replaces the one-hot target with a blend of one-hot and uniform: the true class is asked for about 1 minus epsilon, every other class for epsilon divided by the number of classes.

solid answer

~50 s

Normally the loss compares the softmax output against a one-hot target: 1.0 on the labelled class, 0 everywhere else. Label smoothing mixes that target with a uniform distribution, so with `K` classes and smoothing strength `epsilon` the true class gets `1 - epsilon + epsilon/K` and each of the others gets `epsilon/K`. At `K = 10` and `epsilon = 0.1` that is 0.91 for the label and 0.01 for each of the nine others. The consequence matters more than the formula: cross-entropy against a fixed target is minimised when the softmax equals that target, so the model now has a finite optimum. The correct-versus-other logit gap it wants to reach is `log(0.91/0.01)`, about 4.5, instead of growing without bound as it does under a one-hot target. That bounded gap is the whole regularizing effect — it stops the network from paying for ever more extreme confidence on data it already gets right.

code

python · 16 lines
python
import math

K, eps = 10, 0.1
uniform = eps / K                      # 0.01 on every class
target = [uniform] * K
target[3] += 1.0 - eps                 # class 3 is the observed label
print([round(t, 3) for t in target])   # 0.01 ... 0.91 ... 0.01

# Cross-entropy against a fixed target is minimised when softmax == target,
# so the correct-vs-other logit gap the model aims for is finite:
print(round(math.log(target[3] / uniform), 2))   # 4.51  == log(91)

# A one-hot target has no such ceiling: the gap just keeps growing.
for p_correct in (0.99, 0.999, 0.9999):
    other = (1 - p_correct) / (K - 1)
    print(p_correct, round(math.log(p_correct / other), 2))

go deeper

for a junior

Recall the shape of the change: the label vector goes from 1 and 0s to something like 0.91 and 0.01s, it is applied only during training, and its purpose is to stop the model from becoming too sure.

for a middle

Be ready to write the smoothed target for a stated K and epsilon and explain why the loss now has a finite optimum, including the logit-gap number that follows from it. Name which convention you are using.

for a senior

Show that you know what the change ripples into: probability scales shift so downstream thresholds need refitting, held-out likelihood gets worse by design, and epsilon is a dial you tune against a metric rather than a default you inherit.

for a principal

Own the framing that smoothing is an objective change, not a knob: it trades likelihood quality and representation structure for confidence control and a small accuracy gain, and you should be able to say which of those your product actually pays for.

## The target a one-hot loss fits A `K`-class classifier ends in `K` logits that a softmax turns into a probability vector `p`. Training minimises the cross-entropy between `p` and a target distribution `q`, and with an ordinary hard label `q` is one-hot: 1.0 on the observed class `y`, 0.0 on the other `K - 1`. That target is never actually reachable. A softmax output is a ratio of positive exponentials, so no finite set of logits produces exactly 1.0 and exactly 0.0. On training data the network can separate, the loss therefore keeps falling for as long as you keep pushing `z_y` further above the other logits — there is no finite minimum, only a direction of continuing improvement. The visible symptoms are familiar: training confidence saturates near 1.0, the last layer's weight norm keeps creeping up, and the model reports 0.99 on examples it has no business being sure about. ## What smoothing changes Label smoothing changes the target, not the architecture, not the optimizer, and not what happens at inference. The smoothed target is a mixture of the one-hot vector and the uniform distribution: `q' = (1 - epsilon) * onehot(y) + epsilon * uniform` Component-wise that is `q'_y = 1 - epsilon + epsilon/K` for the labelled class and `q'_k = epsilon/K` for every other class. With `K = 10` and `epsilon = 0.1`, the uniform component contributes 0.01 to all ten entries, so the target vector reads 0.91 on the label and 0.01 on each of the nine others. Be aware of a second convention in common use: the true class gets `1 - epsilon` and the remaining `epsilon` is split over the other `K - 1` classes, giving 0.90 and 0.0111 at the same settings. The two are numerically close and behave the same way, but they are not identical, so state which one you mean when you quote a number. ## Why the logit gap becomes finite Cross-entropy against a fixed target `q'` is minimised, over all valid probability vectors, exactly when `p = q'`. Softmax depends only on logit differences, and `p_y / p_k = exp(z_y - z_k)`. Setting `p = q'` therefore pins the gap the model is trying to reach: `z_y - z_k = log( (1 - epsilon + epsilon/K) / (epsilon/K) )` At `K = 10, epsilon = 0.1` that is `log(91)`, roughly 4.51. Once the network reaches that margin the loss stops rewarding a bigger one; pushing further actually increases the loss. Under a one-hot target the same ratio is `p_y / (1 - p_y)` divided over the wrong classes, which diverges: at 0.99 the gap is about 6.8, at 0.999 about 9.1, at 0.9999 about 11.4, and it never stops. That is the honest one-line description of the method: it converts an objective with an optimum at infinity into one with an optimum at roughly 4.5. It sits alongside penalties that shrink parameters as a way of bounding how extreme the fitted solution can become, but it does it by editing the target rather than by penalising the weights. ## What it buys and what it costs Buys: measurably less extreme confidence, a small but repeatedly reported top-1 accuracy gain on large classification and sequence tasks, and some tolerance of mislabelled examples — a wrong label costs less when the target already concedes probability mass to the other classes. Costs: held-out log-likelihood and per-token perplexity get worse by construction, because the model is being trained never to put full mass on the observed class and those metrics score exactly that. Anything downstream that consumes raw probabilities — a decision threshold, a cascade's abstain rule — has to be re-tuned against the new scale. And because the loss also shapes the hidden representation, the penultimate features change, which matters if some other system reuses them. ## Practical notes - `epsilon` in the 0.05 to 0.2 range covers most reported use; larger values are chosen when labels are known to be unreliable. - Smoothing applies to the training targets only. Evaluation uses the real labels, and inference is unchanged — there is no epsilon at serving time. - The same idea applies to a single-sigmoid binary head: train toward `epsilon` and `1 - epsilon` instead of 0 and 1. - It does not apply to regression targets; there is no probability simplex to soften. - At the optimum it does not change which class wins, because the mass it adds is symmetric across the wrong classes. It still changes the learned function, since it changes every update along the way, so measured accuracy does move.

  • How would you pick epsilon for a crowd-annotated set whose labels are wrong about 8 percent of the time?
    Start near the error rate: under the convention where the label gets `1 - epsilon` and the rest split the remainder, epsilon equal to 0.08 reproduces a uniform 8 percent flip. Then discount it. Noisy hard labels already deliver that soft target in expectation, so stacking epsilon on top compounds — 8 percent noise plus epsilon 0.08 behaves like roughly 0.15. Real annotation error is also concentrated on confusable classes, not uniform, so treat the noise rate as an upper bound and sweep.
  • Does label smoothing change which class the model predicts?
    Not at the optimum. The extra mass is spread symmetrically over the wrong classes, so the ordering of the logits that minimises the smoothed loss is the same ordering that minimises the one-hot loss. It does change the learned function, because it changes every gradient step taken to get there, so measured accuracy shifts — usually slightly upward on large tasks. Any threshold tuned on the old probability scale, though, has to be refit.
  • On separable training data, does one-hot cross-entropy ever reach its minimum?
    No. The loss falls monotonically as the correct-class probability approaches 1, which finite logits never reach, so there is no minimiser — only an escape to infinity. In practice the logits and the final weight norm keep growing as long as you train, which is exactly the overconfidence that smoothing removes by giving the objective a finite optimum at a fixed logit gap.

A one-hot target is a grader who only accepts a perfect score, so you keep studying forever. A smoothed target says 91 out of 100 is full marks, and you stop.

saying these in an interview costs you the question

  • Says the smoothed target puts zero on the wrong classes
  • Claims smoothing flips which class the model predicts
  • Thinks epsilon is applied to the outputs at inference time
  • Says one-hot cross-entropy has a finite minimum on separable data
  • Confuses softening labels with averaging several models' predictions

context