skip to content

Why does smoothing training labels stop a log-loss classifier saturating at 0 and 1?

level: seniorimportance: nice to knowfreq 24%

answer

  1. hard targets have no finite optimum
  2. log loss is minimised at the target
  3. the best log-odds becomes a finite number
  4. eps should match believed labelling error
  5. confidence changes more than accuracy

basics

~20 s

Log loss against a hard 0/1 target keeps falling as the prediction approaches it, so the fit inflates coefficients without limit. Smoothing targets to 0.95 and 0.05 puts the minimum there, capping the log-odds and the weights.

solid answer

~40 s

The per-row log loss `-[t*log(p) + (1-t)*log(1-p)]` is minimised at `p = t`. With hard targets that means `p = 1`, which is only reachable as the log-odds run off to infinity — so on rows the model already separates, the fit keeps growing coefficients to squeeze out a last sliver of loss. That is overfitting to rows it is already right about. Replace the targets with `1-eps` and `eps`, and the best achievable prediction is `1-eps`: the optimal log-odds is the finite number `log((1-eps)/eps)`, and the pressure to inflate weights stops there. Flipping a small random fraction of labels does something similar, more crudely and with more variance. Expect better-calibrated probabilities and smaller coefficients; expect accuracy to be flat or a touch worse. Set eps near the label-error rate you actually believe in.

go deeper

for a junior

Recall that log loss punishes confident mistakes hardest and that hard 0/1 targets push predicted probabilities towards the extremes. Know that softening the targets makes the model less absolute.

for a middle

Explain that log loss against a soft target is minimised when the prediction equals that target, so the optimal log-odds becomes finite and the fit stops inflating coefficients to chase an unreachable 1.

for a senior

Show the diagnosis and the dosing: spot saturation from a score distribution piled at the extremes, tie the smoothing amount to the labelling error rate you believe in, and re-check any tuned decision threshold afterwards.

for a principal

Decide whether the product needs trustworthy probabilities at all, and weigh smoothing everything against investing in label quality — a regularizer that hides annotation error is cheaper than fixing it and worse in the long run.

## Where the saturation comes from A probabilistic binary classifier trained with log loss pays `-log(p)` on a positive row and `-log(1-p)` on a negative one. With hard targets of 1 and 0, that penalty keeps shrinking all the way to `p = 1` and `p = 0`, and those endpoints correspond to log-odds of plus and minus infinity. On any training set the model can separate — which is common with many features and few rows, and in a document-triage set where a handful of tokens give the class away — there is no finite optimum. The fit keeps scaling its coefficients up, because doubling every weight doubles the log-odds and shaves a little more off the loss on rows it already classifies correctly. Nothing is being learned in that phase; the model is simply becoming more strident about what it already believed, and it will be equally strident when it is wrong on a new row. ## What smoothing changes Replace the target for a positive row with `t = 1 - eps` and for a negative row with `t = eps`, for some small eps such as 0.02 or 0.05. Now minimise `-[t*log(p) + (1-t)*log(1-p)]` over p. Differentiating, the minimum sits at `p = t` exactly. The best the model can do on that row is predict `1 - eps`, which corresponds to a log-odds of `log((1-eps)/eps)` — a finite target. Push past it and the loss goes back up. The engine that was driving coefficients upward is switched off at a bounded point. That is the mechanism, and note where it acts. A squared-weight penalty caps the weights by charging for them directly. Smoothing caps them indirectly, by removing the reward for growing them. The two dials often land in a similar place, which is why smoothing belongs in the same family as noise on the inputs and averaging over fits: a regularizer that never appears as a term in the loss. ## Randomly flipping labels The cruder cousin is flipping a small percentage of training labels outright. The expected target for a flipped-with-probability-eps label is the smoothed target, so in expectation you are doing the same thing — but with far more variance, because a specific row is either flipped or not, and a flipped row in a small class does real damage. Smoothing every target deterministically is the lower-variance way to make the same statement, and it is what you should reach for first. ## When smoothing is the honest choice Smoothing is a statement about the labels, so make it when the statement is true. In a document-triage set labelled by human reviewers, some percentage of the labels are simply wrong — reviewers disagree, guidelines shift, ambiguous documents get a coin-flip. With hard targets the fit is required to believe those wrong labels absolutely, and it spends capacity contorting itself to get them right. Smoothing says out loud: this label is *probably* right. Choose eps in the neighbourhood of the error rate you believe in — three percent labelling error suggests an eps in that region, not 0.3. ## What to expect, and what to check Coefficients get smaller and the score distribution stops crowding the two extremes. Probability quality typically improves: the scores mean more as probabilities than they did. Accuracy usually moves very little, and can dip slightly, because the ordering of the scores is driven by the same features either way. Judge the change on a probability-sensitive measure rather than on accuracy alone, and on a held-out set rather than the one you tuned eps on. ## When not to smooth - **When the labels really are exact.** A label read off a completed transaction — the invoice was paid or it was not — carries no annotation error. Smoothing it throws away information for nothing. - **When the product consumes the extreme tail.** If an automated action fires only on the most confident one percent of cases, smoothing compresses precisely the region the product depends on, and you must re-tune the threshold rather than assume the old one still means what it did. - **When the label noise is systematic rather than random.** If one class is mislabelled in a particular, patterned way, smoothing every target uniformly does not address the pattern; it just makes the model quieter about a bias it still has. That case needs a labelling fix, not a regularizer. ## The interview version Say that log loss with hard targets has its optimum at infinite log-odds, so the fit has an unbounded incentive to grow weights; that smoothing moves the optimum to a finite point because the loss is minimised at the target value itself; that the effect is smaller coefficients and better-behaved probabilities rather than higher accuracy; and that eps should be tied to the labelling error rate you actually believe in.

  • How would you choose the smoothing amount?
    Anchor it to the labelling error rate you actually believe in — if roughly three percent of triage labels are wrong, an eps near that is defensible and anything an order of magnitude larger is discarding signal. Then treat it as a hyperparameter and confirm it on a held-out set using a probability-sensitive measure, because accuracy will barely move and will not tell you whether it helped.
  • When would you not smooth at all?
    When the labels carry no annotation error — a label read off a completed transaction rather than a human judgement — and when the product consumes the extreme tail of the score, such as an automated action reserved for the most confident one percent. Smoothing compresses exactly that tail. If the goal is plain variance control rather than honest targets, a weight penalty is the more direct dial.
  • How does flipping a small percentage of labels compare with smoothing every target?
    In expectation they are the same statement: a label flipped with probability eps has expected target 1-eps. Flipping realises that expectation one row at a time, so it adds variance, and a flip inside a small class can be genuinely costly. Deterministic smoothing gets the same regularizing effect without the sampling noise, so prefer it unless you specifically want to stress-test robustness to real mislabelling.

saying these in an interview costs you the question

  • Says deliberately corrupting labels can never help
  • Expects label smoothing to raise accuracy
  • Uses smoothing instead of fixing systematically mislabelled rows
  • Applies a large eps to labels known to be exact
  • Confuses smoothing the targets with jittering the features

context