skip to content

How do you read a multiclass log loss of 1.2 from a 20-language classifier?

level: middleimportance: should knowfreq 44%

answer

  1. only the true class's probability counts
  2. compare against an uninformed baseline
  3. that baseline grows with the class count
  4. log of 20 is about 3.0
  5. exponentiate the loss to make it concrete

basics

~20 s

Multiclass log loss averages the negative log of the probability given to the true class. Uniform guessing over 20 languages scores log 20, about 3.0, so 1.2 beats chance and means an average true-class probability near 0.30.

solid answer

~50 s

Multiclass log loss is `-(1/N) * sum over rows of log p(true class)`. Only the probability assigned to the actual language enters each row's term — the other 19 affect it only through the softmax normaliser that made them sum to one. To read 1.2, compare it to the uninformed baseline: predicting a flat 1/20 everywhere gives `log(20)` ≈ 3.00, and a perfect confident model gives 0. So 1.2 is clearly better than chance but far from certain. Exponentiating back, `exp(-1.2)` ≈ 0.30 is the geometric mean probability the model put on the correct language. The loss is unbounded above: a single row given probability 0.001 to the true language contributes about 6.9, so a handful of confident mistakes can dominate the average. That sensitivity is the point — log loss grades calibration, not just whether the argmax was right.

go deeper

for a junior

Know the direction: lower is better, zero is perfect, and the value is the negative log of the probability given to the correct class, averaged over rows.

for a middle

Explain why only the true-class probability appears in each term, anchor a value against the log(K) uniform baseline, and convert back with exp of the negative loss to state an average true-class probability.

for a senior

Demonstrate that you use it as a calibration signal: know that a handful of confidently-wrong rows can dominate the average, and be able to decide when a better-calibrated model beats a more accurate one.

for a principal

Own the choice of headline metric and its baseline for the organisation: what the probability outputs are used for downstream, and whether a proper scoring rule or a decision-level metric should drive model selection.

## The definition For a classifier that outputs a probability distribution over K classes, multiclass log loss (multiclass cross-entropy) over N rows is ``` log_loss = -(1/N) * sum_i log( p_i[ y_i ] ) ``` where `y_i` is the true class of row i and `p_i[y_i]` is the probability the model assigned to that class. Written the other way, with a one-hot target `t_i`, it is `-(1/N) * sum_i sum_k t_ik * log p_ik`, and since `t_ik` is zero for every class except the true one, all those terms vanish. **Only the true-class probability enters each row's term.** That surprises people: if the true language is Portuguese and the model gave Portuguese 0.30, it does not matter for the loss whether the remaining 0.70 sat entirely on Spanish or was spread thinly over the other 18. The wrong classes influence the loss only indirectly, because the softmax normaliser makes the K probabilities sum to one, so mass given to Spanish is mass taken from Portuguese. ## Reading a number A log-loss value is meaningless without its baselines. - **Perfect and confident**: probability 1 on every true class gives `-log(1) = 0`. - **Uninformed**: a flat `1/20` on every row gives `-log(1/20) = log(20)` ≈ **3.00**. In general `log(K)` — note that this baseline moves with the number of classes, so a log loss of 1.2 is mediocre for 3 classes and good for 20. - **Prior-only**: always predicting the training class frequencies gives the entropy of the label distribution, which is below `log(K)` whenever the classes are unbalanced. This is the honest baseline to beat. So 1.2 on 20 languages sits well below the 3.00 uniform baseline. A useful second reading: because the loss is the mean of `-log p`, `exp(-1.2)` ≈ **0.30** is the **geometric mean** of the probabilities assigned to the true languages. Equivalently, `exp(1.2)` ≈ 3.3 is sometimes called the perplexity — loosely, the model behaves as if it were choosing uniformly among about 3.3 languages rather than 20. ## Why the loss is so sensitive to confident mistakes `-log p` is unbounded as `p` approaches zero. Some values worth memorising: ``` p = 0.5 -> 0.69 p = 0.3 -> 1.20 p = 0.1 -> 2.30 p = 0.01 -> 4.61 p = 0.001 -> 6.91 ``` One row where the model gave the true language 0.001 contributes 6.91 to the sum. In a 1000-row evaluation that single row adds 0.0069 to the average — small — but a hundred such rows add 0.69, which would swamp the difference between two otherwise similar models. This is why a log-loss regression between model versions is often traced to a small slice of confidently-wrong inputs rather than to a broad degradation. Because a probability of exactly zero would make the loss infinite, evaluation clips probabilities into a range like `[1e-15, 1 - 1e-15]` before taking the log. Worth mentioning: it shows you know the metric has a failure mode. ## What log loss tells you that accuracy does not Log loss is a **proper scoring rule**: it is minimised, in expectation, exactly when the reported probabilities equal the true conditional probabilities. Reporting anything other than your honest belief makes it worse. Accuracy only asks whether the argmax was right, so it is blind to confidence. The practical consequence is that the two metrics can disagree. Model A may get more languages right while stating 0.99 on everything, including the ones it gets wrong; model B may be right slightly less often but honest about its uncertainty. A wins on accuracy, B wins on log loss. Which you prefer depends on what happens downstream: if the 20-language identifier's output feeds a routing rule with an abstain threshold, or is combined with other evidence, the calibrated model is worth more than the marginally more accurate one. If a hard label is consumed and nothing else, accuracy may be the metric that matters. A related caution: log loss is an average over rows, so with a heavily skewed language mix it is dominated by whichever languages are common. It tells you nothing per class on its own. ## What to say in an interview Give the formula, say that only the true-class probability enters, anchor the number against `log(K)` and the class-prior baseline, and convert it back with `exp(-loss)` to make it concrete. Then name the two behaviours that make it useful and dangerous: it rewards honest calibration, and it is dominated by confident mistakes.

  • Model A has higher accuracy than model B but a worse log loss. What is going on?
    A is right more often but poorly calibrated — typically it states very high probabilities and pays a large penalty on the rows it gets wrong, since `-log p` explodes as the true-class probability approaches zero. B is more honest about uncertainty. Pick based on the downstream use: abstain thresholds and score combination need B, a bare label may prefer A.
  • Why is log(K) not always the right baseline to beat?
    `log(K)` is the loss of uniform guessing, which assumes the classes are equally likely. If the language mix is skewed, always predicting the training class frequencies already beats it, and that entropy-of-the-prior value is the honest floor. Beating uniform on a skewed dataset can mean the model learned only the base rates.
  • What happens if the model assigns probability zero to the true class?
    The term becomes negative infinity and the average is undefined, so implementations clip probabilities away from 0 and 1 before taking the log. The clip bounds the damage but does not hide it: a clipped row still contributes a very large value, which is exactly the signal that the model was confidently wrong.

saying these in an interview costs you the question

  • Says all class probabilities enter each row's term equally
  • Judges a log loss value without any baseline
  • Thinks log loss is bounded above by one
  • Treats lower log loss as always meaning higher accuracy
  • Ignores that a few confident mistakes dominate the average

context