Validation cross-entropy is rising while validation accuracy still improves — what explains it?
answer
- accuracy counts, loss weighs
- confidence, not correctness
- one wrong example can dominate
- -log p is unbounded below one
- ask what the output is used for
basics
~20 sCross-entropy scores confidence, accuracy scores only the decision. As training continues the model becomes more confident everywhere, so the examples it still gets wrong contribute a large and growing penalty while newly-correct ones add almost nothing.
solid answer
~50 sThe two curves measure different things, so they can move apart. Accuracy counts how many argmax decisions are right and is capped at one; cross-entropy is `-log p` on the true class, which saturates near zero for a confidently correct example but grows without bound as a wrong prediction becomes more confident. Long training pushes probabilities toward the extremes, so stubbornly wrong examples dominate the mean loss while borderline ones flip to correct and lift accuracy. The practical reading is that the model is still improving as a decision rule but getting worse-calibrated. What you do depends on the product: if you threshold or rank on the probabilities, the rising loss is a real regression and points at calibration; if only the top-1 decision ships, the accuracy trace is the one that matters. Before either, check that the rise exceeds the validation split's own epoch-to-epoch noise.
go deeper
Know that accuracy only counts right and wrong decisions while cross-entropy also penalizes how confident a wrong prediction was, so the two curves need not move together.
Explain the mechanism: -log p is unbounded as the true-class probability approaches zero while correct examples saturate near zero loss, so a few confident errors can lift the mean.
Tie the reading to the downstream use — thresholded, ranked or top-1 — and show that you check the split's noise level before treating either trend as real.
Own which metric the team selects models on, and make the calibration requirement explicit for products whose decisions consume probabilities rather than labels.
## Two metrics, two questions Accuracy asks: is the arg-max class the right one? It is a step function of the prediction — nothing changes as a correct probability moves from 0.51 to 0.99. Cross-entropy asks: how much probability mass did you put on the truth? For a single example with true-class probability `p`, the loss is `-log(p)`. It is small and slowly-varying near `p = 1`, and it grows without limit as `p` approaches zero. That asymmetry is the whole phenomenon: one confidently wrong example at `p = 0.001` contributes about 6.9 nats, which the correct examples cannot offset, because their individual losses are already near zero and bounded below by it. ## Why long training produces the divergence Minimizing cross-entropy on training data pushes the model to be more decisive. As it fits the training set more tightly, its output distribution sharpens on held-out data too — including on examples it gets wrong. So on the validation set: - The already-correct examples move from moderately confident to very confident. Their loss falls, but only by a little, since it was already small. - The wrong examples move from moderately wrong to confidently wrong. Their loss rises a lot, and it is unbounded. - A few examples near the decision boundary flip to the correct side. Accuracy rises. Sum it up and the mean loss climbs while the count of correct decisions climbs too. This is a well-known miscalibration effect in long-trained, high-capacity networks: they end up more accurate and more overconfident at the same time. ## What it is not It is not proof that training has gone wrong. It is not the same as a plain widening train/validation gap, where the model is simply memorizing and both loss and accuracy degrade. And it is not a bug in the metric code — although a genuine bug worth excluding first is computing the two quantities on different data, different preprocessing, or a different subset. ## Deciding which trace to trust Ask what the model's output is used for downstream. - **A hard top-1 decision with no thresholding.** Accuracy — or better, the task metric that reflects the real cost of each error type — is the signal that matters. A rising loss is then a calibration observation, not a regression. - **A threshold, an abstain rule, or a cost-weighted decision.** The probability values are the product. A rising cross-entropy means the numbers you threshold on are drifting, and the operating point tuned earlier may no longer be valid. - **Ranking or downstream fusion.** Cross-entropy is only a proxy; measure the ranking metric directly, since calibration can degrade while ordering stays fine. When the probabilities matter and the model is otherwise good, the standard response is to fix calibration separately rather than to abandon the model — a monotone rescaling of the scores fitted on held-out data changes no decision ordering and no accuracy, only the numbers. Label smoothing and stronger regularization also damp the drift at the source, at some cost in peak accuracy. ## Do not skip the noise check Both curves are estimates on a finite split. On a validation set of a few hundred examples, a single example flipping moves accuracy by a third of a point, and one confidently wrong example can move the mean loss visibly on its own. If the epoch-to-epoch swing of the validation trace is larger than the trend you are arguing about, there is no trend to explain. Establish that swing by looking at the variation across several adjacent epochs — or across seeds — before treating any small movement in either curve as a finding. This is the most common way a team spends a week explaining a difference that was never there. ## The senior answer in one shape Name the mechanism (a proper scoring rule punishes confident errors without bound, a counting metric does not), state what it implies (the model is improving as a decision rule and degrading as a probability estimator), tie the response to the downstream use, and gate all of it on the split being large enough for the movement to be real.
- The validation metric moves by 0.3 points between epochs on a 300-example split. Is that a finding?Almost certainly not. A single example is worth a third of a point on 300, so a 0.3-point move is one example changing its mind. Estimate the split's own epoch-to-epoch swing across several adjacent epochs, or across seeds, and refuse to interpret anything smaller than it. If decisions of that size matter, the fix is a bigger or repeated evaluation, not a better story.
- If the probabilities are what your product consumes, how do you respond to the rising loss?Treat it as a calibration problem rather than a training failure. Fit a monotone rescaling of the scores on a held-out split, which leaves accuracy and ranking untouched while restoring the meaning of the numbers, and re-check the operating point that was tuned on the older, better-calibrated outputs. Damping the drift at source — for example with label smoothing — is the alternative, at some cost in peak accuracy.
A student who answers more questions correctly than last term but now writes every wrong answer in bold, underlined, with no hedging: their score improves, their reliability does not.
saying these in an interview costs you the question
- Assumes rising validation loss always means training has failed
- Believes cross-entropy and accuracy must move in the same direction
- Reports accuracy without ever checking calibration
- Treats a tiny loss movement on a small split as signal
- Compares the two metrics computed on different subsets