Why is classification accuracy useless as a training loss for a neural network?
answer
- what shape is the accuracy curve?
- counts jump, they do not slide
- piecewise constant in the scores
- gradient zero almost everywhere
- needs a differentiable stand-in
basics
~20 sAccuracy counts correct predictions, so it is a step function of the model's scores. Nudging a weight usually changes nothing at all, leaving a gradient of zero almost everywhere and no slope for gradient descent to follow.
solid answer
~40 sAccuracy is built from the 0-1 loss: 1 if the prediction is wrong, 0 if it is right. As a function of the weights it is piecewise constant — the count only changes when some example's score crosses the decision boundary, and everywhere else a small weight change moves the score but not the count. So the gradient is exactly zero almost everywhere and undefined at the jumps, and gradient descent has nothing to descend. The fix is a **surrogate**: a differentiable function of the scores that upper-bounds or tracks the 0-1 loss and keeps decreasing as the correct class pulls ahead. Cross-entropy and hinge are the standard ones. The metric you report does not change — you still report accuracy; you just do not differentiate it.
go deeper
Be ready to say in one sentence that accuracy counts correct predictions, so it moves in jumps and a small weight change leaves it flat, giving gradient descent no slope to follow.
Explain the 0-1 loss as a piecewise-constant function of the scores, and describe what makes hinge or cross-entropy a usable surrogate: differentiable, monotone in the right direction, and an upper bound on the error count.
Demonstrate that you keep objective and metric separate in practice: optimise the surrogate, but early-stop, select and threshold on the metric, and investigate when validation loss and validation accuracy start moving in opposite directions.
Own the call on when a bespoke differentiable metric surrogate earns its complexity against a standard loss plus post-processing, including the review and maintenance cost of a custom objective nobody else on the team has seen.
## The metric and the objective are two different objects Every supervised training run has two functions in play. One is the **metric**: the number stakeholders read, such as accuracy, F1 at a fixed threshold, or a revenue figure. The other is the **training loss**: the function whose gradient actually moves the weights. Beginners assume these should be the same function. For almost every interesting metric they cannot be, and understanding why is the entry point to this whole area. ## What accuracy looks like as a function of the weights For a binary classifier producing a real-valued score `s` and a label `y` in {-1, +1}, the 0-1 loss is `L01 = 1 if y*s <= 0 else 0`. Accuracy is one minus the average of that over the dataset. Now imagine holding the data fixed and sweeping a single weight. Each example's score `s` moves smoothly as the weight moves — but `L01` for that example does not move at all until `s` crosses zero, at which point it jumps by a full unit. So accuracy, as a function of the weights, is **piecewise constant**: flat plateaus separated by cliffs. Its derivative is zero on every plateau — that is, almost everywhere — and undefined on the cliffs. Backpropagation multiplies a chain of derivatives together; a zero at the top of that chain zeroes the whole thing. The update is `w <- w - lr * 0`, i.e. no update. Training would stand still no matter how wrong the model was. There is a second, subtler problem. Even setting gradients aside, exactly minimising 0-1 error is combinatorially hard: the objective is non-convex with a huge number of flat regions and no local information pointing toward a better one. Search-based optimisation can attack it for a handful of parameters, but not for millions of weights. ## What a surrogate has to do A surrogate loss replaces the step with a shape that has a useful slope. The requirements are modest but real: - **Differentiable (or subdifferentiable) in the scores**, so a non-zero gradient flows back. - **Monotone in the right direction**: it must decrease as the correct class's score rises relative to the alternatives, so pushing the loss down pushes the model toward correct predictions. - Ideally it **upper-bounds the 0-1 loss**, so driving the surrogate to zero forces the error count to zero, and the surrogate value is a certificate on the error rate. Two classic choices for a binary score `s` with `y` in {-1, +1}: - **Hinge**: `max(0, 1 - y*s)`. At `y*s = 0` it equals 1, matching the 0-1 loss at the boundary; it sits above the step everywhere else in the wrong region. It has a constant-magnitude slope for violators and is exactly flat once the example clears the margin. - **Logistic / cross-entropy**: `log(1 + exp(-y*s))`. Smooth everywhere, never exactly zero, with a gradient of the familiar `p - y` form when written in probability terms. It also upper-bounds the 0-1 loss once rescaled by a constant. Both are *classification-calibrated*: with enough data and capacity, the score that minimises the surrogate makes the same sign decision as the one that minimises the error rate. That is the theoretical licence for the swap — you are not optimising a random proxy, you are optimising something whose optimum agrees with the metric's optimum. ## Where the two can still come apart Agreement in the limit is not agreement in your run. Cross-entropy keeps rewarding extra confidence on examples that are already correct, so a batch containing a few violently wrong, confidently-scored examples can dominate the gradient while accuracy barely moves. Under heavy class imbalance a model can drive cross-entropy down by sharpening its already-correct majority-class predictions and never learn the minority class at all. It is completely normal — and a good diagnostic habit — to watch validation loss and validation accuracy on the same plot and notice when they diverge: loss creeping up while accuracy holds steady usually means the model is getting more confident on the ones it gets wrong. ## The practical stance Train on the surrogate; select, threshold, and report on the metric. Model selection, early stopping and the decision threshold are all free to use the non-differentiable metric directly, because none of them needs a gradient — they only need to compare a handful of candidates. Reserve differentiating a bespoke metric-shaped loss for cases where that split genuinely fails you, because a custom objective is code someone has to review and defend for the life of the model.
- If accuracy has no gradient, why does minimising cross-entropy usually raise accuracy anyway?Because cross-entropy is classification-calibrated: the score that minimises it makes the same decision as the score that minimises error rate, and it upper-bounds the 0-1 loss up to a constant. Pushing it down widens the score gap in the correct direction. The link is asymptotic, though — in a finite, imbalanced run the loss can improve while accuracy stalls, which is why you track both.
- Are there ways to optimise a non-differentiable metric directly?Yes, with gradient-free search: grid or random search, evolutionary methods, or Bayesian optimisation can all optimise a metric they can only evaluate. They scale to tens of parameters, not to millions of weights, so the usual division of labour is gradient descent on a surrogate for the network and direct search for the few decision parameters, such as a threshold, sitting on top of it.
- Your validation loss rises while validation accuracy stays flat. What is happening?The model is becoming more confidently wrong on the examples it already misclassifies without flipping any decisions. Cross-entropy grows without bound as a wrong prediction approaches certainty, while accuracy only counts sign changes. It is a warning about calibration and overfitting rather than about decision quality, and it argues for early-stopping on the metric you actually ship, not on the loss.
Accuracy is a staircase and cross-entropy is a ramp beside it. Standing on a step, you cannot feel which way is down; on the ramp you always can, and it takes you to the same floor.
saying these in an interview costs you the question
- Says accuracy is merely too slow or expensive to compute each step
- Claims accuracy has a gradient, just a very noisy one
- Assumes lower cross-entropy always means higher accuracy
- Cannot distinguish the reported metric from the trained objective
- Thinks any differentiable function will do as a surrogate