Why does squared error on a softmax classification head train so slowly?
answer
- the Jacobian factor survives instead of cancelling
- smallest update where the model is most wrong
- the error gets multiplied by a tiny probability
- convex in the linear case, then it is not
basics
~10 sSquared error keeps the softmax Jacobian's probability factor instead of cancelling it, so an example whose correct class sits at probability 0.01 sends almost no signal. The composed objective is also non-convex, unlike cross-entropy.
solid answer
~50 sSquared error treats the softmax output as a regression target: `L = 0.5 * sum_k (p_k - y_k)^2`. Backpropagating to the logits gives `dL/dz_m = sum_k (p_k - y_k) * p_k * (delta_km - p_m)`, and every term carries a probability factor from the softmax Jacobian. If the correct class currently has `p_c = 0.01`, its contribution is scaled by roughly 0.01 - the update is smallest exactly where the model is most wrong, which is backwards. Cross-entropy's `-log` contributes `1/p_c`, which cancels that factor and leaves `p - y`. Second, composing squared error with a softmax is non-convex even when the logits are a linear function of the inputs, so there are flat regions and poor curvature that the cross-entropy version, which is convex in that same linear case, does not have. The two objectives share a minimiser but not the path to it.
go deeper
Be able to say that classification heads use cross-entropy and that squared error on class labels gives very small updates when the model is badly wrong, so training stalls rather than fails loudly.
Expect to show the mechanism: chain squared error through the softmax Jacobian, point at the probability factor that survives, and contrast it with the logarithm's reciprocal that cancels it in cross-entropy.
Recognise the symptom from a training curve - loss flat well above target, tiny output-layer gradient norms, worst at the start - and explain why tuning the step size is the wrong response.
Own the general principle when specifying objectives: judge a loss by the gradient it delivers away from the optimum, not by where its minimum sits, and defend when a bounded-output regression target genuinely justifies squared error.
## The setup You have a classification head: logits `z`, probabilities `p = softmax(z)`, and a one-hot target `y`. Two candidate objectives: ``` cross-entropy: L_ce = -log p_c (c = index of the true class) squared error: L_mse = 0.5 * sum_k (p_k - y_k)^2 ``` Both are minimised by `p = y`, so at first glance the choice looks cosmetic. It is not, because training does not care where the minimum is - it cares about the gradient everywhere else. ## Where the gradient goes The softmax Jacobian is `dp_k/dz_m = p_k * (delta_km - p_m)`. Chain squared error through it: ``` dL_mse/dz_m = sum_k (p_k - y_k) * p_k * (delta_km - p_m) ``` Every single term is multiplied by a probability `p_k` coming from the Jacobian, and often by a second probability from the `(delta_km - p_m)` factor. Now take the case that matters: the model is badly wrong, with the true class at `p_c = 0.01`. The error `(p_c - 1)` is large in magnitude, about `-0.99`, which is what you want. But it is immediately multiplied by `p_c = 0.01` and then by factors like `(1 - p_c)`, so the surviving signal is of order `10^-2`, and the cross terms are of order `10^-4`. The example the model most needs to learn from produces one of the smallest updates in the batch. Cross-entropy inverts that relationship. `d(-log p_c)/dp_c = -1/p_c` is the exact reciprocal of the leading Jacobian factor, so the two cancel and the whole gradient collapses to `dL_ce/dz = p - y`. At `p_c = 0.01` the correct logit's component is `-0.99` - close to the maximum this loss can produce. Gradient magnitude grows with wrongness under cross-entropy and shrinks with wrongness under squared error. That is the entire practical story: the run does not diverge, it crawls, and it crawls worst at the start when everything is wrong. ## The non-convexity The second problem is about the shape of the objective rather than its slope. Take the simplest possible model, logits that are a linear function of the input. Cross-entropy composed with a softmax over linear logits is **convex** in the weights - a single basin, no spurious local minima. Squared error composed with the same softmax is **not** convex: the composition introduces curvature that can be flat or misleading in regions where the model is confidently wrong, and the objective admits stationary points that are not the global optimum. For a deep network the whole objective is non-convex either way, so this is not a proof that the deep case is safe with cross-entropy. But the linear case is the diagnostic: it tells you the badness comes from composing squared error with a softmax, not from the depth of the model, and every extra layer sits on top of that already-bad output geometry. ## What it looks like in practice The symptom is a run that appears to be working and is not: the loss falls a little and then flattens well above what the task should reach, accuracy hovers near the majority-class rate for a long time, and gradient norms at the output layer are orders of magnitude smaller than they should be. Nothing errors out. It is much easier to spot when you know that the flat phase is worst early - when every example has the true class near the bottom of the distribution - and eases as the model happens to get some examples right, which is the reverse of a normal training curve's dynamics. Raising the learning rate does not rescue it. The vanishing factor is **per example**: the healthy examples get the same multiplier as the sick ones, so a global step-size increase scales both and destabilises the healthy ones before it makes the stuck ones move. And a step size cannot change the geometry of a non-convex plateau. ## When squared error on probabilities is actually fine The argument above is about **class labels**. If the target is a genuine continuous proportion - a fraction of clicks, a mixing coefficient, a rate you actually measured - then it is a regression task whose output happens to lie in `[0, 1]`, and squared error against it is a defensible modelling choice with a bounded, non-explosive gradient. What you must not do is take a one-hot label, call it a regression target, and pay the vanishing-gradient tax without noticing. ## How to answer this in an interview The strong answer has three beats, in this order: (1) the gradient of squared error keeps the softmax Jacobian factor while cross-entropy's logarithm cancels it, (2) the consequence is a near-zero update precisely for the most-wrong examples, since the probability factor multiplies the error rather than the other way round, and (3) the composed objective additionally loses convexity that the cross-entropy version retains in the linear case. The weak answer stops at `cross-entropy is the standard choice for classification`, which tells the interviewer only that you have read the defaults.
- Would a much larger learning rate fix a softmax head trained with squared error?No. The shrinking factor is per example, not global: the badly wrong examples are damped and the nearly correct ones are not, so scaling every step up destabilises the healthy examples long before it unsticks the stuck ones. A step size also cannot remove a non-convex plateau - it only changes how fast you traverse the same surface.
- Both losses are minimised at p equal to the target, so why does the choice matter?Because optimisation never sits at the minimum during training; it follows gradients from wherever it currently is. Two objectives can share a global minimum and still have completely different slopes and curvature everywhere else. Cross-entropy makes the gradient grow with the size of the mistake, squared error through a softmax makes it shrink, and that difference governs whether the run gets there at all.
- Is squared error ever the right loss on an output constrained to [0, 1]?Yes, when the target is a real measured proportion rather than a class label - a click rate, an observed fraction, a mixing weight. Then it is genuinely regression on a bounded quantity and squared error is a reasonable modelling choice. The mistake is applying it to a one-hot label, which is a likelihood question dressed up as a regression.
saying these in an interview costs you the question
- Says both losses share a minimum so the choice is cosmetic
- Blames the learning rate rather than the surviving Jacobian factor
- Claims squared error is not differentiable through a softmax
- Thinks cross-entropy is preferred purely for numerical reasons
- Asserts the composed squared-error objective is convex