skip to content

Why is squared-error loss on a logistic model's sigmoid output not convex in the weights?

level: middleimportance: nice to knowfreq 32%

answer

  1. convex in what, exactly?
  2. an S-curve sits between weights and loss
  3. log undoes the exponential inside sigmoid
  4. sigmoid derivative near zero when saturated
  5. log loss gradient: (p - y) times x

basics

~20 s

Squaring the gap between a label and a sigmoid composes a bowl with an S-curve, producing flat saturated regions and stationary points that are not the global best. Cross-entropy instead collapses to a convex function of the linear score.

solid answer

~50 s

For one row the squared error is `(sigma(z) - y)^2` with `z = w*x + b`. The sigmoid is S-shaped — it curves one way below zero and the other way above — so squaring its distance to the label gives a surface that is not a single bowl in the weights: it flattens out at both saturated ends and can hold stationary points that are not the global optimum. Cross-entropy avoids this because `-[y*log(p) + (1-y)*log(1-p)]` simplifies to `log(1 + e^z) - y*z`, which curves upward in `z`, and `z` is linear in the weights. There is a gradient consequence as well. The squared-error gradient is `2*(sigma(z) - y)*sigma'(z)*x`, and `sigma'(z)` is nearly zero exactly when the model is confidently wrong, so the worst rows barely move the weights. With log loss the sigmoid derivative cancels and the gradient is `(p - y)*x`, proportional to how wrong the prediction was.

go deeper

for a junior

Recall that cross-entropy is the standard loss for a sigmoid output and squared error is not, and that the reason involves the sigmoid sitting between the weights and the prediction. You are not expected to derive it.

for a middle

Explain both halves: composing a bowl with an S-curve is no longer a bowl in the weights, and the squared-error gradient carries a sigmoid-derivative factor that vanishes exactly on confidently-wrong rows. Write the two gradients side by side.

for a senior

Recognise the symptom in a real fit — a loss plateau above zero with saturated, badly misclassified rows barely moving — and connect it to the loss choice rather than blaming the data or the optimiser.

for a principal

Frame it as matching a loss to an output layer and to what the business scores you on. Be able to separate the training objective from the evaluation measure, and to justify when a proper scoring rule belongs in reporting but not in the descent.

## The two candidate losses A binary classifier with a sigmoid output turns a linear score `z = w*x + b` into a probability `p = sigma(z) = 1 / (1 + e^-z)`. Two obvious ways to score that against a label `y` in {0, 1}: ``` squared error : (p - y)^2 cross-entropy : -[ y*log(p) + (1-y)*log(1-p) ] ``` Squared error looks like the safer, more familiar choice — it is what regression uses, and it is convex in `p`. That is the trap: convexity in the predicted probability is not what matters. What matters is the shape of the loss as a function of the **weights**, because the weights are what the optimiser moves. ## Why the composition breaks convexity Convexity survives composing a convex function with a *linear* map. It does not survive composing a convex function with an S-shaped one. The sigmoid curves upward for negative scores and downward for positive scores, so `(sigma(z) - y)^2` inherits that change of curvature. Trace it for `y = 1`: at very negative `z` the loss is close to 1 and almost flat; through the middle it drops steeply; at very positive `z` it is close to 0 and almost flat again. That S-shaped descent is not a bowl. Add several rows pulling in different directions and the summed surface can hold stationary points that are not the global minimum, plus large plateaus where the gradient is nearly zero for reasons that have nothing to do with being at the optimum. ## Why cross-entropy stays convex Substitute the sigmoid into the cross-entropy and it simplifies dramatically: ``` -[ y*log(p) + (1-y)*log(1-p) ] = log(1 + e^z) - y*z ``` The first term is softplus, which curves upward everywhere. The second is linear in `z`, and adding a linear term cannot introduce a downward bend. So the per-row loss curves upward in `z`, and because `z` is linear in the weights, the summed training loss is a single bowl in the weights. The logarithm in cross-entropy is precisely what undoes the exponential inside the sigmoid — that cancellation is the whole reason the pairing works. ## The gradient argument, which interviewers like more Even where the squared-error surface is locally well behaved, its gradient is badly scaled. Differentiating: ``` squared error : d/dw = 2*(p - y) * sigma'(z) * x, sigma'(z) = p*(1 - p) cross-entropy : d/dw = (p - y) * x ``` The factor `p*(1-p)` is at most 0.25, and it collapses toward zero as `p` approaches 0 or 1. Consider a row with label 1 that the model currently predicts at `p = 0.001`. This is the most wrong the model can be, and it is exactly where you want a large correction. Squared error multiplies the error by `0.001 * 0.999 = 0.000999`, producing an update roughly a thousand times smaller than the same row would get from log loss, where the derivative cancels and the update size is simply `(p - y)` times the feature. A model initialised badly, or saturated on a subset of rows, can sit there for a very long time. ## What is actually true about each claim - "Squared error is convex" — true in the prediction, false in the weights once a sigmoid sits in between. Always ask *convex in what*. - "Squared error cannot be used at all" — too strong. It can be optimised, and on easy data it often finds a good solution. You simply lose the guarantee, and you inherit the vanishing-update behaviour. - "Cross-entropy is convex, so any classifier trained with it is convex" — false. The convexity comes from the score being linear in the weights. The loss function alone is not enough. - "Squared error on the probability is the same as squared error in linear regression" — no. Linear regression has no sigmoid between the weights and the prediction, which is exactly what preserves the bowl. ## The practical takeaway Cross-entropy is the matched loss for a sigmoid output: it makes the training objective a single bowl in the weights, and it makes the size of each update proportional to the size of the mistake. That pairing is why it is the default, not tradition. If you are ever asked to defend the choice, give both halves — the convexity argument and the gradient argument — because the second one is what you actually feel during training.

  • Squared error is convex in the predicted probability, so why does that not settle it?
    Because the optimiser moves the weights, not the probability. The relevant question is the shape of the loss as a function of the parameters, and the sigmoid sitting between the weights and the probability is what destroys that shape. Convexity is never a property of a loss alone — always of a loss composed with a model.
  • What does the vanishing update look like during an actual fit?
    A subset of rows sits saturated at the wrong end — predicted near 0 with label 1, or the reverse — and the loss plateaus at a clearly non-zero value while the weights barely move. Switching to cross-entropy on the same data makes those rows produce full-size updates immediately, and the plateau disappears.
  • Is there any setting where squaring the error against a probability is deliberate?
    Yes, as an evaluation measure rather than a training objective. The mean squared error between a predicted probability and a 0/1 outcome is the Brier score, a proper scoring rule used to assess probability quality. Using it to score a model is uncontroversial; using it as the thing you descend on is what costs you convexity.

saying these in an interview costs you the question

  • Says squared error is convex, without saying convex in what
  • Claims the sigmoid itself is a convex function
  • Thinks cross-entropy is convex for any model
  • Cannot state the gradient of log loss with a sigmoid
  • Argues squared error simply converges more slowly, nothing else

context