skip to content

Why is logistic regression trained with cross-entropy loss rather than squared error?

level: middleimportance: must knowfreq 76%

answer

  1. start from the Bernoulli likelihood
  2. logarithm turns a product into a sum
  3. the sigmoid slope cancels in one loss
  4. gradient on the score is predicted minus actual
  5. confidently wrong at 0.999 with a 0.002 gradient

basics

~20 s

Cross-entropy is the Bernoulli negative log-likelihood, and its gradient on the linear score is simply predicted minus actual. Squared error through a sigmoid is non-convex in the weights and its gradient nearly vanishes exactly where the model is confidently wrong.

solid answer

~50 s

For one row with label `y` in {0, 1} and predicted probability `p`, cross-entropy is `-[y*log(p) + (1-y)*log(1-p)]`, summed over rows. That is exactly the negative log-likelihood of a Bernoulli model, so minimising it is maximum-likelihood fitting, and it is convex in the weights, so the optimum is unique and reachable. Its gradient with respect to the linear score is `p - y` - the sigmoid's slope cancels out of the derivation. Squared error `(p - y)^2` gives a gradient of `2*(p - y)*p*(1 - p)`, which carries the sigmoid's slope as a factor. On a row predicted at 0.999 whose true label is 0, cross-entropy contributes a loss of about 6.9 and a gradient of 0.999, while squared error contributes 0.998 and a gradient of 0.002 - roughly 500 times weaker, precisely where the model is most wrong.

code

python · 18 lines
python
import math

def sigmoid(z):
    return 1.0 / (1.0 + math.exp(-z))

z, y = 6.9, 0.0          # score pushed far positive, true label is 0
p = sigmoid(z)

ce_loss = -(y * math.log(p) + (1 - y) * math.log(1 - p))
sq_loss = (p - y) ** 2

ce_grad = p - y                      # d(cross-entropy)/dz
sq_grad = 2 * (p - y) * p * (1 - p)  # d(squared error)/dz

print("p        ", round(p, 4))          # 0.999
print("losses   ", round(ce_loss, 3), round(sq_loss, 3))   # 6.901 0.998
print("gradients", round(ce_grad, 4), round(sq_grad, 6))   # 0.999 0.002009
print("ratio    ", round(ce_grad / sq_grad, 1))            # ~497

go deeper

for a junior

Recall the per-row formula, -[y*log(p) + (1-y)*log(1-p)], and that only one term is live per row: positives cost -log(p), negatives cost -log(1 - p).

for a middle

Derive it. Start from the Bernoulli likelihood, take the log, negate, and differentiate to reach p - y on the score. Then contrast with the squared-error gradient and name the extra p*(1-p) factor as the culprit.

for a senior

Demonstrate that you have seen this bite: a fit that stalls with a near-maximum loss, a handful of noisy labels dominating an unbounded objective, or an optimiser that behaves differently from a different starting point because the surface was not convex.

for a principal

Own the choice of objective as a design decision. Argue when the plain likelihood objective is right and when the business cost of an error is asymmetric enough that the objective, not just a downstream decision rule, should reflect it.

## What the objective is A sigmoid classifier predicts `p = sigmoid(w*x + b)`, a number in (0, 1). Training means choosing the weights that make the observed labels most plausible. Write the model's assumption explicitly: each row's label is a Bernoulli draw with success probability `p`. The probability the model assigns to a single observed label `y` is then ``` P(y | x) = p^y * (1 - p)^(1 - y) ``` which reads as `p` when `y = 1` and `1 - p` when `y = 0`. Taking the logarithm and negating gives the per-row loss: ``` loss = -[ y*log(p) + (1 - y)*log(1 - p) ] ``` That is cross-entropy, also called the log loss or binary cross-entropy. Summed (or averaged) over rows it is the objective the fit minimises. Because the log turns the product over rows into a sum, minimising it is the same as maximising the likelihood of the data under the model. Only one of the two terms is ever live: for a positive row the loss is `-log(p)`, for a negative row it is `-log(1 - p)`. In both cases the loss is 0 when the model is perfectly right and rises without bound as the prediction moves towards the wrong end. Predicting 0.5 for everything costs `-log(0.5) = 0.693` per row - the reference number a warranty-claim model starts from when it has learned nothing beyond a balanced base rate, and the number it must beat as the weights move. ## Why not squared error Squared error, `(p - y)^2`, is the natural reflex from linear regression, and it is a bad fit here for three separate reasons. **1. The gradient vanishes where the model is most wrong.** Differentiate both losses with respect to the linear score `z`, using `dp/dz = p*(1 - p)`: ``` cross-entropy: d loss / dz = p - y squared error: d loss / dz = 2*(p - y)*p*(1 - p) ``` The cross-entropy derivation cancels the sigmoid's slope against the `1/p` from the logarithm, leaving the residual alone. Squared error keeps `p*(1 - p)` as a multiplier. Now take a row that the model scores at 0.999 whose true label is 0 - a confidently wrong prediction, exactly the row you most want to fix. Cross-entropy gives a loss of 6.9 and a gradient of 0.999. Squared error gives a loss of 0.998 - almost its maximum of 1, so it *knows* the row is wrong - but a gradient of about 0.002, because `p*(1 - p) = 0.001` at that point. The update is roughly 500 times smaller. The model is stuck in the saturated tail and crawls out of it. **2. It is non-convex in the weights.** Cross-entropy composed with a sigmoid on a linear score is convex in `w` and `b`: there is a single optimum and no local minima to get trapped in. Squared error composed with a sigmoid is not convex, so the surface can have flat regions and local minima that depend on where you started. For a model whose whole selling point is a reliable, reproducible fit, that is a real loss. **3. It is not the likelihood.** Squared error is the maximum-likelihood objective for Gaussian noise on a continuous target. The target here is a 0/1 label, and the model's noise assumption is Bernoulli. Fitting with squared error means optimising a criterion the model does not correspond to, and forfeits the statistical machinery that follows from a likelihood. ## The gradient in weight space Applying the chain rule one more step, for weight `j`: ``` d loss / d w_j = (p - y) * x_j ``` Summed over rows this is the whole gradient, and it has a satisfying reading: each row pushes each weight in proportion to its own feature value times how far the prediction missed. A row the model already gets right contributes almost nothing; a badly missed row dominates. This form is why logistic regression trains stably with plain gradient descent, and why the update rule looks identical to the linear-regression one despite the non-linear squashing in between. ## Asymmetry, and the price of confidence Cross-entropy is unbounded above. A single row predicted at 0.001 whose label is 1 contributes `-log(0.001) = 6.9`; predicted at 0.000001 it contributes 13.8. Squared error caps out at 1 per row no matter how wrong. That unboundedness is the intended behaviour - it makes the fit strongly averse to confident mistakes and, in the tails, it is the only thing keeping the optimiser awake. The practical consequence is that a handful of mislabelled rows can dominate the objective, so if labels are noisy the objective will happily distort the fit to accommodate them. ## Common mistakes - Saying squared error "doesn't notice" a confidently wrong row. It does - the loss is near its maximum. What is missing is the *gradient*, not the loss value. - Saying cross-entropy is preferred because it is a better report of model quality. Here it is the objective the fit minimises; whether it is also the right number to report is a separate matter. - Writing the loss as `-log(p)` for every row. That term applies only to positive rows; negative rows use `-log(1 - p)`.

  • Write the gradient of the cross-entropy objective with respect to a single weight.
    For weight `j` it is the sum over rows of `(p - y) * x_j`, where `p` is that row's predicted probability. Each row nudges the weight in proportion to its own feature value scaled by the size and sign of the miss, so rows the model already predicts well contribute almost nothing. The sigmoid's derivative has cancelled out, which is why the update looks like the linear-regression one.
  • Is the cross-entropy objective convex in the weights, and does that guarantee a unique solution?
    It is convex, so gradient descent cannot land in a spurious local minimum. Uniqueness is not automatic: with perfectly separable classes the weights can grow without bound and no finite minimiser exists, and with perfectly collinear features many weight vectors give the same fit. A penalty on the weights restores a unique, finite optimum in both cases.
  • Does the same saturation argument apply when the model is confidently right?
    Yes, but there it is desirable. A row predicted at 0.999 whose label is 1 contributes a loss of about 0.001 and a gradient of -0.001 under cross-entropy, so it barely moves the weights - correct behaviour, since there is nothing to fix. The problem with squared error is that it treats the confidently wrong row almost as gently as the confidently right one.

saying these in an interview costs you the question

  • Claims squared error fails because its loss stays small when wrong
  • Says cross-entropy is chosen just because it is standard
  • Writes the per-row loss as -log(p) regardless of the label
  • Thinks the sigmoid derivative survives in the cross-entropy gradient
  • Believes squared error through a sigmoid is convex in the weights

context