skip to content

Why does the bias-variance decomposition not add up cleanly under 0-1 classification loss?

level: seniorimportance: nice to knowfreq 24%

answer

  1. the proof only ever used the square
  2. counting mistakes is a step, not a curve
  3. wobbles that stay on one side are free
  4. variance can carry a negative sign here

basics

~20 s

Squared error is quadratic, so expanding it leaves exactly three additive non-negative terms. Counting misclassifications is a step function of the prediction error, so no such algebra exists, and instability across refits can even lower the error rate.

solid answer

~50 s

The clean split falls out of the algebra of squaring: expand `(y - prediction)^2` and the cross terms vanish, because label noise has mean zero and deviations from the average prediction have mean zero by construction, leaving squared bias, variance and noise, all non-negative and all additive. The 0-1 loss counts only whether the predicted class is wrong, so it is a step function of the deviation: a wobble that never crosses the decision boundary costs nothing, and a small one that does costs the full amount. Several 0-1 decompositions have been proposed and they disagree, but they share one striking feature - the variance contribution can be negative. If the average prediction sits on the wrong side of the boundary, instability pushes some refits onto the correct side and lowers the expected error rate. For classifiers, use the framing qualitatively, or decompose squared error on the predicted probability instead.

go deeper

for a junior

You are unlikely to be asked this. Knowing that the neat three-term split is specific to squared error, and not a universal law about every metric, is already a good answer at this level.

for a middle

Be able to point at the derivation: the cross terms cancel because the loss is a square. Say that counting misclassifications is a step function, so the same cancellation is unavailable.

for a senior

Deliver the counter-intuitive consequence with a concrete case - instability helping when the average prediction sits on the wrong side of the boundary - and propose decomposing squared error on the predicted probability instead.

for a principal

Own the reporting standard. Decide what your teams are allowed to claim from a bias-variance argument on classification metrics, and insist that a proper scoring rule backs any quantitative version of it.

## Where the clean split comes from Under squared error the decomposition is not a modelling insight, it is arithmetic. Write the error at a fixed input as the sum of three pieces - label noise, the offset between the truth and the average prediction, and the offset between the average prediction and this particular refit - then square and take expectations. Each square becomes one term. Every cross term dies, because the noise averages to zero and is independent of the training data, and the deviation of a refit from the average of all refits averages to zero by the definition of that average. Three terms, all non-negative, adding exactly. That proof uses one property of the loss and one only: it is the square of the deviation. Change the loss and the proof evaporates. ## Why 0-1 loss breaks it Consider automated weld inspection, where each weld is labelled pass or fail. The 0-1 loss charges 1 if the predicted class is wrong and 0 if it is right. As a function of how far the model's underlying score sits from the truth, that is a step: flat, then a cliff at the decision boundary, then flat again. Two consequences follow immediately. **Deviations below the cliff are free.** If the average model output for a weld is 0.80 and refits scatter between 0.70 and 0.92, all of them classify it as pass. Under squared error that scatter is a real variance cost; under 0-1 loss it costs exactly nothing. So the same instability maps to a large variance term in one accounting and zero in the other, which is already fatal to any hope of a shared formula. **Instability can be an asset.** Take a genuinely failing weld where the average model output is 0.45 - on the wrong side of a 0.5 boundary, so the average prediction is wrong. A perfectly stable model gets it wrong every single time: error rate 1.0. A model whose refits scatter around 0.45 gets it *right* whenever a refit lands above 0.5, so the expected error rate drops below 1.0. Variance helped. Under squared error that can never happen - variance is a sum of squares and only ever adds. This is why every proposed 0-1 decomposition carries a signed variance contribution, usually described as variance being harmful where the average prediction is already correct and helpful where it is not. Several such decompositions exist in the literature and they define bias and variance differently from each other; none of them is *the* decomposition in the way the squared-error one is. ## What to do instead **Decompose a squared-error loss on the probability.** If the model outputs a probability rather than a hard label, apply squared error to that probability against the 0/1 outcome. That loss is quadratic again, so the standard decomposition applies exactly, and the terms mean what you expect. The squared error of a predicted probability against a binary outcome is the Brier score, and it decomposes cleanly. This is usually the honest way to run a bias-variance analysis on a classifier. **Use the vocabulary qualitatively.** Saying "this classifier looks unstable across refits" and "this classifier is systematically wrong on this segment every time" remains useful and true. What you must not do is present accuracy as a sum of three numbers, or claim that reducing variance necessarily improves the error rate. **Watch the threshold.** Because the loss depends only on which side of the boundary a score falls, changing the decision threshold changes the whole picture, moving errors between classes without any change in the model. Any statement about bias and variance under 0-1 loss is implicitly a statement about a particular threshold. ## The trap in the question Candidates who have only ever seen the decomposition as a slogan will confidently write `error = bias^2 + variance + noise` for accuracy. The tell that someone has actually worked with it is that they can name where the derivation used the square, and can produce the counter-intuitive consequence: for a classifier that is systematically wrong at a point, adding instability improves expected accuracy there. That result sounds wrong and is correct, which makes it a good discriminating question - though it is a differentiator, not a screener.

  • Give a concrete case where more instability lowers a classifier's error rate.
    Take an input whose true class is fail and where the average model output over refits is 0.45 against a 0.5 threshold. A perfectly stable model predicts pass every time and is wrong every time. A model whose refits scatter around 0.45 lands above 0.5 some of the time and is right on those refits, so its expected error rate at that input is strictly lower. The bias was the problem; the noise around it partially rescued the prediction.
  • If you wanted a rigorous bias-variance analysis of a classifier, what would you actually decompose?
    The squared error of the predicted probability against the 0/1 outcome - the Brier score. It is quadratic, so the standard three-term decomposition holds exactly and the terms keep their usual meaning. You lose the direct link to the headline accuracy number, but you gain an analysis that is actually valid, and probability quality is what most downstream decisions depend on anyway.

saying these in an interview costs you the question

  • States that accuracy equals squared bias plus variance plus noise
  • Assumes every loss decomposes into three additive non-negative terms
  • Claims a stable classifier always beats an unstable one
  • Cannot say which step of the derivation needs the square
  • Ignores that the decision threshold changes the whole accounting

context