skip to content

Imbalance, Surrogates and Trade-offs

When the plain objective fails the problem: rare positives lost in the mean, a business metric nothing can differentiate, and several losses competing inside one sum. These questions test judgment.

on this pageshow

explore

questions

16

Why is classification accuracy useless as a training loss for a neural network?

level: juniorimportance: must knowfreq 62%

answer

  1. what shape is the accuracy curve?
  2. counts jump, they do not slide
  3. piecewise constant in the scores
  4. gradient zero almost everywhere
  5. needs a differentiable stand-in

basics

~20 s

Accuracy counts correct predictions, so it is a step function of the model's scores. Nudging a weight usually changes nothing at all, leaving a gradient of zero almost everywhere and no slope for gradient descent to follow.

solid answer

~40 s

Accuracy is built from the 0-1 loss: 1 if the prediction is wrong, 0 if it is right. As a function of the weights it is piecewise constant — the count only changes when some example's score crosses the decision boundary, and everywhere else a small weight change moves the score but not the count. So the gradient is exactly zero almost everywhere and undefined at the jumps, and gradient descent has nothing to descend. The fix is a **surrogate**: a differentiable function of the scores that upper-bounds or tracks the 0-1 loss and keeps decreasing as the correct class pulls ahead. Cross-entropy and hinge are the standard ones. The metric you report does not change — you still report accuracy; you just do not differentiate it.

go deeper

for a junior

Be ready to say in one sentence that accuracy counts correct predictions, so it moves in jumps and a small weight change leaves it flat, giving gradient descent no slope to follow.

for a middle

Explain the 0-1 loss as a piecewise-constant function of the scores, and describe what makes hinge or cross-entropy a usable surrogate: differentiable, monotone in the right direction, and an upper bound on the error count.

for a senior

Demonstrate that you keep objective and metric separate in practice: optimise the surrogate, but early-stop, select and threshold on the metric, and investigate when validation loss and validation accuracy start moving in opposite directions.

for a principal

Own the call on when a bespoke differentiable metric surrogate earns its complexity against a standard loss plus post-processing, including the review and maintenance cost of a custom objective nobody else on the team has seen.

## The metric and the objective are two different objects Every supervised training run has two functions in play. One is the **metric**: the number stakeholders read, such as accuracy, F1 at a fixed threshold, or a revenue figure. The other is the **training loss**: the function whose gradient actually moves the weights. Beginners assume these should be the same function. For almost every interesting metric they cannot be, and understanding why is the entry point to this whole area. ## What accuracy looks like as a function of the weights For a binary classifier producing a real-valued score `s` and a label `y` in {-1, +1}, the 0-1 loss is `L01 = 1 if y*s <= 0 else 0`. Accuracy is one minus the average of that over the dataset. Now imagine holding the data fixed and sweeping a single weight. Each example's score `s` moves smoothly as the weight moves — but `L01` for that example does not move at all until `s` crosses zero, at which point it jumps by a full unit. So accuracy, as a function of the weights, is **piecewise constant**: flat plateaus separated by cliffs. Its derivative is zero on every plateau — that is, almost everywhere — and undefined on the cliffs. Backpropagation multiplies a chain of derivatives together; a zero at the top of that chain zeroes the whole thing. The update is `w <- w - lr * 0`, i.e. no update. Training would stand still no matter how wrong the model was. There is a second, subtler problem. Even setting gradients aside, exactly minimising 0-1 error is combinatorially hard: the objective is non-convex with a huge number of flat regions and no local information pointing toward a better one. Search-based optimisation can attack it for a handful of parameters, but not for millions of weights. ## What a surrogate has to do A surrogate loss replaces the step with a shape that has a useful slope. The requirements are modest but real: - **Differentiable (or subdifferentiable) in the scores**, so a non-zero gradient flows back. - **Monotone in the right direction**: it must decrease as the correct class's score rises relative to the alternatives, so pushing the loss down pushes the model toward correct predictions. - Ideally it **upper-bounds the 0-1 loss**, so driving the surrogate to zero forces the error count to zero, and the surrogate value is a certificate on the error rate. Two classic choices for a binary score `s` with `y` in {-1, +1}: - **Hinge**: `max(0, 1 - y*s)`. At `y*s = 0` it equals 1, matching the 0-1 loss at the boundary; it sits above the step everywhere else in the wrong region. It has a constant-magnitude slope for violators and is exactly flat once the example clears the margin. - **Logistic / cross-entropy**: `log(1 + exp(-y*s))`. Smooth everywhere, never exactly zero, with a gradient of the familiar `p - y` form when written in probability terms. It also upper-bounds the 0-1 loss once rescaled by a constant. Both are *classification-calibrated*: with enough data and capacity, the score that minimises the surrogate makes the same sign decision as the one that minimises the error rate. That is the theoretical licence for the swap — you are not optimising a random proxy, you are optimising something whose optimum agrees with the metric's optimum. ## Where the two can still come apart Agreement in the limit is not agreement in your run. Cross-entropy keeps rewarding extra confidence on examples that are already correct, so a batch containing a few violently wrong, confidently-scored examples can dominate the gradient while accuracy barely moves. Under heavy class imbalance a model can drive cross-entropy down by sharpening its already-correct majority-class predictions and never learn the minority class at all. It is completely normal — and a good diagnostic habit — to watch validation loss and validation accuracy on the same plot and notice when they diverge: loss creeping up while accuracy holds steady usually means the model is getting more confident on the ones it gets wrong. ## The practical stance Train on the surrogate; select, threshold, and report on the metric. Model selection, early stopping and the decision threshold are all free to use the non-differentiable metric directly, because none of them needs a gradient — they only need to compare a handful of candidates. Reserve differentiating a bespoke metric-shaped loss for cases where that split genuinely fails you, because a custom objective is code someone has to review and defend for the life of the model.

  • If accuracy has no gradient, why does minimising cross-entropy usually raise accuracy anyway?
    Because cross-entropy is classification-calibrated: the score that minimises it makes the same decision as the score that minimises error rate, and it upper-bounds the 0-1 loss up to a constant. Pushing it down widens the score gap in the correct direction. The link is asymptotic, though — in a finite, imbalanced run the loss can improve while accuracy stalls, which is why you track both.
  • Are there ways to optimise a non-differentiable metric directly?
    Yes, with gradient-free search: grid or random search, evolutionary methods, or Bayesian optimisation can all optimise a metric they can only evaluate. They scale to tens of parameters, not to millions of weights, so the usual division of labour is gradient descent on a surrogate for the network and direct search for the few decision parameters, such as a threshold, sitting on top of it.
  • Your validation loss rises while validation accuracy stays flat. What is happening?
    The model is becoming more confidently wrong on the examples it already misclassifies without flipping any decisions. Cross-entropy grows without bound as a wrong prediction approaches certainty, while accuracy only counts sign changes. It is a warning about calibration and overfitting rather than about decision quality, and it argues for early-stopping on the metric you actually ship, not on the loss.

Accuracy is a staircase and cross-entropy is a ramp beside it. Standing on a step, you cannot feel which way is down; on the ramp you always can, and it takes you to the same floor.

saying these in an interview costs you the question

  • Says accuracy is merely too slow or expensive to compute each step
  • Claims accuracy has a gradient, just a very noisy one
  • Assumes lower cross-entropy always means higher accuracy
  • Cannot distinguish the reported metric from the trained objective
  • Thinks any differentiable function will do as a surrogate

context

open as a page

What does per-class loss weighting change when one intent has 200,000 training examples and others have 20?

level: juniorimportance: must knowfreq 70%

basics

~20 s

Per-class weighting multiplies each example's loss by a factor set by its class, so errors on rare intents contribute more gradient. It rebalances the effective class prior the model fits. It adds no new information about rare classes.

open as a page

How do you set the weights when a multi-task loss sums a depth error in metres and a segmentation cross-entropy?

level: middleimportance: must knowfreq 58%

basics

~20 s

Set weights so each term contributes comparable gradient, not so the raw numbers match. A depth error in metres dwarfs a cross-entropy in nats, so it dominates by accident of units. Tune the weights on per-task validation metrics, never on total loss.

open as a page

What is the difference between epistemic and aleatoric uncertainty in a model's predictions?

level: middleimportance: must knowfreq 70%

basics

~20 s

Aleatoric uncertainty is noise inherent in the data — a noisy sensor, identical inputs with different labels — and more data will not remove it. Epistemic uncertainty is the model's ignorance of regions it barely saw, and more data shrinks it.

open as a page

Why can a trained classifier output a 99.9% softmax score on pure noise?

level: middleimportance: must knowfreq 62%

basics

~20 s

Softmax normalizes scores across the trained classes, so its output ranks them and must sum to one — never evidence that the input belongs to any. Nothing in training penalised confident nonsense, and far from the data the gaps can grow.

open as a page

In focal loss with gamma = 2, what happens to an example already given probability 0.99 for its true class?

level: middleimportance: must knowfreq 60%

basics

~10 s

Focal loss multiplies cross-entropy by (1 - p_t)^gamma. At gamma = 2 and p_t = 0.99 that factor is 0.0001, so the example contributes one ten-thousandth of its cross-entropy and stops steering training.

open as a page

What does hinge loss's margin do that cross-entropy on the same scores does not?

level: middleimportance: should knowfreq 41%

basics

~20 s

Hinge loss demands a cushion, not just a correct sign: with max(0, 1 - y*s), an example scored 0.9 toward the right class still pays 0.1, and only past 1.0 does its gradient hit zero. Cross-entropy never stops pushing.

open as a page

Why attach an auxiliary loss head to an intermediate layer, and why anneal its coefficient toward zero?

level: middleimportance: should knowfreq 42%

basics

~20 s

An auxiliary head attached partway up gives the early layers a short gradient path and their own training signal, which eases optimisation of a deep stack. Its coefficient is annealed toward zero so the real head governs the final model.

open as a page

For a rare-positive classifier missing its F1 target, do you train soft-F1 or tune the threshold?

level: seniorimportance: should knowfreq 44%

basics

~20 s

Tune the threshold first: train with cross-entropy, then pick the operating point on a held-out split. A differentiable soft-F1 loss, built from summed probabilities instead of hard counts, is the fallback when no fixed threshold works.

open as a page

How would you get an epistemic uncertainty estimate out of an already-trained network?

level: seniorimportance: should knowfreq 48%

basics

~20 s

Sample several plausible models and measure their disagreement on the input. Keeping dropout active for thirty stochastic forward passes is the cheap option; five independently seeded networks are the stronger one. The spread, not the average, is the signal.

open as a page

Reweighting the loss versus oversampling the rare class: what is identical and what actually differs?

level: seniorimportance: should knowfreq 42%

basics

~10 s

Both target the same rebalanced objective and give the same expected gradient. They differ in gradient variance, mini-batch composition, what one epoch means, compute per pass, and how badly the rare rows get memorised.

open as a page

Two tasks' gradients meet a shared trunk at negative cosine similarity — how do you decide what to sacrifice?

level: principalimportance: should knowfreq 34%

basics

~20 s

Negative cosine between two task gradients means an update helping one hurts the other, and no choice of weights removes that. Name the task the product is judged on, fix the degradation you will accept elsewhere, then pick a mechanism to enforce it.

open as a page

Why can inverse-frequency class weights destabilise training, and what does effective-number weighting fix?

level: middleimportance: nice to knowfreq 30%

basics

~10 s

Inverse-frequency weights scale with the raw count ratio, so a 200,000-versus-20 split gives a 10,000-to-1 weight and wildly noisy updates. Effective-number weighting uses (1 - beta^n)/(1 - beta), which saturates and caps the ratio.

open as a page

How do you train a search ranker when the target metric NDCG has zero gradient?

level: seniorimportance: nice to knowfreq 27%

basics

~20 s

Change the unit of the loss from the item to the pair. NDCG reads only the sort order, so it is flat; a logistic loss on the score difference inside a should-outrank pair is smooth and pushes the order right.

open as a page

In uncertainty weighting of a multi-task loss, how are the per-task weights learned and what stops them collapsing to zero?

level: seniorimportance: nice to knowfreq 30%

basics

~20 s

Each task carries a learned scalar noise parameter, and its loss weight is the inverse of that scalar, so hard-to-fit tasks are downweighted automatically. An added log-noise penalty per task is what stops every weight sliding to zero.

open as a page

How do you set the abstention threshold for a chest-radiograph triage model that routes uncertain cases to a radiologist?

level: principalimportance: nice to knowfreq 34%

basics

~20 s

Treat it as a coverage-versus-risk trade, not a modelling choice. Sweep the threshold on held-out data, then pick the point where the deferred volume fits real radiologist capacity and the residual error is acceptable — and check that trade per subgroup.

open as a page