Why does a soft-margin SVM minimise hinge loss instead of the 0/1 error rate?
answer
- a step function has no useful slope
- convex upper bound on the count
- zero once the point clears the margin
- slack equals the loss at the optimum
- kink at margin one, use a subgradient
basics
~20 sThe 0/1 error rate is a step function: flat almost everywhere, so its gradient carries no direction and minimising it directly is intractable. Hinge loss, max(0, 1 - y*score), is a convex upper bound on it that an optimiser can actually descend.
solid answer
~50 sCode the label as `y = +1` or `-1` and let `f(x)` be the raw score. The 0/1 loss is `1` when `y*f(x) < 0` and `0` otherwise — piecewise constant, so its gradient is zero wherever it is defined and nothing tells the optimiser which way to move; minimising it directly over a linear model is non-convex and intractable. Hinge loss is `max(0, 1 - y*f(x))`: zero once the point clears the margin, then growing linearly as the point falls short. It is convex, it upper-bounds the 0/1 loss, and at the optimum each point's slack variable equals its hinge loss. Its flat region is the interesting part — a correctly classified point three margins away contributes exactly nothing to the objective or its subgradient, which is why the solution depends only on points on or inside the margin. The price is a kink at `y*f(x) = 1`, so solvers use a subgradient there.
code
python · 22 lines# hinge loss = max(0, 1 - y * f(x)); y is -1 or +1, f(x) is the raw score
points = [(+1, 2.30), (+1, 0.40), (-1, -1.70), (-1, 0.80), (+1, 1.00)]
total = 0.0
for y, f in points:
margin = y * f
slack = max(0.0, 1.0 - margin)
total += slack
if margin > 1.0:
role = "beyond the margin, contributes nothing"
elif margin == 1.0:
role = "exactly on the margin"
elif slack < 1.0:
role = "inside the margin, still correct"
else:
role = "misclassified"
print("y*f=%+.2f slack=%.2f %s" % (margin, slack, role))
C = 10.0
w_norm_sq = 4.0 # squared length of the weight vector, for illustration
print("total slack =", total)
print("objective =", 0.5 * w_norm_sq + C * total)go deeper
Be ready to write max(0, 1 - y*f(x)) with the label coded as plus or minus one, and to say it is zero once the point is correctly classified past the margin and grows linearly otherwise.
Explain the surrogate argument end to end: the error rate is piecewise constant with no usable gradient, hinge is convex and sits above it, and at the optimum each slack variable equals that point's hinge loss.
Demonstrate the consequences you have felt in practice: sparsity of the solution in the training points, the kink handled by a subgradient, and the fact that the score is uncalibrated so a downstream threshold or probability needs its own held-out step.
Own the choice of surrogate as a modelling decision. Argue when a saturating margin loss is the right objective for a business problem versus a loss that keeps penalising confident-but-wrong predictions, and what each choice implies about noisy labels.
## Two losses, side by side Code the label as `y = +1` or `y = -1` and let `f(x) = w . x + b` be the raw, unsquashed score. Define the **margin of a point** as `m = y * f(x)`: positive when the prediction is correct, and larger the more confidently correct it is. - **0/1 loss**: `1` if `m < 0`, else `0`. This is exactly what accuracy measures. - **Hinge loss**: `max(0, 1 - m)`. Zero once `m >= 1`, then rising linearly at slope 1 as `m` falls. Evaluate hinge on a few points to feel its shape. `m = 2.3` gives `0`. `m = 1.0` gives `0` — the point sits exactly on the margin. `m = 0.4` gives `0.6` — correctly classified, but inside the margin. `m = -0.8` gives `1.8` — misclassified, and charged more the deeper it is on the wrong side. ## Why the error rate cannot be the training objective The 0/1 loss is piecewise constant. Nudge `w` a little and either nothing changes or one point flips and the total jumps by a whole unit. Its derivative is zero everywhere it exists and undefined at the jumps, so a gradient tells the optimiser nothing about which direction reduces error. The summed 0/1 loss over a linear model is also non-convex with many flat plateaus and, in general, minimising it exactly is computationally intractable — there is no efficient algorithm that finds the error-minimising hyperplane on arbitrary non-separable data. The standard escape is a **convex surrogate**: replace the loss you care about with a convex function that sits above it, minimise the surrogate, and inherit a bound on the thing you wanted. Hinge loss qualifies. It is convex (a maximum of two affine functions), and it dominates the 0/1 loss everywhere: at `m = 0` the 0/1 loss is at most 1 and hinge is exactly 1; for `m < 0` hinge is `1 - m > 1`; for `0 < m < 1` hinge is positive while 0/1 is zero. So the average hinge loss upper-bounds the training error rate, and driving it down squeezes the error rate from above. That also explains the residual gap: the two objectives do not have the same minimiser. A model can trade one deep violation for several shallow ones, lowering total hinge while leaving the misclassification count flat or even slightly worse. ## The unit of margin, and why hinge is not just "relu of the error" Hinge charges nothing only when `m >= 1`, not when `m >= 0`. That `1` is the whole point: it demands a *buffer*, not merely a correct sign. Because `w` can be rescaled freely, fixing the required margin at 1 and letting `||w||` float is what turns "be confidently right" into the concrete penalty `0.5*||w||^2`. The full objective ``` 0.5*||w||^2 + C * sum_i max(0, 1 - y_i * f(x_i)) ``` is therefore an ordinary regularised empirical-risk problem: an L2 penalty on the weights plus a convex data-fitting loss, with `C` weighting the second against the first. The slack variables in the constrained form are not an extra modelling choice — at the optimum, `slack_i` equals exactly `max(0, 1 - m_i)`, the point's hinge loss. Three regimes follow directly: `slack = 0` means the point is on or beyond the margin; `0 < slack <= 1` means it is inside the margin but still on the correct side; `slack > 1` means it is misclassified. ## Consequences of the flat region The zero region is where hinge earns its keep. A point with `m = 3` contributes zero to the objective and zero to the subgradient. Delete it from the training set and the solution does not move. Only points with `m <= 1` — those on or inside the margin — determine the fitted weights, so the solution is **sparse in the training data**. This is a genuine difference from losses that stay strictly positive for every point and therefore let every observation, however comfortably classified, exert a pull on the boundary. The cost is a kink at `m = 1`, where hinge is continuous but not differentiable. Solvers handle this with a subgradient: any value in `[-1, 0]` times `y*x` works at the kink, and the objective remains convex, so convergence guarantees survive. ## What hinge loss does not give you It is not a proper scoring rule. Because it saturates at zero, it stops caring how confidently a point is classified once it clears the margin, so the raw score is a signed distance-like quantity, not a probability. If a calibrated probability is required, fit a monotone map from score to probability on held-out data as a separate step; reading the raw score as a probability is a common and costly mistake. Hinge also grows only linearly in the violation, so a single far-out mislabelled point is charged proportionally rather than quadratically. That makes it less explosive than a squared penalty on the same violation, but it is still unbounded: one badly wrong label placed far across the boundary can still drag the solution, especially at large `C`.
- What is the relationship between a point's slack variable and its hinge loss?At the optimum they are the same number: slack equals max(0, 1 - y*f(x)). Nothing is gained by making a slack larger than it has to be, so the constrained and unconstrained forms describe the same problem. Reading the value tells you the regime: zero means on or beyond the margin, between zero and one means inside it but correct, above one means misclassified.
- Why does hinge loss make the solution depend on only a subset of the training points?Its flat region gives exactly zero loss and zero subgradient contribution for any point with y*f(x) > 1. Those points can be deleted without moving the fitted weights, so only points on or inside the margin shape the boundary. A loss that stays strictly positive everywhere never has that property — every observation keeps some pull.
- Can you read the raw SVM score as a probability of the positive class?No. Hinge loss is not a proper scoring rule; it stops penalising a point once it clears the margin, so the score conveys a signed distance rather than a calibrated likelihood. If you need probabilities, fit a monotone map from score to probability on held-out data as a separate step, and validate the calibration on data neither step has seen.
- How do solvers cope with hinge loss being non-differentiable?The kink is at y*f(x) = 1 and the function is convex, so a subgradient stands in for the gradient there — any slope between the flat piece and the linear piece is a valid descent direction. Convexity means the usual convergence guarantees still hold; only the smoothness assumption is lost.
The 0/1 error is a pass/fail stamp: it tells you that you missed, never by how much. Hinge loss reports how far you missed the mark by, which is the only thing that tells an optimiser which way to move.
saying these in an interview costs you the question
- Says minimising hinge loss and minimising the error rate give the same optimum
- Claims the 0/1 error rate can be optimised by gradient descent
- Thinks hinge loss is differentiable everywhere
- Says hinge charges zero as soon as the point is on the correct side
- Treats the raw margin score as a probability