In a one-vs-one SVM router, two classes tie on votes — why not compare margin scores?
answer
- the raw output is a signed distance
- each pair has its own weight norm
- different pairs, different rulers
- hinge loss is not a proper scoring rule
- fit a sigmoid on held-out folds
basics
~20 sEach pairwise SVM produces a signed distance measured in its own weight-norm units, learned from its own two classes. Those numbers share no common scale and are not probabilities, so comparing them across pairs is arbitrary.
solid answer
~50 sA pairwise SVM's raw output is `w·x + b` — a signed distance to that pair's hyperplane, scaled by that pair's `||w||`. Each of the 780 models was fit on a different subset with different separability and different class sizes, so 1.4 from one pair and 0.9 from another carry no comparable meaning. Nor is either a probability: hinge loss is not a proper scoring rule, so the value is at best monotone in confidence within one pair. Summing decision values is the common tiebreak and works acceptably, but it is a heuristic with nothing behind it. The principled fix is a calibration map per pairwise model, fitted on held-out data — Platt scaling fits a sigmoid from decision value to probability — and then combined into class probabilities. Note that the calibrated winner can differ from the vote winner.
go deeper
Know that an SVM outputs a signed distance to its boundary, not a probability, and that a bigger number is not automatically a more confident prediction when it comes from a different model.
Explain why the scales differ: each pairwise model has its own weight norm set by how separable that pair was, and hinge loss gives the output no probabilistic meaning to begin with.
Demonstrate having shipped one. Talk about where the tie rule lives, calibrating on held-out folds rather than training data, and the fact that a calibrated winner can contradict the vote winner.
Own the product answer. Decide whether the router should emit a single category at all, and design the abstention and review path so that ambiguous cases are escalated rather than resolved by an arbitrary rule.
## What the number actually is Every binary SVM outputs `f(x) = w·x + b` before the sign is taken. Geometrically, `f(x)/||w||` is the signed perpendicular distance from the point to that model's hyperplane; the raw `f(x)` is that distance multiplied by that model's `||w||`. In a pairwise decomposition over 40 categories there are 780 such functions, each with its own `w`, its own `b`, and its own support vectors. That is where the comparison breaks. `||w||` is set by how separable that particular pair happened to be — a pair of visually distinct categories yields a wide margin and a small norm, a pair of near-duplicates yields a cramped margin and a large one. So a decision value of 1.4 from the "phone accessories vs cables" model and 0.9 from the "kitchenware vs cookware" model are two numbers on two different rulers. Ranking them picks the model with the bigger ruler, not the more confident prediction. Three further sources of incomparability stack on top: - **Different training sets.** Each pairwise model saw only its two categories' rows, with different sizes and different class balance, so the same nominal `C` regularises each of them differently in effect. - **Different noise levels.** A pair that overlaps heavily produces small decision values for almost every point; a cleanly separated pair produces large ones for almost every point. The spread differs before you look at any individual prediction. - **The loss is not probabilistic.** Hinge loss is not a proper scoring rule — it is optimised by getting points past the margin, not by reporting the right frequency. Nothing in training pushes the decision value toward anything with a probabilistic interpretation, unlike a log-loss-trained model whose output does estimate a probability by construction. ## Why ties happen at all Pairwise voting counts, for each incoming listing, how many of the 39 duels each category won. Ties are ordinary at 40 categories, and worse, the pairwise outcomes need not be transitive: A beats B, B beats C, C beats A. There is no total ordering to fall back on, because the voting scheme never produced one — it produced a tournament. ## The tiebreaks, in order of principle **Sum the decision values.** The common practical rule: for each category, add up the (signed) decision values of the duels it took part in, and take the largest. It works acceptably in practice and costs nothing, but it is exactly the comparison across rulers described above, so treat it as a heuristic, not a fix. **Calibrate each pairwise model, then combine.** Fit a mapping from decision value to probability for each pairwise classifier on held-out data — the standard method is Platt scaling, which fits a sigmoid `1/(1 + exp(a*f(x) + b))` with the two parameters estimated on data the SVM did not train on. Fitting the sigmoid on the training data itself produces badly optimistic estimates, so cross-validated folds are required. Once each duel reports a probability instead of a distance, the pairwise probabilities can be combined into a distribution over the 40 categories. This is the principled route, at the price of an extra fitting stage per pairwise model and a warning: the class with the highest calibrated probability is not always the class that won the vote, so the two ways of reading the model can disagree. **Design the tie away.** Often the best senior answer. A marketplace router does not have to emit one category. It can return the top three for a seller to confirm, or abstain when the vote is tied or the margin between the top two is thin and send the listing to human review. If the downstream system has a review queue at all, converting ambiguity into an explicit abstention is more valuable than any tiebreak rule, because the tied cases are exactly the ones where an arbitrary choice is most likely to be wrong. ## What this implies for monitoring Because the raw scores are not comparable, a threshold like "only auto-route when the score exceeds 1.0" cannot be set once and applied across categories — it means something different for every pair. Any confidence threshold in a router built this way needs to be defined on calibrated outputs, or set per category from held-out data. ## The short version A margin is a geometric quantity in per-model units, not a confidence. Vote ties reveal that the decomposition never produced a comparable score in the first place, and the choices are a cheap heuristic, a calibration stage, or an abstention path — not a rescaling of numbers that were never on the same scale.
- What is the usual practical tiebreak, and why is it only a heuristic?Summing the signed decision values across the duels each class took part in and taking the largest. It behaves reasonably, but it adds numbers produced by separately trained models with different weight norms and different separability, so nothing guarantees the sum reflects confidence. It is a convention that works, not a justified rule.
- How would you get a real probability out of a pairwise SVM?Fit a calibration map from decision value to probability for each pairwise model — Platt scaling fits a two-parameter sigmoid — and estimate it on held-out folds rather than the training data, which would give badly optimistic estimates. Then combine the pairwise probabilities into a distribution over classes. Expect the calibrated argmax to occasionally disagree with the vote winner.
- Can you avoid needing a tiebreak at all?Often, yes, and it is usually the better system design. Return the top few categories for the seller to confirm, or abstain when the top two are close and route the listing to human review. Tied cases are precisely where an arbitrary pick is most likely wrong, so converting ambiguity into an explicit escalation beats inventing a rule.
It is like ranking athletes by how far ahead they finished, when every race was a different length on a different track. The margins are real, but they are not on one scale.
saying these in an interview costs you the question
- Treats the decision value as a probability
- Compares raw scores across separately trained models
- Sets one confidence threshold for all pairs
- Fits the calibration sigmoid on the training data
- Assumes pairwise vote outcomes must be transitive