Would you fit one multinomial softmax model or 40 one-vs-rest models for 40 categories?
answer
- one fit versus many independent fits
- who shares a normaliser and who does not
- 40, or 40, or 780 models
- independent scores need renormalising
- adding a class refits everything, or just one
basics
~20 sOne multinomial fit trains all 40 classes jointly through a shared normaliser, so its probabilities already sum to one. Forty one-vs-rest fits are independent, so their scores need renormalising, while one-vs-one would need 780 models.
solid answer
~50 sDefault to the single multinomial fit when the learner supports it. Multinomial softmax optimises all 40 classes against one shared normaliser, so raising one class's score directly lowers the others, and the output is a genuine distribution over the 40 leaf categories. One-vs-rest instead trains 40 independent binary models — category k against everything else — each seeing a badly imbalanced problem, and their 40 scores have no reason to sum to one; you have to renormalise before calling them probabilities, and after renormalising they are usually worse calibrated. One-vs-one would fit `k(k-1)/2 = 780` pairwise models and vote, which is a lot of models for 40 classes though each trains on a small subset. I would pick one-vs-rest only for a concrete reason: the base learner has no multinomial form, or I need to add and retrain single categories independently.
go deeper
Be able to describe the two constructions in plain words: one model that scores all 40 categories together, versus 40 yes-or-no models run in parallel and compared.
Explain the shared normaliser and what it does — the classes compete for one unit of probability — and why 40 independently fitted scores need renormalising before they can be read as a distribution.
Bring the operational tradeoff: parallel training and per-class retraining on a churning taxonomy versus coherent, better-calibrated probabilities from a single fit, and say how you would measure which loss actually hurts the product.
Own the taxonomy-versus-model coupling as a design decision: how often categories change, whether the platform can afford a full refit each time, and whether downstream consumers depend on probabilities being coherent across classes.
## Three ways to get K classes out of a binary-shaped world With 40 marketplace leaf categories there are three standard constructions. **Multinomial (softmax) fit — one model.** The model holds one weight vector per class, produces 40 logits, and converts them with `p_k = exp(z_k) / sum_j exp(z_j)`. It is trained by minimising multiclass cross-entropy over all 40 classes at once. **One-vs-rest (OvR) — 40 models.** For each category k, train a binary classifier: "is this listing category k, yes or no?", pooling the other 39 categories into the negative class. At prediction time you run all 40 and take the highest score. **One-vs-one (OvO) — 780 models.** Train one binary classifier for every pair of categories, `40 * 39 / 2 = 780` of them, each on only the rows belonging to those two categories. At prediction time every model casts a vote and the category with the most votes wins. ## What the multinomial fit buys you **Coupling.** The shared denominator means the 40 classes compete. During training, the gradient for the true class pushes its logit up while every other class's logit is pushed down; the classes are fit against each other rather than each against a blur of "everything else". That is usually the right inductive bias when the categories are genuinely mutually exclusive. **Probabilities that already sum to one.** No post-processing. You can threshold, abstain, or aggregate sibling categories by adding their probabilities and the arithmetic is meaningful. **One optimisation, one set of hyperparameters.** One regularisation strength, one convergence check, one artefact to ship and monitor. ## What one-vs-rest costs you **Scores that do not sum to one.** Each of the 40 binary models was fit in ignorance of the others. Consider a soil-type classifier over 7 mutually exclusive soil classes: the seven one-vs-rest models might output `0.7, 0.6, 0.2, 0.1, 0.1, 0.05, 0.05` — a total of 1.8. You can divide by the total to force a distribution, but that renormalisation is a patch, not an estimate: nothing in the fitting procedure ever tried to make the seven numbers coherent, and the renormalised values are typically less well calibrated than a multinomial fit's. **Per-model imbalance.** With 40 roughly balanced categories, every one-vs-rest model faces a 1-vs-39 problem, so each binary fit sees around 2.5% positives. That is a harder, more skewed problem than the multinomial fit ever sees, and it can shift each model's operating point in a way that makes the 40 scores non-comparable at prediction time. **Argmax over incomparable scores.** Taking the highest of 40 independently-calibrated scores assumes the scores are on a common scale. They are not, in general — one model may be systematically optimistic. This is the practical failure mode of OvR: it works, but the winner is decided by whichever model happens to be most enthusiastic. ## What one-vs-rest buys you **Independence, and that is genuinely useful.** The 40 fits are embarrassingly parallel. A new leaf category can be added by training one more binary model, leaving the existing 39 untouched — with a multinomial fit, adding a class changes the normaliser and requires refitting everything. Similarly, a single misbehaving category can be retrained in isolation. If the taxonomy churns weekly, that operational property can outweigh the calibration loss. **It works with any binary learner.** Some learners have no natural multiclass objective at all; one-vs-rest is the generic wrapper that gives them one. ## Where one-vs-one fits The 780 pairwise models for 40 classes look alarming, but each one trains on only the rows of two categories — roughly 1/20 of the data each if the classes are balanced — so total training work is not 780 times a full fit. Its real advantages are that each subproblem is balanced and often much easier to separate than "this class versus everything". Its real costs are the model count to store and evaluate, the votes being harder to turn into probabilities, and ties in the voting. For 40 classes with a learner that has a native multinomial form, it is rarely the right call. ## How to answer in an interview Say the default first — one multinomial fit, because the classes are mutually exclusive and you want a coherent distribution over them — then name the specific conditions that would flip your choice: a base learner with no multinomial objective, a taxonomy where categories are added and retrained constantly, or a need to parallelise fits across machines. Do not present one-vs-rest as simply "the same thing, decomposed": the decomposition changes what is being optimised, and the scores it returns are not probabilities until you renormalise them, and only loosely so afterwards.
- A new leaf category is added to the taxonomy every week. Does that change your answer?It pushes toward one-vs-rest. Adding a class to a multinomial model changes the shared normaliser, so the whole model must be refit and revalidated. With one-vs-rest you train one more binary model and leave the other 40 alone. You pay for it in calibration, so I would still measure whether the renormalised probabilities are good enough for the downstream decision.
- Why does each one-vs-rest model face an imbalance problem that the multinomial fit does not?Each one-vs-rest model pools the other 39 categories into a single negative class, so with balanced categories it sees roughly 2.5% positives. The multinomial fit never constructs that pooled negative: it compares the 40 classes directly against each other through one normaliser, so no subproblem is skewed by the decomposition itself.
- If you must use one-vs-rest, how do you turn the 40 scores into something you can threshold?Divide each score by the total so they sum to one, but treat that as a rescaling rather than an estimate. Then check calibration on held-out data — for example whether listings given probability 0.8 are correct about 80% of the time — and refit a calibration mapping if they are not. Never assume renormalising alone made them probabilities.
saying these in an interview costs you the question
- Claims one-vs-rest scores automatically sum to one
- Says the two approaches optimise the same objective
- Ignores that a new class forces a full multinomial refit
- Thinks 40 classes need 40 one-vs-one models
- Assumes renormalised one-vs-rest scores are calibrated