Why use one sigmoid per label instead of a softmax over the same output layer?
answer
- Ask whether two labels can co-occur
- Independent decisions or one shared budget
- Multi-hot targets, not a class index
- Outputs need not sum to one
- Threshold each output, do not argmax
basics
~20 sPer-label sigmoids fit tasks where one example can carry several labels at once: each output is its own independent probability and the outputs need not sum to one. A softmax makes the classes share a single probability budget, which only fits mutually exclusive classes.
solid answer
~40 sThe choice follows the label structure of the task, not the architecture. If exactly one label is true per example, a softmax head is right. If labels can co-occur, you want one output unit per label with a sigmoid applied to each, so the head emits N independent probabilities that do not have to sum to one. A chest-radiograph findings head is the standard case: one scan can be simultaneously cardiomegaly, effusion and atelectasis, and forcing those through a softmax would make each finding suppress the others as the shared budget gets split. Targets become a multi-hot vector rather than a single class index, and inference changes too: you threshold each output independently instead of taking one argmax. The output width is the same in both designs; what differs is whether the units compete.
code
python · 17 linesimport math
# three findings on one chest radiograph, raw head outputs (logits)
logits = [2.0, 1.0, -1.0]
# per-label sigmoid head: each label independent
sigmoid = [1 / (1 + math.exp(-z)) for z in logits]
# softmax head over the same logits: one shared budget
m = max(logits)
exps = [math.exp(z - m) for z in logits]
softmax = [e / sum(exps) for e in exps]
print([round(p, 3) for p in sigmoid]) # [0.881, 0.731, 0.269]
print(round(sum(sigmoid), 3)) # 1.881 -> not a distribution
print([round(p, 3) for p in softmax]) # [0.705, 0.259, 0.035]
print(round(sum(softmax), 3)) # 1.0 -> forced to sharego deeper
Recall the one-line rule and be able to apply it to an example: labels that can co-occur get one sigmoid each, labels that are mutually exclusive get a softmax. Know that sigmoid outputs are not required to sum to one.
Explain what changes beyond the squashing function — multi-hot targets, a per-label objective, thresholding instead of argmax, and per-label metrics instead of accuracy. Be able to say why a softmax cannot represent three simultaneous findings.
Show you check the label structure against the data rather than the schema, and that you notice when apparent exclusivity is an artefact of an annotation guideline. Be ready to describe how you would migrate a mislabelled softmax head to a per-label head without breaking downstream consumers.
Own the framing that head design encodes a claim about the world's label structure, and that claim outlives any one model. Argue when to enforce a one-of-N constraint architecturally versus learn it from data, and what it costs to change that decision later.
## What the head is A network's *head* is the final layer that converts the trunk's hidden activations into whatever the task needs to emit. For classification that means a linear layer producing one raw score (a *logit*) per label, followed by a squashing function that turns logits into probabilities. The interesting design decision is which squashing function, and that decision is determined by the label structure of the data, not by the size of the network. ## Two shapes of label structure **Mutually exclusive (single-label).** Exactly one label is true per example — a photo is a cat or a dog or a bird, never two. Here a softmax over the whole output vector is correct: it produces a proper distribution over the classes, one shared probability budget of 1.0 divided among them. **Multi-label.** Several labels can be true simultaneously. A chest radiograph can show cardiomegaly, effusion and atelectasis at once; it can also show none of them. Here the right head is *N independent binary decisions*: one output unit per label, with the logistic sigmoid `p_i = 1 / (1 + exp(-z_i))` applied to each unit's logit separately. Each `p_i` answers the question "is label i present?" without reference to the others. The N probabilities can sum to 0.2 or to 4.7 — that is not an error, because they are not a distribution over one variable, they are N separate Bernoulli parameters. ## Why the softmax is wrong for co-occurring labels A softmax normalises across the output vector: raising one logit necessarily lowers every other probability. That coupling is exactly what you want when the classes compete for one true answer, and exactly what you do not want when findings co-occur. With three findings genuinely present, the best a softmax head can do is roughly 0.33 each, which no downstream threshold can turn back into "all three are present". The all-negative case is worse: a softmax must place probability 1.0 somewhere, so a scan with no findings still gets confidently assigned one, unless you add an explicit "no finding" class and lose the ability to express partial certainty about the real ones. Per-label sigmoids handle both cases naturally — all outputs low means nothing detected, several high means several present. ## What changes end-to-end 1. **Targets.** Single-label training uses an integer class index (or a one-hot row). Multi-label training uses a multi-hot vector: `[1, 0, 1, 1, 0, ...]`, one 0/1 per label, with as many ones as apply. 2. **Output layer.** The linear layer is the same shape either way — width equals the number of labels. Only the squashing function and the target format change. 3. **Objective.** A per-label head is trained with the standard per-label binary objective summed or averaged across labels; the details of that objective belong to the loss discussion, but the key structural point is that each label contributes its own term. 4. **Inference.** Single-label heads take the argmax. Per-label heads compare each probability to a threshold and emit the set of labels that clear it — which is why threshold choice becomes its own design problem. 5. **Metrics.** You stop reporting plain accuracy and start reporting per-label precision/recall or set-level measures, because "the prediction" is now a set rather than one class. ## Mixed and hierarchical cases Real label sets are often part exclusive, part independent. A product image might have exactly one *category* and any number of *attributes*. The clean answer is multiple heads off one trunk: a softmax head over the exclusive group and a sigmoid head over the independent group, each with its own target format. Cramming both groups into one squashing function is the usual mistake — a softmax over the union makes independent attributes compete, and a single sigmoid bank over the union throws away the known constraint that exactly one category is true, letting the model predict two categories or none. ## The diagnostic Ask one question of the data: *can two labels be true for the same example?* If yes, independent sigmoids. If no — and you are confident the exclusivity is a real property of the world rather than an artefact of how the data was collected — a softmax, and you get the exclusivity constraint enforced for free.
- The output layer has the same width either way — so what actually differs?Only the squashing function and everything downstream of it. A sigmoid applies elementwise, so each logit maps to its own probability; a softmax normalises across the vector, so the logits are coupled. That difference propagates into the target format (multi-hot versus class index), the objective, and inference (per-label thresholds versus a single argmax).
- How would you handle a label set that is one exclusive category plus many independent attributes?Two heads off the same trunk: a softmax head over the category group, whose exclusivity is a real constraint worth enforcing, and a sigmoid head over the attribute group. Each head gets its own target format. Merging the groups into one squashing function either makes independent attributes compete or discards the known one-of-N constraint.
- If a multi-label sigmoid head outputs all values below 0.1, is something wrong?Not necessarily — that is the head's legitimate way of saying no label applies, which a softmax head cannot express. It is only a problem if you expected positives, in which case suspect heavy label sparsity pushing every unit's bias negative, or an evaluation cutoff that is simply too high for these score distributions.
A softmax is a single ballot where voters must pick one candidate; per-label sigmoids are a checklist where each box is ticked or not, independently of the others.
saying these in an interview costs you the question
- Claims sigmoid outputs must sum to one like a softmax
- Uses softmax with top-k selection for genuinely co-occurring labels
- Says softmax and sigmoid heads differ only in output width
- Takes an argmax over a multi-label sigmoid head at inference
- Cannot state that a multi-label target is a multi-hot vector