skip to content

Heads and Training Loop

The last layer that turns hidden activations into logits, probabilities or a bare number, and the epoch-batch-step loop that consumes them. Interviewers ask which head fits which task.

on this pageshow

explore

questions

11

Why use one sigmoid per label instead of a softmax over the same output layer?

level: juniorimportance: must knowfreq 80%

answer

  1. Ask whether two labels can co-occur
  2. Independent decisions or one shared budget
  3. Multi-hot targets, not a class index
  4. Outputs need not sum to one
  5. Threshold each output, do not argmax

basics

~20 s

Per-label sigmoids fit tasks where one example can carry several labels at once: each output is its own independent probability and the outputs need not sum to one. A softmax makes the classes share a single probability budget, which only fits mutually exclusive classes.

solid answer

~40 s

The choice follows the label structure of the task, not the architecture. If exactly one label is true per example, a softmax head is right. If labels can co-occur, you want one output unit per label with a sigmoid applied to each, so the head emits N independent probabilities that do not have to sum to one. A chest-radiograph findings head is the standard case: one scan can be simultaneously cardiomegaly, effusion and atelectasis, and forcing those through a softmax would make each finding suppress the others as the shared budget gets split. Targets become a multi-hot vector rather than a single class index, and inference changes too: you threshold each output independently instead of taking one argmax. The output width is the same in both designs; what differs is whether the units compete.

code

python · 17 lines
python
import math

# three findings on one chest radiograph, raw head outputs (logits)
logits = [2.0, 1.0, -1.0]

# per-label sigmoid head: each label independent
sigmoid = [1 / (1 + math.exp(-z)) for z in logits]

# softmax head over the same logits: one shared budget
m = max(logits)
exps = [math.exp(z - m) for z in logits]
softmax = [e / sum(exps) for e in exps]

print([round(p, 3) for p in sigmoid])   # [0.881, 0.731, 0.269]
print(round(sum(sigmoid), 3))           # 1.881  -> not a distribution
print([round(p, 3) for p in softmax])   # [0.705, 0.259, 0.035]
print(round(sum(softmax), 3))           # 1.0    -> forced to share

go deeper

for a junior

Recall the one-line rule and be able to apply it to an example: labels that can co-occur get one sigmoid each, labels that are mutually exclusive get a softmax. Know that sigmoid outputs are not required to sum to one.

for a middle

Explain what changes beyond the squashing function — multi-hot targets, a per-label objective, thresholding instead of argmax, and per-label metrics instead of accuracy. Be able to say why a softmax cannot represent three simultaneous findings.

for a senior

Show you check the label structure against the data rather than the schema, and that you notice when apparent exclusivity is an artefact of an annotation guideline. Be ready to describe how you would migrate a mislabelled softmax head to a per-label head without breaking downstream consumers.

for a principal

Own the framing that head design encodes a claim about the world's label structure, and that claim outlives any one model. Argue when to enforce a one-of-N constraint architecturally versus learn it from data, and what it costs to change that decision later.

## What the head is A network's *head* is the final layer that converts the trunk's hidden activations into whatever the task needs to emit. For classification that means a linear layer producing one raw score (a *logit*) per label, followed by a squashing function that turns logits into probabilities. The interesting design decision is which squashing function, and that decision is determined by the label structure of the data, not by the size of the network. ## Two shapes of label structure **Mutually exclusive (single-label).** Exactly one label is true per example — a photo is a cat or a dog or a bird, never two. Here a softmax over the whole output vector is correct: it produces a proper distribution over the classes, one shared probability budget of 1.0 divided among them. **Multi-label.** Several labels can be true simultaneously. A chest radiograph can show cardiomegaly, effusion and atelectasis at once; it can also show none of them. Here the right head is *N independent binary decisions*: one output unit per label, with the logistic sigmoid `p_i = 1 / (1 + exp(-z_i))` applied to each unit's logit separately. Each `p_i` answers the question "is label i present?" without reference to the others. The N probabilities can sum to 0.2 or to 4.7 — that is not an error, because they are not a distribution over one variable, they are N separate Bernoulli parameters. ## Why the softmax is wrong for co-occurring labels A softmax normalises across the output vector: raising one logit necessarily lowers every other probability. That coupling is exactly what you want when the classes compete for one true answer, and exactly what you do not want when findings co-occur. With three findings genuinely present, the best a softmax head can do is roughly 0.33 each, which no downstream threshold can turn back into "all three are present". The all-negative case is worse: a softmax must place probability 1.0 somewhere, so a scan with no findings still gets confidently assigned one, unless you add an explicit "no finding" class and lose the ability to express partial certainty about the real ones. Per-label sigmoids handle both cases naturally — all outputs low means nothing detected, several high means several present. ## What changes end-to-end 1. **Targets.** Single-label training uses an integer class index (or a one-hot row). Multi-label training uses a multi-hot vector: `[1, 0, 1, 1, 0, ...]`, one 0/1 per label, with as many ones as apply. 2. **Output layer.** The linear layer is the same shape either way — width equals the number of labels. Only the squashing function and the target format change. 3. **Objective.** A per-label head is trained with the standard per-label binary objective summed or averaged across labels; the details of that objective belong to the loss discussion, but the key structural point is that each label contributes its own term. 4. **Inference.** Single-label heads take the argmax. Per-label heads compare each probability to a threshold and emit the set of labels that clear it — which is why threshold choice becomes its own design problem. 5. **Metrics.** You stop reporting plain accuracy and start reporting per-label precision/recall or set-level measures, because "the prediction" is now a set rather than one class. ## Mixed and hierarchical cases Real label sets are often part exclusive, part independent. A product image might have exactly one *category* and any number of *attributes*. The clean answer is multiple heads off one trunk: a softmax head over the exclusive group and a sigmoid head over the independent group, each with its own target format. Cramming both groups into one squashing function is the usual mistake — a softmax over the union makes independent attributes compete, and a single sigmoid bank over the union throws away the known constraint that exactly one category is true, letting the model predict two categories or none. ## The diagnostic Ask one question of the data: *can two labels be true for the same example?* If yes, independent sigmoids. If no — and you are confident the exclusivity is a real property of the world rather than an artefact of how the data was collected — a softmax, and you get the exclusivity constraint enforced for free.

  • The output layer has the same width either way — so what actually differs?
    Only the squashing function and everything downstream of it. A sigmoid applies elementwise, so each logit maps to its own probability; a softmax normalises across the vector, so the logits are coupled. That difference propagates into the target format (multi-hot versus class index), the objective, and inference (per-label thresholds versus a single argmax).
  • How would you handle a label set that is one exclusive category plus many independent attributes?
    Two heads off the same trunk: a softmax head over the category group, whose exclusivity is a real constraint worth enforcing, and a sigmoid head over the attribute group. Each head gets its own target format. Merging the groups into one squashing function either makes independent attributes compete or discards the known one-of-N constraint.
  • If a multi-label sigmoid head outputs all values below 0.1, is something wrong?
    Not necessarily — that is the head's legitimate way of saying no label applies, which a softmax head cannot express. It is only a problem if you expected positives, in which case suspect heavy label sparsity pushing every unit's bias negative, or an evaluation cutoff that is simply too high for these score distributions.

A softmax is a single ballot where voters must pick one candidate; per-label sigmoids are a checklist where each box is ticked or not, independently of the others.

saying these in an interview costs you the question

  • Claims sigmoid outputs must sum to one like a softmax
  • Uses softmax with top-k selection for genuinely co-occurring labels
  • Says softmax and sigmoid heads differ only in output width
  • Takes an argmax over a multi-label sigmoid head at inference
  • Cannot state that a multi-label target is a multi-hot vector

context

open as a page

Why must a softmax classification head have exactly one output unit per class?

level: juniorimportance: must knowfreq 78%

basics

~20 s

A softmax head emits one score per class and normalizes those scores into probabilities that sum to one. K classes need K units: with fewer, some class can never be predicted; extra units create classes no label ever selects.

open as a page

With 50,000 training examples and batch size 128, how many iterations and updates do 3 epochs take?

level: juniorimportance: must knowfreq 78%

basics

~20 s

An epoch is one full pass over the data. At batch size 128, 50,000 examples give 390 full batches plus a ragged batch of 80, so 391 iterations per epoch, one update each, and 1,173 updates over 3 epochs.

open as a page

Narrate one training step for a batch of 64 — which tensors change and which do not?

level: middleimportance: must knowfreq 68%

basics

~20 s

The forward pass turns the 64 inputs into activations and one loss number, the backward pass fills a gradient for every parameter, and the update writes new parameter values. Only parameters, optimizer state and gradient buffers change; the batch, targets and architecture do not.

open as a page

Your resale-price regression head predicts -30 for cheap items — how do you fix it?

level: middleimportance: should knowfreq 55%

basics

~20 s

A plain linear output head is unbounded by construction, so negative predictions are expected, not a bug. Fix it in the head: predict in log space and exponentiate, pass the output through a positive-valued transform, or accept the negatives and clip only at serving.

open as a page

Can you skip the softmax at inference and take the argmax of a classifier's raw scores?

level: middleimportance: should knowfreq 48%

basics

~20 s

Yes, for a top-1 or top-k label. Softmax exponentiates each score and divides them all by the same positive total, so it never changes their order. You need the normalized values only when something downstream consumes the number itself.

open as a page

How do you choose decision thresholds for a 50-tag multi-label audio tagging head?

level: seniorimportance: should knowfreq 52%

basics

~20 s

Tune one threshold per tag on a held-out set rather than applying 0.5 everywhere. Tags differ in base rate, score distribution and the cost of a mistake, so 'live recording' may fire best at 0.2 while 'acoustic guitar' needs 0.7.

open as a page

A deployed 20-way product classifier needs a 21st category. What happens to the softmax head?

level: seniorimportance: should knowfreq 58%

basics

~20 s

The head gains one weight row and one bias for the new class; the trunk keeps its shape. All 21 classes then share one probability budget, so old outputs shift and old thresholds no longer transfer.

open as a page

What breaks when a validation pass is run in training mode instead of evaluation mode?

level: seniorimportance: should knowfreq 60%

basics

~20 s

Dropout stays active, so every validation number is a noisy sample of a randomly thinned network, and normalization layers use the validation batch's own statistics and refresh their stored ones. The metric becomes irreproducible, batch-order dependent, and validation data leaks into the model.

open as a page

Would you serve one shared-trunk model with steering and brake heads, or two separate models?

level: principalimportance: nice to knowfreq 36%

basics

~20 s

Share a trunk when the two tasks need the same perception features and you want one forward pass; keep them separate when they need independent release cadences. Sharing buys compute and transfer, and costs you a coupled artefact where retraining for one task can regress the other.

open as a page

Your model sees a 2-billion-token stream once — how do you plan and report the run without epochs?

level: principalimportance: nice to knowfreq 38%

basics

~20 s

Count the run in optimizer steps and tokens consumed, not epochs — an epoch that happens once carries no progress information. Fix a step budget up front, hang evaluation and checkpointing on step intervals, and compare runs at equal tokens seen.

open as a page