When does a network's output head need binary cross-entropy rather than categorical cross-entropy?
answer
- ask whether labels can co-occur
- one budget of probability, or K separate ones
- sum over units versus normalise over classes
- independent Bernoullis versus one categorical
basics
~20 sUse binary cross-entropy when each output is an independent yes/no decision, so several can be true at once. Use categorical cross-entropy when exactly one of the classes is correct and the outputs must compete for a single unit of probability.
solid answer
~50 sThe question is whether the labels are mutually exclusive. Categorical cross-entropy is `-log p_c`, where `p` is a softmax over the class axis: the probabilities sum to one, so pushing one class up necessarily pushes the others down. That is right when exactly one label is true - one spoken digit, one phoneme, one language. Binary cross-entropy is `-(y*log p + (1-y)*log(1-p))` applied independently to each output unit, each squashed by its own sigmoid, and the per-unit losses are summed. Nothing couples the units, so an example can carry three true tags and no true tags equally well. So a multi-label tagger gets K sigmoid outputs and summed binary cross-entropy; a single-winner classifier gets one softmax and categorical cross-entropy. Both are the same negative-log-likelihood idea under different likelihood assumptions: K independent Bernoullis versus one categorical.
go deeper
Be ready to state the rule in one sentence and name the activation that goes with each loss: sigmoid per output for binary cross-entropy, softmax across classes for categorical cross-entropy.
Expect to write both formulas and explain the coupling - why summing independent per-unit terms leaves labels free, while normalising over classes makes them compete for one unit of probability.
Show you can diagnose the mismatch from symptoms: a multi-label task under a softmax trains happily while per-tag recall sits at a ceiling. Also handle the no-true-label case and explain why it needs a background class under a softmax.
Own the framing that loss choice is a claim about the label-generating process. Be able to argue when to reshape a task - collapsing overlapping tags into exclusive classes, or splitting one head into several - rather than patching the loss.
## The one question that decides it Both losses are negative log-likelihood. The only thing that changes is what probability model you are claiming your outputs describe, and that follows from the labels: **can more than one label be true for the same example?** - **No, exactly one is true** (a phoneme, a digit, a language, an intent): the outputs describe a single **categorical** distribution over classes. Use a softmax over the class axis and categorical cross-entropy. - **Yes, any subset can be true** (an audio clip tagged both `speech` and `music`; a document tagged `finance`, `legal` and `europe`): the outputs describe **K independent Bernoulli** decisions. Use one sigmoid per output and binary cross-entropy summed across outputs. ## Categorical cross-entropy With class probabilities `p_1 ... p_K` that sum to one and a target class `c`, the loss for one example is ``` L = -log p_c ``` Written against a one-hot target vector `y` it is `L = -sum_k y_k * log p_k`, but every term with `y_k = 0` drops out, so only the correct class's probability appears. Because the probabilities are tied to a fixed budget of one, this loss is **competitive**: the only way to raise `p_c` is to take mass away from the other classes. That is exactly the inductive bias you want when the classes really do exclude each other. A useful consequence: the loss can only be driven to zero by making the correct class approach probability one, and it is unbounded above - an example the model assigns probability 0.001 to contributes about 6.9, while a confident correct example contributes almost nothing. ## Binary cross-entropy With one probability `p` for a single yes/no output and label `y` in {0, 1}, ``` L = -( y*log p + (1-y)*log(1-p) ) ``` Exactly one of the two terms is active per output. For a multi-label head with K outputs you evaluate this for every output and sum (or average) over K: ``` L = sum_j -( y_j*log p_j + (1-y_j)*log(1-p_j) ) ``` There is no normalisation across `j`. Output 3 rising does not force output 7 down. The loss also supplies a gradient for the **negative** labels, which the categorical form does not have a separate term for - under a softmax, negatives are suppressed only indirectly, through the normaliser. ## Why a multi-label task breaks under a softmax Suppose an example truly has two tags out of twenty. A softmax head can put at most a total of one across both, so the best it can do is roughly 0.5 and 0.5. The target cannot be represented at all if you keep a one-hot convention, and if you use a target of 0.5/0.5 you have quietly changed the task into predicting a mixture rather than a set. At inference you take the top class and the second true tag is never emitted, so recall against multi-tag examples has a hard ceiling. The failure is silent: training loss falls, single-tag examples look fine, and only per-tag recall shows the problem. ## Why a single-label task under independent sigmoids is merely weaker The reverse mistake is less catastrophic. K sigmoids on a mutually exclusive task still trains - you just threw away a constraint the task actually satisfies. The model must learn exclusivity from data instead of getting it for free, the outputs no longer sum to one so you cannot read them as a distribution without renormalising, and there is nothing preventing two outputs from both reading 0.9. ## The binary case, done either way A plain two-outcome problem can legitimately be written either as one sigmoid output or as a two-class softmax, and they are the same model. A softmax over two logits `z_0, z_1` depends only on their difference, since adding a constant to both leaves the probabilities unchanged; the sigmoid parameterises that difference directly with half the output parameters. The two-class softmax carries a redundant degree of freedom - a flat direction in parameter space along which the loss does not change. Nothing goes wrong in practice (weight decay pins it down), but it is worth being able to say out loud, because it is the cleanest way to show you understand that cross-entropy is about the likelihood you are asserting, not about how many output units you happen to have wired up. ## Practical checklist 1. Can two labels co-occur on one example? If yes, per-output sigmoid plus binary cross-entropy. 2. Do you need the outputs to be a probability distribution you can sample from or take an expectation over? If yes, softmax plus categorical cross-entropy. 3. Is there a genuine `none of the above` outcome? A softmax cannot express it without an explicit extra class; independent sigmoids express it naturally as all outputs low.
- Is a single sigmoid output the same model as a two-class softmax?Effectively yes. A softmax over two logits depends only on their difference, because adding the same constant to both leaves the probabilities untouched, so one of the two output degrees of freedom is redundant. A sigmoid parameterises that difference directly with half the parameters. Same function class, same optimum; the two-class version just carries a flat direction that regularisation ends up pinning down.
- What goes wrong if you train a multi-label tagger with categorical cross-entropy?The softmax forces the tag probabilities to sum to one, so the tags compete for a fixed budget. An example with two true tags cannot be fitted - the model splits mass between them, or with a one-hot target it is actively taught that one true tag is wrong. At inference the top-1 read-out emits a single tag, so recall on multi-tag examples is capped no matter how long you train.
- Which loss handles an example with no true labels at all?Binary cross-entropy handles it directly: every output has target zero and every unit gets a push downward. A softmax head cannot represent `none of these` at all, since its probabilities always sum to one - you would have to add an explicit background class so the empty case has somewhere to put its mass.
Categorical cross-entropy is a single ballot where you must pick one candidate; binary cross-entropy is a row of separate yes/no referendum questions, each answered on its own.
saying these in an interview costs you the question
- Picks the loss by counting output units rather than by label exclusivity
- Says categorical cross-entropy handles multi-label targets fine
- Thinks binary cross-entropy only applies to two-class problems
- Applies a softmax across outputs meant to be independent sigmoids
- Cannot say that both losses are negative log-likelihood