skip to content

Activations and Loss Functions

You will learn why ReLU displaced sigmoid, when GELU or SiLU is preferred, and how to match a loss to its task. The favourite probe - 'why not squared error for classification?' - is answered here.

on this pageshow

explore

questions

page 1 of 2

When does a network's output head need binary cross-entropy rather than categorical cross-entropy?

level: juniorimportance: must knowfreq 76%

answer

  1. ask whether labels can co-occur
  2. one budget of probability, or K separate ones
  3. sum over units versus normalise over classes
  4. independent Bernoullis versus one categorical

basics

~20 s

Use binary cross-entropy when each output is an independent yes/no decision, so several can be true at once. Use categorical cross-entropy when exactly one of the classes is correct and the outputs must compete for a single unit of probability.

solid answer

~50 s

The question is whether the labels are mutually exclusive. Categorical cross-entropy is `-log p_c`, where `p` is a softmax over the class axis: the probabilities sum to one, so pushing one class up necessarily pushes the others down. That is right when exactly one label is true - one spoken digit, one phoneme, one language. Binary cross-entropy is `-(y*log p + (1-y)*log(1-p))` applied independently to each output unit, each squashed by its own sigmoid, and the per-unit losses are summed. Nothing couples the units, so an example can carry three true tags and no true tags equally well. So a multi-label tagger gets K sigmoid outputs and summed binary cross-entropy; a single-winner classifier gets one softmax and categorical cross-entropy. Both are the same negative-log-likelihood idea under different likelihood assumptions: K independent Bernoullis versus one categorical.

go deeper

for a junior

Be ready to state the rule in one sentence and name the activation that goes with each loss: sigmoid per output for binary cross-entropy, softmax across classes for categorical cross-entropy.

for a middle

Expect to write both formulas and explain the coupling - why summing independent per-unit terms leaves labels free, while normalising over classes makes them compete for one unit of probability.

for a senior

Show you can diagnose the mismatch from symptoms: a multi-label task under a softmax trains happily while per-tag recall sits at a ceiling. Also handle the no-true-label case and explain why it needs a background class under a softmax.

for a principal

Own the framing that loss choice is a claim about the label-generating process. Be able to argue when to reshape a task - collapsing overlapping tags into exclusive classes, or splitting one head into several - rather than patching the loss.

## The one question that decides it Both losses are negative log-likelihood. The only thing that changes is what probability model you are claiming your outputs describe, and that follows from the labels: **can more than one label be true for the same example?** - **No, exactly one is true** (a phoneme, a digit, a language, an intent): the outputs describe a single **categorical** distribution over classes. Use a softmax over the class axis and categorical cross-entropy. - **Yes, any subset can be true** (an audio clip tagged both `speech` and `music`; a document tagged `finance`, `legal` and `europe`): the outputs describe **K independent Bernoulli** decisions. Use one sigmoid per output and binary cross-entropy summed across outputs. ## Categorical cross-entropy With class probabilities `p_1 ... p_K` that sum to one and a target class `c`, the loss for one example is ``` L = -log p_c ``` Written against a one-hot target vector `y` it is `L = -sum_k y_k * log p_k`, but every term with `y_k = 0` drops out, so only the correct class's probability appears. Because the probabilities are tied to a fixed budget of one, this loss is **competitive**: the only way to raise `p_c` is to take mass away from the other classes. That is exactly the inductive bias you want when the classes really do exclude each other. A useful consequence: the loss can only be driven to zero by making the correct class approach probability one, and it is unbounded above - an example the model assigns probability 0.001 to contributes about 6.9, while a confident correct example contributes almost nothing. ## Binary cross-entropy With one probability `p` for a single yes/no output and label `y` in {0, 1}, ``` L = -( y*log p + (1-y)*log(1-p) ) ``` Exactly one of the two terms is active per output. For a multi-label head with K outputs you evaluate this for every output and sum (or average) over K: ``` L = sum_j -( y_j*log p_j + (1-y_j)*log(1-p_j) ) ``` There is no normalisation across `j`. Output 3 rising does not force output 7 down. The loss also supplies a gradient for the **negative** labels, which the categorical form does not have a separate term for - under a softmax, negatives are suppressed only indirectly, through the normaliser. ## Why a multi-label task breaks under a softmax Suppose an example truly has two tags out of twenty. A softmax head can put at most a total of one across both, so the best it can do is roughly 0.5 and 0.5. The target cannot be represented at all if you keep a one-hot convention, and if you use a target of 0.5/0.5 you have quietly changed the task into predicting a mixture rather than a set. At inference you take the top class and the second true tag is never emitted, so recall against multi-tag examples has a hard ceiling. The failure is silent: training loss falls, single-tag examples look fine, and only per-tag recall shows the problem. ## Why a single-label task under independent sigmoids is merely weaker The reverse mistake is less catastrophic. K sigmoids on a mutually exclusive task still trains - you just threw away a constraint the task actually satisfies. The model must learn exclusivity from data instead of getting it for free, the outputs no longer sum to one so you cannot read them as a distribution without renormalising, and there is nothing preventing two outputs from both reading 0.9. ## The binary case, done either way A plain two-outcome problem can legitimately be written either as one sigmoid output or as a two-class softmax, and they are the same model. A softmax over two logits `z_0, z_1` depends only on their difference, since adding a constant to both leaves the probabilities unchanged; the sigmoid parameterises that difference directly with half the output parameters. The two-class softmax carries a redundant degree of freedom - a flat direction in parameter space along which the loss does not change. Nothing goes wrong in practice (weight decay pins it down), but it is worth being able to say out loud, because it is the cleanest way to show you understand that cross-entropy is about the likelihood you are asserting, not about how many output units you happen to have wired up. ## Practical checklist 1. Can two labels co-occur on one example? If yes, per-output sigmoid plus binary cross-entropy. 2. Do you need the outputs to be a probability distribution you can sample from or take an expectation over? If yes, softmax plus categorical cross-entropy. 3. Is there a genuine `none of the above` outcome? A softmax cannot express it without an explicit extra class; independent sigmoids express it naturally as all outputs low.

  • Is a single sigmoid output the same model as a two-class softmax?
    Effectively yes. A softmax over two logits depends only on their difference, because adding the same constant to both leaves the probabilities untouched, so one of the two output degrees of freedom is redundant. A sigmoid parameterises that difference directly with half the parameters. Same function class, same optimum; the two-class version just carries a flat direction that regularisation ends up pinning down.
  • What goes wrong if you train a multi-label tagger with categorical cross-entropy?
    The softmax forces the tag probabilities to sum to one, so the tags compete for a fixed budget. An example with two true tags cannot be fitted - the model splits mass between them, or with a one-hot target it is actively taught that one true tag is wrong. At inference the top-1 read-out emits a single tag, so recall on multi-tag examples is capped no matter how long you train.
  • Which loss handles an example with no true labels at all?
    Binary cross-entropy handles it directly: every output has target zero and every unit gets a push downward. A softmax head cannot represent `none of these` at all, since its probabilities always sum to one - you would have to add an explicit background class so the empty case has somewhere to put its mass.

Categorical cross-entropy is a single ballot where you must pick one candidate; binary cross-entropy is a row of separate yes/no referendum questions, each answered on its own.

saying these in an interview costs you the question

  • Picks the loss by counting output units rather than by label exclusivity
  • Says categorical cross-entropy handles multi-label targets fine
  • Thinks binary cross-entropy only applies to two-class problems
  • Applies a softmax across outputs meant to be independent sigmoids
  • Cannot say that both losses are negative log-likelihood

context

open as a page

What is a logit in a neural classifier, and why do losses take logits rather than probabilities?

level: juniorimportance: must knowfreq 72%

basics

~20 s

A logit is the raw, unbounded score a network's final layer emits before any squashing function turns it into a probability. Losses take logits so the squashing and the logarithm are computed together, which avoids overflow and log-of-zero.

open as a page

Why is classification accuracy useless as a training loss for a neural network?

level: juniorimportance: must knowfreq 62%

basics

~20 s

Accuracy counts correct predictions, so it is a step function of the model's scores. Nudging a weight usually changes nothing at all, leaving a gradient of zero almost everywhere and no slope for gradient descent to follow.

open as a page

What is ReLU, and why is it the default hidden-layer activation in deep networks?

level: juniorimportance: must knowfreq 84%

basics

~20 s

ReLU computes max(0, z): positive pre-activations pass through unchanged and negative ones become zero. Its derivative is exactly 1 wherever the unit is active, so gradients flow back unshrunk, and evaluating it costs a single comparison.

open as a page

What does per-class loss weighting change when one intent has 200,000 training examples and others have 20?

level: juniorimportance: must knowfreq 70%

basics

~20 s

Per-class weighting multiplies each example's loss by a factor set by its class, so errors on rare intents contribute more gradient. It rebalances the effective class prior the model fits. It adds no new information about rare classes.

open as a page

What happens to a sigmoid hidden unit's gradient when its input is far from zero?

level: juniorimportance: must knowfreq 70%

basics

~20 s

A sigmoid unit's gradient collapses toward zero. The curve flattens at both extremes, so its derivative - at most 0.25, at the origin - becomes almost nothing in the tails, and the layers below stop learning.

open as a page

How do GELU and SiLU treat a negative input differently from ReLU?

level: juniorimportance: must knowfreq 58%

basics

~20 s

ReLU hard-zeroes every negative input. GELU and SiLU multiply the input by a smooth gate between 0 and 1 - a Gaussian CDF for GELU, a sigmoid for SiLU - so small negatives survive and the derivative stays continuous.

open as a page

Why does the gradient of softmax cross-entropy with respect to the logits reduce to p - y?

level: middleimportance: must knowfreq 68%

basics

~20 s

The softmax derivative cancels against the derivative of the log instead of multiplying into it. Differentiating -log of the correct class's probability through the softmax leaves predicted probability minus one-hot target, one subtraction per logit, bounded between -1 and 1.

open as a page

Why does a numerically stable log-sum-exp subtract the row maximum before exponentiating?

level: middleimportance: must knowfreq 58%

basics

~20 s

Subtracting the row maximum caps every exponent at zero, so the exponential cannot overflow. Adding it back outside the logarithm leaves the value provably identical: factoring it out of the sum turns it into an added constant.

open as a page

How do you set the weights when a multi-task loss sums a depth error in metres and a segmentation cross-entropy?

level: middleimportance: must knowfreq 58%

basics

~20 s

Set weights so each term contributes comparable gradient, not so the raw numbers match. A depth error in metres dwarfs a cross-entropy in nats, so it dominates by accident of units. Tune the weights on per-task validation metrics, never on total loss.

open as a page

What is the difference between epistemic and aleatoric uncertainty in a model's predictions?

level: middleimportance: must knowfreq 70%

basics

~20 s

Aleatoric uncertainty is noise inherent in the data — a noisy sensor, identical inputs with different labels — and more data will not remove it. Epistemic uncertainty is the model's ignorance of regions it barely saw, and more data shrinks it.

open as a page

Why can a trained classifier output a 99.9% softmax score on pure noise?

level: middleimportance: must knowfreq 62%

basics

~20 s

Softmax normalizes scores across the trained classes, so its output ranks them and must sum to one — never evidence that the input belongs to any. Nothing in training penalised confident nonsense, and far from the data the gaps can grow.

open as a page

With squared-error loss, why can one outlier target own almost the entire mini-batch gradient?

level: middleimportance: must knowfreq 62%

basics

~20 s

Squared error's per-example gradient grows linearly with the residual, so an example whose error is 50 times larger contributes about 50 times the gradient. Batch averaging divides everyone equally, so it never dilutes that imbalance.

open as a page

In a ReLU network, what makes a hidden unit die, and why does it stay dead?

level: middleimportance: must knowfreq 66%

basics

~20 s

A ReLU unit dies when its pre-activation is negative for every input in the data. It then outputs zero, its local derivative is zero, and so the loss sends exactly zero gradient to its weights and bias — nothing ever moves them again.

open as a page

In focal loss with gamma = 2, what happens to an example already given probability 0.99 for its true class?

level: middleimportance: must knowfreq 60%

basics

~10 s

Focal loss multiplies cross-entropy by (1 - p_t)^gamma. At gamma = 2 and p_t = 0.99 that factor is 0.0001, so the example contributes one ten-thousandth of its cross-entropy and stops steering training.

open as a page

Why is tanh usually preferred over sigmoid for a network's hidden layers?

level: middleimportance: must knowfreq 62%

basics

~20 s

Tanh is zero-centred and steeper. Its outputs span (-1, 1) rather than (0, 1), so the next layer's weight gradients are not all forced to share a sign, and its peak derivative is 1.0 against sigmoid's 0.25.

open as a page

Why does squared error on a softmax classification head train so slowly?

level: middleimportance: should knowfreq 52%

basics

~10 s

Squared error keeps the softmax Jacobian's probability factor instead of cancelling it, so an example whose correct class sits at probability 0.01 sends almost no signal. The composed objective is also non-convex, unlike cross-entropy.

open as a page

What does hinge loss's margin do that cross-entropy on the same scores does not?

level: middleimportance: should knowfreq 41%

basics

~20 s

Hinge loss demands a cushion, not just a correct sign: with max(0, 1 - y*s), an example scored 0.9 toward the right class still pays 0.1, and only past 1.0 does its gradient hit zero. Cross-entropy never stops pushing.

open as a page

Why attach an auxiliary loss head to an intermediate layer, and why anneal its coefficient toward zero?

level: middleimportance: should knowfreq 42%

basics

~20 s

An auxiliary head attached partway up gives the early layers a short gradient path and their own training signal, which eases optimisation of a deep stack. Its coefficient is annealed toward zero so the real head governs the final model.

open as a page

Squared-error loss is maximum likelihood under what assumption about the target?

level: middleimportance: should knowfreq 47%

basics

~20 s

That the target is Gaussian around the network's output with a constant, input-independent variance. Minimising squared error is exactly maximising that Gaussian likelihood, which is why the trained output estimates the conditional mean of the target.

open as a page

Why is SiLU's dip below zero acceptable when most activations are monotone?

level: middleimportance: should knowfreq 41%

basics

~20 s

Monotonicity is not required for training or for function approximation. SiLU dips to about -0.278 near x = -1.278 and then rises; its derivative is continuous and its output is bounded below, so gradient descent handles it fine.

open as a page

Why is chaining a softmax into a separate log less stable than one fused log-softmax?

level: seniorimportance: should knowfreq 44%

basics

~20 s

The chain materializes a probability first, and that intermediate destroys information: a tail value rounds to exactly zero, so its logarithm is negative infinity. Fusing computes the log-probability as a subtraction that never forms the probability at all.

open as a page

For a rare-positive classifier missing its F1 target, do you train soft-F1 or tune the threshold?

level: seniorimportance: should knowfreq 44%

basics

~20 s

Tune the threshold first: train with cross-entropy, then pick the operating point on a held-out split. A differentiable soft-F1 loss, built from summed probabilities instead of hard counts, is the fallback when no fixed threshold works.

open as a page

How would you get an epistemic uncertainty estimate out of an already-trained network?

level: seniorimportance: should knowfreq 48%

basics

~20 s

Sample several plausible models and measure their disagreement on the input. Keeping dropout active for thirty stochastic forward passes is the cheap option; five independently seeded networks are the stronger one. The spread, not the average, is the signal.

open as a page

When would you replace ReLU with Leaky ReLU, PReLU, or ELU, and what does each cost?

level: seniorimportance: should knowfreq 46%

basics

~20 s

Leaky ReLU gives the negative branch a small fixed slope so units cannot go permanently silent. PReLU learns that slope, usually one per channel, at the cost of extra parameters. ELU saturates toward a negative bound, centring activations but paying for an exponential.

open as a page

Reweighting the loss versus oversampling the rare class: what is identical and what actually differs?

level: seniorimportance: should knowfreq 42%

basics

~10 s

Both target the same rebalanced objective and give the same expected gradient. They differ in gradient variance, mini-batch composition, what one epoch means, compute per pass, and how badly the rare rows get memorised.

open as a page

Your sigmoid hidden layer outputs 0.999 for every training example - what caused it?

level: seniorimportance: should knowfreq 40%

basics

~20 s

The pre-activations are far too large and positive, because the input features were fed in unscaled. Decibel-scale loudness values in the tens, summed over many channels, push every unit deep into the sigmoid's flat tail.

open as a page

Would you swap GELU for ReLU in a battery-powered on-device vision backbone?

level: seniorimportance: should knowfreq 34%

basics

~20 s

Only if measurement says the activation is genuinely costing you. GELU evaluates a transcendental per element where ReLU does a comparison, and on a small backbone with large feature maps that can be a real share of frame time.

open as a page

Two tasks' gradients meet a shared trunk at negative cosine similarity — how do you decide what to sacrifice?

level: principalimportance: should knowfreq 34%

basics

~20 s

Negative cosine between two task gradients means an update helping one hurts the other, and no choice of weights removes that. Name the task the product is judged on, fix the degradation you will accept elsewhere, then pick a mechanism to enforce it.

open as a page

Your demand targets carry rare 50x sensor spikes — how do you choose between Huber, a target transform, and a data fix?

level: principalimportance: should knowfreq 38%

basics

~20 s

Decide what the spikes are before touching the loss. Corruption belongs in the data pipeline, where the fix is auditable. A robust loss like Huber stabilises training but shifts the statistic you estimate, so it under-predicts spikes that turn out to be real.

open as a page

showing 1–30 of 36