skip to content

Activations and Loss Functions

You will learn why ReLU displaced sigmoid, when GELU or SiLU is preferred, and how to match a loss to its task. The favourite probe - 'why not squared error for classification?' - is answered here.

on this pageshow

explore

questions

page 2 of 2

Why can inverse-frequency class weights destabilise training, and what does effective-number weighting fix?

level: middleimportance: nice to knowfreq 30%

basics

~10 s

Inverse-frequency weights scale with the raw count ratio, so a 200,000-versus-20 split gives a 10,000-to-1 weight and wildly noisy updates. Effective-number weighting uses (1 - beta^n)/(1 - beta), which saturates and caps the ratio.

open as a page

Your character language model reports 1.35 bits per character - what does that number mean?

level: seniorimportance: nice to knowfreq 38%

basics

~20 s

Bits per character is the training objective itself in base 2: the mean negative log-likelihood the model assigns to each true next character on held-out text. Exponentiating gives a perplexity of about 2.55 characters of effective choice.

open as a page

How do you train a search ranker when the target metric NDCG has zero gradient?

level: seniorimportance: nice to knowfreq 27%

basics

~20 s

Change the unit of the loss from the item to the pair. NDCG reads only the sort order, so it is flat; a logistic loss on the score difference inside a should-outrank pair is smooth and pushes the order right.

open as a page

In uncertainty weighting of a multi-task loss, how are the per-task weights learned and what stops them collapsing to zero?

level: seniorimportance: nice to knowfreq 30%

basics

~20 s

Each task carries a learned scalar noise parameter, and its loss weight is the inverse of that scalar, so hard-to-fit tasks are downweighted automatically. An added log-noise penalty per task is what stops every weight sliding to zero.

open as a page

How does a regression head that outputs a mean and a log-variance train under Gaussian negative log-likelihood?

level: seniorimportance: nice to knowfreq 26%

basics

~20 s

A head emits two numbers per input, a mean and s = log variance, trained with 0.5 * (s + (y - mean)^2 * exp(-s)). The exponential weights each residual by one over the predicted variance; the s term punishes inflating it.

open as a page

How do you set the abstention threshold for a chest-radiograph triage model that routes uncertain cases to a radiologist?

level: principalimportance: nice to knowfreq 34%

basics

~20 s

Treat it as a coverage-versus-risk trade, not a modelling choice. Sweep the threshold on held-out data, then pick the point where the deferred volume fits real radiologist capacity and the residual error is acceptable — and check that trade per subgroup.

open as a page

showing 31–36 of 36