Activations and Loss Functions
You will learn why ReLU displaced sigmoid, when GELU or SiLU is preferred, and how to match a loss to its task. The favourite probe - 'why not squared error for classification?' - is answered here.
on this pageshowhide
explore
- Hidden Unit Nonlinearities9 questions
- Matching Objective to Task11 questions
- Cross-Entropy and NLL4 questions
- Logits and Log-Sum-Exp3 questions
- Squared Error and Huber4 questions
- Imbalance, Surrogates and Trade-offs16 questions
- Weighting and Focal Loss4 questions
- Surrogates for Hard Metrics4 questions
- Multi-Task Loss Weighting4 questions
- Predictive Uncertainty and Abstention4 questions
questions
page 2 of 2Why can inverse-frequency class weights destabilise training, and what does effective-number weighting fix?
basics
~10 sInverse-frequency weights scale with the raw count ratio, so a 200,000-versus-20 split gives a 10,000-to-1 weight and wildly noisy updates. Effective-number weighting uses (1 - beta^n)/(1 - beta), which saturates and caps the ratio.
Your character language model reports 1.35 bits per character - what does that number mean?
basics
~20 sBits per character is the training objective itself in base 2: the mean negative log-likelihood the model assigns to each true next character on held-out text. Exponentiating gives a perplexity of about 2.55 characters of effective choice.
How do you train a search ranker when the target metric NDCG has zero gradient?
basics
~20 sChange the unit of the loss from the item to the pair. NDCG reads only the sort order, so it is flat; a logistic loss on the score difference inside a should-outrank pair is smooth and pushes the order right.
In uncertainty weighting of a multi-task loss, how are the per-task weights learned and what stops them collapsing to zero?
basics
~20 sEach task carries a learned scalar noise parameter, and its loss weight is the inverse of that scalar, so hard-to-fit tasks are downweighted automatically. An added log-noise penalty per task is what stops every weight sliding to zero.
How does a regression head that outputs a mean and a log-variance train under Gaussian negative log-likelihood?
basics
~20 sA head emits two numbers per input, a mean and s = log variance, trained with 0.5 * (s + (y - mean)^2 * exp(-s)). The exponential weights each residual by one over the predicted variance; the s term punishes inflating it.
How do you set the abstention threshold for a chest-radiograph triage model that routes uncertain cases to a radiologist?
basics
~20 sTreat it as a coverage-versus-risk trade, not a modelling choice. Sweep the threshold on held-out data, then pick the point where the deferred volume fits real radiologist capacity and the residual error is acceptable — and check that trade per subgroup.
showing 31–36 of 36