skip to content

In focal loss with gamma = 2, what happens to an example already given probability 0.99 for its true class?

level: middleimportance: must knowfreq 60%

answer

  1. cross-entropy times a modulating factor
  2. the factor uses the current predicted probability
  3. gamma = 0 gives you back cross-entropy
  4. 0.01 squared is one ten-thousandth
  5. hard examples keep almost all their loss

basics

~10 s

Focal loss multiplies cross-entropy by (1 - p_t)^gamma. At gamma = 2 and p_t = 0.99 that factor is 0.0001, so the example contributes one ten-thousandth of its cross-entropy and stops steering training.

solid answer

~50 s

Focal loss replaces `-log(p_t)` with `-(1 - p_t)^gamma * log(p_t)`, where `p_t` is the probability the model currently assigns to that example's true class and gamma is the focusing exponent. At gamma = 2 an example already at `p_t = 0.99` picks up a factor of `(1 - 0.99)^2 = 0.0001`, so it contributes one ten-thousandth of its plain cross-entropy; an example at `p_t = 0.5` keeps a quarter of its. The mass of the loss therefore shifts to examples the model still gets wrong. Two things distinguish it from a class weight. First, gamma = 0 recovers plain cross-entropy, so it is a strict generalisation. Second, the factor depends on the current prediction, so it is recomputed every step and an example that becomes easy quietly demotes itself. Focal loss is usually paired with a separate per-class factor alpha: alpha addresses class frequency, gamma addresses easy-versus-hard.

go deeper

for a junior

Recall the shape: cross-entropy multiplied by one minus the predicted probability of the true class, raised to gamma. Know that it shifts effort toward examples the model still gets wrong.

for a middle

Compute the factor out loud at gamma = 2 for a couple of probabilities, say what gamma = 0 recovers, and separate the per-class alpha term from the per-example gamma term as fixes for two different imbalances.

for a senior

Demonstrate that you know when not to use it: noisy or weakly-labelled corpora, where up-weighting the hard tail up-weights the annotation errors. Explain how you would sweep gamma and what held-out signal would tell you to stop.

for a principal

Own the call between changing the objective and changing the data. Focal loss is cheap and reversible; label cleanup and targeted collection cost money but move the ceiling. Be able to justify which one your team spends the quarter on.

## The mechanism Cross-entropy for one example is `-log(p_t)`, where `p_t` is the probability the model assigns to that example's true class. Focal loss multiplies it by a modulating factor: `FL = -(1 - p_t)^gamma * log(p_t)` The exponent gamma (the focusing parameter) controls how sharply well-classified examples are suppressed. Plugging in numbers at gamma = 2: - `p_t = 0.99`: factor `0.01^2 = 0.0001`. One ten-thousandth of its cross-entropy. - `p_t = 0.9`: factor `0.1^2 = 0.01`. One hundredth. - `p_t = 0.5`: factor `0.5^2 = 0.25`. A quarter. - `p_t = 0.1`: factor `0.9^2 = 0.81`. Nearly untouched. So the loss - and the gradient with it - is redistributed from the examples the model has already solved onto the ones it has not. At gamma = 0 the factor is 1 everywhere and you are back to plain cross-entropy exactly; gamma is a dial, not a switch, and values around 1 to 5 are the usual range with 2 the common starting point. ## Why this is a different lever from class weighting A per-class weight is computed from counts before training and is constant. It says "this label matters more". The focal factor is computed from the model's current output and changes every single step. It says "this *example* is not yet learned". Those are orthogonal, which is why the published form of focal loss carries both: an alpha factor that depends on the class, multiplied by the `(1 - p_t)^gamma` factor that depends on the prediction. The distinction matters in a case interviewers like. Suppose a class is rare but trivially separable - a template-generated intent, say. A class weight will keep hammering on it long after the model has nailed it, because the weight cannot see that it is solved. The focal factor demotes it automatically as `p_t` rises. Conversely, if a class is common but genuinely hard, a class weight will downweight it while the focal factor keeps it in play. In practice, when the imbalance is in the label distribution you want alpha; when the imbalance is between easy and hard examples inside a class, you want gamma. ## Relation to hard-example mining A blunter way to spend the gradient on hard examples is to mine them: compute the per-example loss for the batch, keep the top-k largest, and backpropagate only through those. Focal loss is the smooth version of the same instinct. Mining makes a hard cut - an example is in or out - and throws away the rest of the batch entirely; focal loss keeps every example but scales it continuously, which gives a smoother objective, no `k` to tune, and no discontinuity as an example crosses the cut. Both are per-example, prediction-dependent reweightings, and both suffer from the same failure mode. ## The failure mode: label noise Anything that upweights whatever the model finds hard will upweight whatever is impossible, and a mislabelled row is permanently impossible. Its `p_t` stays near zero, its focal factor stays near 1 while everything the model has learned is suppressed toward zero, and so its *relative* influence on the gradient grows without bound as training proceeds. The same holds for genuinely ambiguous examples - two intents that a human annotator also cannot separate - and for outliers. On a clean, curated dataset focal loss with gamma = 2 is often a free win. On a scraped or weakly-labelled corpus, raising gamma can make things measurably worse, and the correct response is to audit the highest-loss examples rather than to raise gamma further. There is a second, milder caution. Very large gamma starves training: if almost every example is suppressed, the effective batch size shrinks to the handful of hard rows, gradient variance rises, and progress stalls. Sweeping gamma over a couple of values and watching held-out macro recall is the standard defence. ## Its effect on the output scores Plain cross-entropy is a proper scoring rule: the objective is minimised exactly when the model outputs the true conditional probability. Focal loss with gamma greater than 0 is not - its minimiser is a deliberately distorted version of that probability, which is the whole point of the modulating factor. So the raw outputs of a focal-trained model are best treated as scores rather than probabilities. If a downstream consumer needs a probability, fit a calibration map on held-out data drawn at the natural rate rather than reading the head's output directly.

  • What does gamma = 0 recover, and why does that matter?
    At gamma = 0 the modulating factor `(1 - p_t)^0` is 1 for every example, so focal loss is exactly cross-entropy (times the alpha factor if one is used). It matters because it makes focal loss a strict generalisation with a free baseline: you can sweep gamma from 0 upward and know that 0 reproduces your current run.
  • Which dataset problem does focal loss make worse?
    Label noise. A mislabelled example keeps `p_t` near zero forever, so its modulating factor stays near 1 while everything correctly learned is crushed toward zero - its relative share of the gradient grows as training proceeds. On weakly-labelled data, raising gamma amplifies annotation errors, and the fix is to inspect the highest-loss rows.
  • How does focal loss relate to keeping only the top-k hardest losses in a batch?
    They are the smooth and the hard version of one idea. Top-k mining makes a binary in-or-out cut and discards the rest of the batch; focal loss keeps every example and scales it continuously by `(1 - p_t)^gamma`. Focal loss removes the `k` hyperparameter and the discontinuity, but shares mining's sensitivity to mislabelled rows.
  • Do you still need a class weight if you use focal loss?
    Often yes. Gamma addresses easy-versus-hard, not head-versus-tail; a rare class whose examples are hard gets help from gamma, but a rare class that is merely outvoted does not. The published form carries a per-class alpha alongside gamma precisely because the two imbalances are different problems.

A class weight is a fixed volume knob per label, set before the concert starts. The focal factor is an automatic mixer that turns down whichever instrument is already in tune, every bar.

saying these in an interview costs you the question

  • Says focal loss down-weights the majority class
  • Cannot state what gamma = 0 recovers
  • Thinks the modulating factor is fixed before training
  • Ignores that mislabelled rows gain influence as gamma rises
  • Treats focal-trained outputs as calibrated probabilities

context