skip to content

Why can inverse-frequency class weights destabilise training, and what does effective-number weighting fix?

level: middleimportance: nice to knowfreq 30%

answer

  1. the weight ratio equals the count ratio
  2. rare batches, enormous single-row gradients
  3. overlapping samples cover less new ground
  4. the effective count saturates at a ceiling
  5. beta dials between uniform and inverse frequency

basics

~10 s

Inverse-frequency weights scale with the raw count ratio, so a 200,000-versus-20 split gives a 10,000-to-1 weight and wildly noisy updates. Effective-number weighting uses (1 - beta^n)/(1 - beta), which saturates and caps the ratio.

solid answer

~50 s

Inverse-frequency weighting sets `w_c` proportional to `1 / n_c`, so with a 200,000-example head class and a 20-example tail class the weight ratio is 10,000 to 1. Any mini-batch that happens to contain a tail row produces a gradient thousands of times larger than a typical step, which shows up as loss spikes, learning-rate sensitivity, overfitting of those 20 rows, and huge leverage for any mislabelled one among them. Effective-number weighting replaces the count `n` with `E_n = (1 - beta^n)/(1 - beta)` and weights by `1 / E_n`. The idea is that samples overlap, so each additional example adds less new coverage; `E_n` grows sub-linearly and saturates at `1/(1 - beta)`. `beta = 0` gives `E_n = 1` for every class, which is no reweighting at all; `beta` approaching 1 gives `E_n` approaching `n`, recovering inverse frequency. So `beta` is a single dial between the two extremes with a bounded weight ratio in between.

code

python · 10 lines
python
def effective_num(n, beta):
    return (1.0 - beta ** n) / (1.0 - beta)

counts = {"head": 200000, "tail": 20}

for beta in (0.0, 0.9, 0.99, 0.999):
    w = {c: 1.0 / effective_num(n, beta) for c, n in counts.items()}
    print(beta, round(w["tail"] / w["head"], 1))

print("inverse frequency", counts["head"] / counts["tail"])

go deeper

for a junior

Know that the simple weight is one over the class count, and that when counts differ by four orders of magnitude that ratio is four orders of magnitude too. Recognising the danger is enough at this level.

for a middle

Be able to write (1 - beta^n)/(1 - beta), say that it saturates at 1/(1 - beta), and state which weighting each endpoint of beta recovers. Explain why a saturating count bounds the weight ratio.

for a senior

Show that you would reach for the square-root or clipped weight first and justify the extra hyperparameter only when counts span orders of magnitude. Tie the choice to held-out macro recall and to observed loss-spiking, not to taste.

for a principal

Frame it as a policy: any static weighting scheme is a claim about how much a tail class is worth relative to the head, and that claim belongs in the product's evaluation contract rather than in a hyperparameter sweep nobody revisits.

## Why the obvious weight is too aggressive The reflex choice for class weights is `w_c` proportional to `1 / n_c`: give every class the same total loss mass. On mild imbalance that is fine. On a long-tailed corpus it is violent. With a head intent at 200,000 examples and tail intents at 20, the weight ratio is exactly the count ratio, 10,000 to 1. Think about what that does to a mini-batch. Batches are drawn from the natural distribution, so most contain no tail row at all, and their gradients look like the head class's gradient. Occasionally a batch contains one tail row, and that single row contributes a term thousands of times larger than everything else combined. The sequence of updates is therefore mostly ordinary steps punctuated by rare enormous ones. Concretely you see: a loss curve with spikes, a run that diverges at a learning rate that was previously safe, gradient clipping firing constantly, and normalisation-layer statistics jerked around by outsized activations. The statistical problems are worse than the numerical ones. Extreme weights mean the network fits those 20 rows almost exactly, so training macro-recall improves while held-out macro-recall does not. And if one of the 20 is mislabelled, you have just made a single annotation error the single most influential row in the entire dataset. ## The effective-number idea Effective-number weighting starts from a different question: how much *distinct* information does `n` examples of a class actually carry? Real examples overlap - two utterances of the same tail intent are often near-paraphrases, and standard augmentation makes the overlap worse. Under a simple model where each new sample covers some fixed fraction of already-covered space, the expected covered volume after `n` samples is `E_n = (1 - beta^n) / (1 - beta)` with `beta` in `[0, 1)`. The class weight is then proportional to `1 / E_n` instead of `1 / n`. Read off the two ends. At `beta = 0`, `E_n = 1` for every `n`, so all classes get the same weight - the unweighted loss. As `beta` approaches 1, `E_n` approaches `n` (the sum `1 + beta + ... + beta^(n-1)` tends to `n`), so you recover inverse-frequency weighting. In between, `E_n` grows sub-linearly and saturates at `1/(1 - beta)`, which is what bounds the damage: no matter how large the head class is, its effective number cannot exceed that ceiling. The numbers make it concrete. Take `beta = 0.999`, so the ceiling is 1000. The head class with 200,000 rows sits at the ceiling, `E = 1000`. The tail class with 20 rows gets `E = (1 - 0.999^20)/0.001`, about 19.8 - barely below its raw count, because 20 samples overlap very little. The weight ratio is therefore about 1000/19.8, roughly 50 to 1, instead of 10,000 to 1. The tail is still strongly favoured; the head is no longer effectively erased. ## Choosing beta, and the cheaper alternatives `beta` is a hyperparameter, and the useful mental model is that `1/(1 - beta)` is the count at which you stop believing extra examples add much. Values are usually quoted in the 0.9 to 0.9999 range and tuned on held-out macro-averaged recall; note that the ceiling is very sensitive near 1, so treat it as a logarithmic dial rather than a linear one. It is also worth knowing the plainer options, because interviewers ask what you would try first. Weighting by `1 / sqrt(n_c)` is a one-line change that compresses a 10,000-to-1 ratio to 100-to-1 and often captures most of the benefit. Simply clipping the weight ratio at some maximum is even blunter and works. Effective-number weighting earns its keep when you want a principled, single-parameter family that spans from no reweighting to full inverse frequency, and when the class counts span several orders of magnitude so that ad-hoc clipping would have to be re-tuned every time the corpus grows. ## What it does not change Effective-number weighting is still a *static* weight computed from counts before training begins. It knows nothing about which examples are hard, it does not adapt as the model improves, and it shifts the model's implied class prior exactly as any other class weight does, so probability outputs still need recalibrating if a downstream consumer treats them as rates. It also does not conjure variety: a class with 20 examples has 20 examples whatever weight you give it, and if the real problem is that the tail is unlearnable from 20 rows, the answer is more data or a different decomposition of the label space, not a better weighting formula.

  • What do beta = 0 and beta approaching 1 each recover?
    `beta = 0` makes `E_n = 1` for every class, so all weights are equal and you are back to the plain unweighted loss. As `beta` approaches 1, `E_n` approaches `n` and the weight approaches `1 / n_c`, which is ordinary inverse-frequency weighting. Everything useful lives strictly between those two endpoints.
  • What is the cheap alternative if you do not want another hyperparameter?
    Weight by `1 / sqrt(n_c)`, or clip the weight ratio at a fixed maximum. The square root turns a 10,000-to-1 ratio into 100-to-1 with no tuning and usually captures most of the gain. Effective-number weighting is worth the extra dial mainly when counts span several orders of magnitude and keep changing.
  • Why is a 10,000-to-1 weight dangerous even if the average gradient is correct?
    Because it is the variance, not the mean, that breaks training. Most batches contain no tail row; the rare ones that do produce a step thousands of times larger than normal. The expected update may be right while the actual trajectory spikes, clips, and destabilises the normalisation statistics.

saying these in an interview costs you the question

  • Assumes inverse frequency is the only principled weighting
  • Cannot say what beta = 0 recovers
  • Thinks effective number counts duplicate rows literally
  • Ignores gradient variance and looks only at the mean
  • Believes a better weight formula compensates for 20 examples

context