skip to content

Reweighting the loss versus oversampling the rare class: what is identical and what actually differs?

level: seniorimportance: should knowfreq 42%

answer

  1. same objective, different estimator
  2. expected gradient identical, variance is not
  3. who is in the batch changes
  4. batch statistics see a different mix
  5. one pass no longer means one dataset

basics

~10 s

Both target the same rebalanced objective and give the same expected gradient. They differ in gradient variance, mini-batch composition, what one epoch means, compute per pass, and how badly the rare rows get memorised.

solid answer

~50 s

Sampling a rare class more often and multiplying its loss by a matching weight optimise the same rebalanced objective, so in expectation the gradient is the same. Everything that differs is a property of the sampling process, and on a mini-batch trainer those properties matter. Weighting keeps every batch at the natural class mix, so the rare class is absent from most batches and then arrives with a huge multiplier - low bias, high gradient variance, spiky updates. Oversampling puts rare rows in nearly every batch, which smooths the updates but means the model sees the same handful of rows over and over, so it memorises them unless augmentation supplies real variety. Batch composition also changes, so any layer that normalises over the batch sees different statistics. And an epoch stops being comparable: with oversampling a pass is longer and revisits rare rows many times, so epoch-keyed schedules and early stopping shift under you.

go deeper

for a junior

Know that both approaches aim at the same rebalanced objective, and that oversampling changes who is in each batch while loss weighting changes only how much each row counts.

for a middle

Explain that the expected gradient is the same and the variance is not, and give the concrete consequences: spiky updates under heavy weights, memorised duplicates under oversampling.

for a senior

Show the operational detail an interviewer is fishing for - epoch semantics breaking, batch-normalised statistics shifting, augmentation being what makes oversampling pay off - and describe how you would combine the two at partial strength.

for a principal

Own the reporting standard. Decide up front that runs are compared on gradient steps and on held-out macro recall at the natural rate, so that a rebalancing experiment cannot quietly become a training-budget experiment.

## The part that is genuinely the same Suppose you want to optimise a rebalanced objective in which each class contributes loss mass proportional to some target rather than to its raw count. There are two ways to get there. Draw examples from the natural distribution and multiply each one's loss by a weight; or change the sampler so that examples are drawn at the target rate and leave the loss alone. Written as expectations, these are the same integral - multiplying a term by `w` and multiplying its sampling probability by `w` both scale its contribution by `w`. So the expected gradient is identical, and any claim that one of them "optimises something different" is wrong at the level of the objective. The interesting differences all live in the estimator, not the target. A deep network is trained with noisy mini-batch estimates of that gradient, and the two routes produce estimates with very different behaviour. ## Gradient variance and batch composition With loss weighting, the sampler is untouched, so a class that makes up one row in ten thousand appears in roughly one batch in thirty at batch size 256 - and when it appears, its term is multiplied by a large weight. The estimate is unbiased and noisy: most steps carry no signal about that class at all, and the occasional step carries a very loud one. Symptoms are loss spikes, clipping firing, and sensitivity to the learning rate. With oversampling, rare rows appear in nearly every batch at weight 1, so each step carries a modest, consistent signal. Variance drops. The price is that there are still only 20 distinct rare rows; oversampling shows the same rows repeatedly, and repeated exposure at full weight is a fast route to memorising them. Augmentation is what converts repetition into something useful, which is why oversampling tends to be much more effective for images and audio, where strong augmentation exists, than for short text where it does not. There is a specifically deep-learning consequence too. Any layer that normalises across the batch dimension - batch normalisation computes a per-channel mean and variance over the examples in the batch (and over spatial positions for a convolutional feature map) - sees a different input distribution when you change the batch's class mix. Resampling therefore changes those statistics and the running estimates used at inference; loss weighting does not. If you are debugging why a resampled run behaves differently at inference than during training, this is the first place to look. ## What an epoch means This one trips people up in real projects. With loss weighting, one epoch is one pass over the real dataset, and comparing runs is easy. With oversampling by duplication, one pass is longer - potentially much longer - and it revisits the rare rows many times. Anything keyed to epochs silently changes meaning: the learning-rate schedule, the early-stopping patience, the checkpoint cadence, the number of times a given rare row has been fitted. Two runs reported as "20 epochs" may differ by an order of magnitude in gradient steps and by far more in exposure to the rare class. Compare on gradient steps or on examples consumed, not on epochs. Undersampling the majority instead is the mirror image: passes become cheap and fast, variance is low and the batch mix is controlled, but you throw away most of your data, and on a head class with genuine internal diversity that is a real loss of information. Reweighting always keeps every row. ## How to choose in practice A reasonable default is: start with a moderate loss weight, because it is a one-line change that leaves the data pipeline, the epoch semantics and the batch statistics alone. Move to a rebalanced sampler when the weights you need are extreme enough to make training unstable, since the sampler converts a variance problem into a repetition problem, and repetition is easier to manage with augmentation than spikes are to manage with clipping. Many practical setups do both at partial strength: sample toward the target rate to keep the rare class present in every batch, and apply a residual weight for the remaining gap, which keeps both the variance and the duplication moderate. Whichever you pick, two things stay true. Both shift the model's implied class prior, so probability outputs need recalibrating on data drawn at the natural rate before any downstream system reads them as rates. And neither invents new rare-class examples - they only change how often and how loudly the ones you have are heard.

  • Why compare the two runs on gradient steps rather than epochs?
    Because oversampling by duplication lengthens a pass and revisits the same rare rows within it. Two runs labelled twenty epochs can differ by an order of magnitude in updates and far more in exposure to the rare class, so anything keyed to epochs - schedule, patience, checkpointing - has silently changed meaning.
  • Which of the two interacts with batch-normalised layers, and how?
    Resampling does. Batch normalisation computes a per-channel mean and variance over the examples in the mini-batch, so changing the class mix changes those statistics and the running estimates carried into inference. Loss weighting leaves the batch composition untouched and only rescales loss terms.
  • When would you undersample the majority instead of either option?
    When the head class is genuinely redundant and compute is the binding constraint - passes get much cheaper, variance is low and batch mix is controlled. Avoid it when the head has real internal diversity, since discarded rows are information you paid to collect and cannot get back inside the run.
  • Can you use both at once?
    Yes, and it is often the best answer. Sample partway toward the target rate so the rare class is present in most batches, then apply a residual loss weight for the remaining gap. That keeps gradient variance moderate without duplicating the same rows enough times to memorise them.

saying these in an interview costs you the question

  • Claims the two optimise different objectives
  • Says oversampling adds information the rare class lacked
  • Compares a resampled run to a weighted run by epoch count
  • Ignores that duplicated rows get memorised without augmentation
  • Overlooks batch-normalised statistics shifting with the batch mix

context