skip to content

What is the difference between random oversampling and random undersampling of an imbalanced training set?

level: juniorimportance: must knowfreq 74%

answer

  1. one copies rows, one deletes rows
  2. neither adds new information
  3. duplicates invite memorisation
  4. discarded majority rows are gone for that fit
  5. target ratio is a tunable knob

basics

~10 s

Random oversampling duplicates existing minority rows until the classes are closer to balanced. Random undersampling deletes majority rows to reach the same ratio. Oversampling keeps every row but repeats information; undersampling throws information away.

solid answer

~40 s

Both change the class ratio the learner sees, and neither adds information about the rare class. Random oversampling samples minority rows with replacement and appends the copies, so a telecom table with 10,000 churners and 200,000 non-churners grows instead of shrinking; the risk is that a flexible model memorises the duplicated rows, since a duplicate sits exactly on top of the original and widens nothing. Random undersampling drops majority rows at random, so balancing that same table means discarding about 95% of the non-churners; training gets fast and cheap, but genuine majority variation disappears and the boundary moves between runs. I treat undersampling as reasonable when the majority is huge and redundant, oversampling when data is scarce, and I tune the target ratio rather than assuming 1:1.

go deeper

for a junior

Be ready to say in one sentence each what oversampling and undersampling do to the row count, and to state that neither creates new information about the rare class.

for a middle

Explain the mechanics: duplicates multiply a row's weight in the loss or in node class counts, while deletion removes genuine majority variation and raises run-to-run variance.

for a senior

Show that you check whether imbalance is actually hurting before intervening, tune the resampled ratio instead of defaulting to 1:1, and recover discarded majority rows through ensemble undersampling.

for a principal

Own the tradeoff between training cost and information loss across a portfolio of models, and be clear that resampling shifts the operating point rather than adding evidence about the rare class.

## The problem being attacked An imbalanced training set is one where the label you care about is rare: 10,000 churners against 200,000 non-churners in a telecom table, or 400 defective parts in 50,000 inspected units. Most learners minimise an average loss over rows, so a class that supplies 5% of the rows supplies about 5% of the pressure. The model can score well by leaning toward the common label. **Resampling** attacks this by changing the composition of the training rows so the class ratio the learner optimises against is closer to balanced. The two crudest forms are random oversampling and random undersampling. ## Random oversampling: duplicate the minority Draw minority rows uniformly at random **with replacement** and append the copies until the target ratio is reached. Nothing is invented — repeating the same 400 defective parts twenty times still describes exactly 400 distinct defects, and gets you to 8,000 minority rows against 49,600 good ones, which is not even balanced. Mechanically, a duplicated row contributes to the objective once per copy. In a tree, node class counts change, so different splits are chosen and a deep, pure leaf drawn tightly around a duplicated point becomes easy to justify — including around a mislabelled one. In a margin- or distance-based model, the copies occupy the *same* coordinates as the original, so they cannot enlarge the region the minority covers; they only raise the price of getting that one location wrong. Arithmetically, appending a row twenty times is the same as counting it twenty times in the loss. Consequences worth naming: - **Memorisation.** With enough capacity the model fits the individual duplicated rows, so training scores look excellent and held-out scores do not move. - **Cost.** The training set grows. Balancing the telecom table by oversampling produces 400,000 rows instead of 210,000. - **No new coverage.** Minority regions that were never sampled stay unsampled. ## Random undersampling: delete the majority Drop majority rows at random until the ratio is reached. Balancing the telecom table means keeping 10,000 of the 200,000 non-churners and discarding roughly 95% of them. - **It is fast.** The training set falls to 20,000 rows, which matters when you are fitting many candidate models. - **The loss is real and one-directional.** Unusual-but-genuine majority rows — the rare corners of "normal" — may vanish entirely, so the model's picture of the common class gets thinner and the boundary shifts noticeably between random draws. That is added variance, not just lost accuracy. - **The standard mitigation is not to discard permanently.** Train several models on different random majority subsets and combine them, so every majority row is seen by some ensemble member. Ensemble undersampling methods such as EasyEnsemble are built on exactly this idea. ## What neither one does Neither adds information about the rare class. What they change is the ratio the objective is averaged over, which makes the model more willing to emit the rare label. A direct consequence: the scores a resampled model produces reflect the resampled ratio, not the population rate — so a raw probability from such a model should not be read as a real-world frequency. ## Choosing how far to go Full balance is a convention, not a requirement, and on extreme ratios it is a violent intervention: a 0.8% defect rate needs about 124 copies of every defect to reach 1:1. Partial resampling — lifting the minority to 10% or 20% — is usually the better starting point, and the target ratio deserves to be tuned like any other hyperparameter and judged on evaluation data you never resampled. ## How to answer in an interview State the mechanics in one line each (copy versus delete), then name the cost of each (memorisation and size versus lost majority variation and instability), then say what you would actually do: check whether the imbalance is even hurting before intervening, prefer undersampling when the majority is large and redundant, prefer oversampling when every row is precious, and treat the resampled ratio as a tuned knob rather than a fixed 1:1 rule. SMOTE, which interpolates new minority points instead of copying them, is the natural next step from random oversampling.

  • When is random undersampling the right call despite throwing data away?
    When the majority class is large and largely redundant, and training cost is the binding constraint. Dropping 190,000 near-identical non-churners rarely removes much signal, and the model trains in a fraction of the time. To recover what was discarded, train several models on different random majority subsets and combine them, so every majority row still influences some member.
  • What does duplicating a minority row do inside a decision tree?
    It multiplies that row's contribution to the class counts at every node it reaches, so impurity calculations change and different splits win. The tree can then justify a deep, pure leaf wrapped tightly around that single location. If the duplicated row happens to be mislabelled, oversampling has amplified the error twenty-fold.
  • Should you always resample all the way to a 1:1 class ratio?
    No. 1:1 is a convention, not a target. On a 0.8% positive rate it means roughly 124 copies of every positive, which is a huge distortion of the training distribution. Lifting the minority to 10% or 20% is often enough, and the ratio should be tuned like any hyperparameter and scored on evaluation data that was left untouched.

Oversampling is photocopying the same few pages so the pile looks thicker; undersampling is binning most of the other book so the two piles match.

saying these in an interview costs you the question

  • Says oversampling adds new information about the minority class
  • Claims undersampling is free because the majority is redundant
  • Assumes a balanced training set always improves the model
  • Thinks duplicated rows cannot cause overfitting because they are real data
  • Treats 1:1 as a required target rather than a tunable ratio

context