skip to content

Imbalance and Feature Selection

Class weights versus resampling for a rare positive class, SMOTE and its failure modes, and filter, wrapper and embedded feature selection. Interviewers probe which remedy fits and its cost.

on this pageshow

explore

questions

13

What is the difference between random oversampling and random undersampling of an imbalanced training set?

level: juniorimportance: must knowfreq 74%

answer

  1. one copies rows, one deletes rows
  2. neither adds new information
  3. duplicates invite memorisation
  4. discarded majority rows are gone for that fit
  5. target ratio is a tunable knob

basics

~10 s

Random oversampling duplicates existing minority rows until the classes are closer to balanced. Random undersampling deletes majority rows to reach the same ratio. Oversampling keeps every row but repeats information; undersampling throws information away.

solid answer

~40 s

Both change the class ratio the learner sees, and neither adds information about the rare class. Random oversampling samples minority rows with replacement and appends the copies, so a telecom table with 10,000 churners and 200,000 non-churners grows instead of shrinking; the risk is that a flexible model memorises the duplicated rows, since a duplicate sits exactly on top of the original and widens nothing. Random undersampling drops majority rows at random, so balancing that same table means discarding about 95% of the non-churners; training gets fast and cheap, but genuine majority variation disappears and the boundary moves between runs. I treat undersampling as reasonable when the majority is huge and redundant, oversampling when data is scarce, and I tune the target ratio rather than assuming 1:1.

go deeper

for a junior

Be ready to say in one sentence each what oversampling and undersampling do to the row count, and to state that neither creates new information about the rare class.

for a middle

Explain the mechanics: duplicates multiply a row's weight in the loss or in node class counts, while deletion removes genuine majority variation and raises run-to-run variance.

for a senior

Show that you check whether imbalance is actually hurting before intervening, tune the resampled ratio instead of defaulting to 1:1, and recover discarded majority rows through ensemble undersampling.

for a principal

Own the tradeoff between training cost and information loss across a portfolio of models, and be clear that resampling shifts the operating point rather than adding evidence about the rare class.

## The problem being attacked An imbalanced training set is one where the label you care about is rare: 10,000 churners against 200,000 non-churners in a telecom table, or 400 defective parts in 50,000 inspected units. Most learners minimise an average loss over rows, so a class that supplies 5% of the rows supplies about 5% of the pressure. The model can score well by leaning toward the common label. **Resampling** attacks this by changing the composition of the training rows so the class ratio the learner optimises against is closer to balanced. The two crudest forms are random oversampling and random undersampling. ## Random oversampling: duplicate the minority Draw minority rows uniformly at random **with replacement** and append the copies until the target ratio is reached. Nothing is invented — repeating the same 400 defective parts twenty times still describes exactly 400 distinct defects, and gets you to 8,000 minority rows against 49,600 good ones, which is not even balanced. Mechanically, a duplicated row contributes to the objective once per copy. In a tree, node class counts change, so different splits are chosen and a deep, pure leaf drawn tightly around a duplicated point becomes easy to justify — including around a mislabelled one. In a margin- or distance-based model, the copies occupy the *same* coordinates as the original, so they cannot enlarge the region the minority covers; they only raise the price of getting that one location wrong. Arithmetically, appending a row twenty times is the same as counting it twenty times in the loss. Consequences worth naming: - **Memorisation.** With enough capacity the model fits the individual duplicated rows, so training scores look excellent and held-out scores do not move. - **Cost.** The training set grows. Balancing the telecom table by oversampling produces 400,000 rows instead of 210,000. - **No new coverage.** Minority regions that were never sampled stay unsampled. ## Random undersampling: delete the majority Drop majority rows at random until the ratio is reached. Balancing the telecom table means keeping 10,000 of the 200,000 non-churners and discarding roughly 95% of them. - **It is fast.** The training set falls to 20,000 rows, which matters when you are fitting many candidate models. - **The loss is real and one-directional.** Unusual-but-genuine majority rows — the rare corners of "normal" — may vanish entirely, so the model's picture of the common class gets thinner and the boundary shifts noticeably between random draws. That is added variance, not just lost accuracy. - **The standard mitigation is not to discard permanently.** Train several models on different random majority subsets and combine them, so every majority row is seen by some ensemble member. Ensemble undersampling methods such as EasyEnsemble are built on exactly this idea. ## What neither one does Neither adds information about the rare class. What they change is the ratio the objective is averaged over, which makes the model more willing to emit the rare label. A direct consequence: the scores a resampled model produces reflect the resampled ratio, not the population rate — so a raw probability from such a model should not be read as a real-world frequency. ## Choosing how far to go Full balance is a convention, not a requirement, and on extreme ratios it is a violent intervention: a 0.8% defect rate needs about 124 copies of every defect to reach 1:1. Partial resampling — lifting the minority to 10% or 20% — is usually the better starting point, and the target ratio deserves to be tuned like any other hyperparameter and judged on evaluation data you never resampled. ## How to answer in an interview State the mechanics in one line each (copy versus delete), then name the cost of each (memorisation and size versus lost majority variation and instability), then say what you would actually do: check whether the imbalance is even hurting before intervening, prefer undersampling when the majority is large and redundant, prefer oversampling when every row is precious, and treat the resampled ratio as a tuned knob rather than a fixed 1:1 rule. SMOTE, which interpolates new minority points instead of copying them, is the natural next step from random oversampling.

  • When is random undersampling the right call despite throwing data away?
    When the majority class is large and largely redundant, and training cost is the binding constraint. Dropping 190,000 near-identical non-churners rarely removes much signal, and the model trains in a fraction of the time. To recover what was discarded, train several models on different random majority subsets and combine them, so every majority row still influences some member.
  • What does duplicating a minority row do inside a decision tree?
    It multiplies that row's contribution to the class counts at every node it reaches, so impurity calculations change and different splits win. The tree can then justify a deep, pure leaf wrapped tightly around that single location. If the duplicated row happens to be mislabelled, oversampling has amplified the error twenty-fold.
  • Should you always resample all the way to a 1:1 class ratio?
    No. 1:1 is a convention, not a target. On a 0.8% positive rate it means roughly 124 copies of every positive, which is a huge distortion of the training distribution. Lifting the minority to 10% or 20% is often enough, and the ratio should be tuned like any hyperparameter and scored on evaluation data that was left untouched.

Oversampling is photocopying the same few pages so the pile looks thicker; undersampling is binning most of the other book so the two piles match.

saying these in an interview costs you the question

  • Says oversampling adds new information about the minority class
  • Claims undersampling is free because the majority is redundant
  • Assumes a balanced training set always improves the model
  • Thinks duplicated rows cannot cause overfitting because they are real data
  • Treats 1:1 as a required target rather than a tunable ratio

context

open as a page

What does a class weight do to a classifier's loss when the positive class is 2% of the rows?

level: middleimportance: must knowfreq 72%

basics

~20 s

A class weight multiplies every loss term from that class by a constant, so each rare positive counts like many rows. The optimiser trades more errors on the common class for fewer on the rare one, and fitted probabilities rise.

open as a page

How do filter, wrapper and embedded feature selection differ in cost and in what they can see?

level: middleimportance: must knowfreq 68%

basics

~20 s

Filters rank columns with a cheap statistic and never train the model. Wrappers train it many times on candidate subsets and keep the best. Embedded selection falls out of one model fit. Cost and model-specificity rise across the three.

open as a page

How does SMOTE create a synthetic minority sample from existing training rows?

level: middleimportance: must knowfreq 78%

basics

~20 s

SMOTE picks a minority row, finds its k nearest neighbours among other minority rows, chooses one of them at random, and places a new point at a random fraction of the way along the straight line between the two.

open as a page

Why can a correlation or mutual-information filter keep two 0.97-correlated columns yet drop a feature that only matters in combination?

level: middleimportance: should knowfreq 52%

basics

~20 s

A correlation or mutual-information filter scores each column against the target on its own. Two near-duplicates both score well and survive; a feature that matters only alongside another scores near zero alone and is cut. The blind spot is univariate scoring.

open as a page

What goes wrong when SMOTE interpolates over one-hot encoded categorical features?

level: middleimportance: should knowfreq 46%

basics

~10 s

Interpolation produces fractional values on indicator columns, so a synthetic row comes out 0.5 on two different category flags. It represents no real category, and no such row could ever arrive at prediction time.

open as a page

Your class-weighted model outputs scores averaging 0.5 when only 2% of rows are positive — why?

level: seniorimportance: should knowfreq 46%

basics

~20 s

The weights changed the objective, so the fit targets the reweighted class mix rather than the real one. Balanced weights make both classes equally heavy, pulling the optimum toward 0.5. The scores still rank correctly but are no longer probabilities.

open as a page

Recursive feature elimination over 400 marketing-attribution columns runs for hours — what is it doing and how do you cut the cost?

level: seniorimportance: should knowfreq 44%

basics

~20 s

Recursive feature elimination refits the model, ranks columns by its weights, drops the weakest, and repeats — hundreds of fits, times cross-validation folds. Cut cost by dropping a block per round, pre-screening with a cheap filter, or ranking with a faster model.

open as a page

Forward stepwise selection picked 8 of 60 sensor channels; why is the winning subset's cross-validated score optimistic?

level: seniorimportance: should knowfreq 38%

basics

~20 s

Forward stepwise selection evaluated hundreds of candidate subsets and reported the winner. The maximum of many noisy estimates is biased upward, so part of the winner's margin is luck — even when the selector saw only training rows.

open as a page

When does adding SMOTE samples make a classifier's decision boundary worse?

level: seniorimportance: should knowfreq 50%

basics

~20 s

SMOTE hurts when the classes overlap or the minority contains mislabelled points. Interpolating between two minority rows that sit on opposite sides of a majority cloud plants synthetic positives inside majority territory, pushing the boundary outward and costing precision.

open as a page

With a 2% positive rate, how do you decide between class weights, resampling, and moving the operating point?

level: principalimportance: should knowfreq 58%

basics

~20 s

Decide by what is actually broken. If the model already ranks cases well and only the cut-off is wrong, move the operating point - it is free and reversible. Reweight when the fit itself ignores the rare class.

open as a page

What does a variance-threshold filter remove from a feature matrix, and when does it drop something useful?

level: juniorimportance: nice to knowfreq 30%

basics

~20 s

A variance-threshold filter drops every column whose spread across rows falls below a cutoff, removing constant and near-constant features. It ignores the target and depends on units, so it can discard a rare binary flag that predicts strongly.

open as a page

When would you give each training row its own weight rather than one weight per class?

level: middleimportance: nice to knowfreq 32%

basics

~20 s

When the cost of an error varies row by row, not just class by class. In insurance-claim triage, weighting each claim by the euro value at risk makes the model spend its capacity where the money is.

open as a page