skip to content

With a 2% positive rate, how do you decide between class weights, resampling, and moving the operating point?

level: principalimportance: should knowfreq 58%

answer

  1. three levers, three moments in the pipeline
  2. diagnose ranking separately from cut-off
  3. the free lever is the post-training one
  4. reweight only when the fit is degenerate
  5. never stack all three tilts at once

basics

~20 s

Decide by what is actually broken. If the model already ranks cases well and only the cut-off is wrong, move the operating point - it is free and reversible. Reweight when the fit itself ignores the rare class.

solid answer

~50 s

Start from an unweighted fit and separate two questions: is the ranking good, and is the operating point right? If ranking is acceptable and the only complaint is that too few cases are flagged, the answer is the operating point — it is a post-training decision, costs one number, keeps the scores interpretable as probabilities, and can be revisited weekly without retraining. Reweighting is for when the fit itself is degenerate: a learner whose optimum ignores the rare class entirely, a tree that never isolates it, or a genuine cost asymmetry you want the model to internalise. Resampling earns its place mainly on scale and pipeline grounds — undersampling a huge majority to make training tractable, at the price of discarded information and a shifted prior. Whatever you pick, choose the weight ratio or the cut-off on a validation split and evaluate on data with the untouched natural class balance.

go deeper

for a junior

Know that there are three distinct remedies for a rare positive class and that changing where you cut the score is the cheapest of them because it needs no retraining.

for a middle

Be able to place each lever in the pipeline — before, during, after training — and to explain that only the post-training one leaves the fitted model and its probabilities untouched.

for a senior

Demonstrate the diagnosis: check whether the ranking is good before touching the fit, apply training-time changes inside the training fold only, and evaluate on the natural class balance at matched operating points.

for a principal

Own the policy. Decide which lever the team reaches for by default, insist the cost ratio is written down rather than inferred from class frequencies, and make sure nobody stacks three tilts and calls the result an improvement.

## The three levers are not interchangeable They act at three different moments and have three different costs. - **Moving the operating point** happens after training. The model and its scores are untouched; you only change the score above which you act. Nothing about the fit is at risk, and the change is reversible in seconds. - **Reweighting** happens during training. It changes the objective, so it changes the fitted parameters. The scores are no longer on the natural probability scale. - **Resampling** happens before training. It changes the data the model sees, so it changes the fit *and* the effective prior, and it can throw information away or reuse the same rows many times. An interviewer asking this wants to hear the ordering logic, not a preference. ## The diagnosis that comes first Fit the model with no intervention, then ask two separate questions on a validation split. 1. **Does it rank?** Look at whether high-scoring cases really do contain most of the positives. If the ordering is good, the model has learnt the signal, and the low flag rate is purely a cut-off problem. 2. **Is the operating point wrong?** If yes, that is a decision-layer problem with a decision-layer fix. A large majority of "the model does not detect the rare class" complaints resolve at step 2. That is the single most useful thing to say: the default cut-off of 0.5 is an arbitrary inheritance, and on a 2% base rate almost nothing will ever exceed it. You do not have a training problem; you have a cut-off that nobody chose deliberately. Which value to choose, and from what — expected cost, a required recall, the size of the review queue — is a threshold-selection exercise in its own right. ## When reweighting is the right lever Reweighting earns its place in three situations. First, when the fit is genuinely degenerate rather than just conservatively scored. Some learners collapse under extreme imbalance: a shallow tree may never make a split that isolates the rare class, a margin-based learner may place the boundary where no positive is on the correct side, and there is then no cut-off that recovers useful behaviour because the ordering itself is uninformative. Second, when you have a real cost asymmetry and you want the model to internalise it during fitting, so that it spends its limited capacity on the errors that matter. A cut-off can only re-slice the ranking the model already produced; the weight can change what the model chose to learn. Third, when the downstream consumer wants a single decision, not a score, and calibration is irrelevant. The cost is that the outputs stop being calibrated, and that a class with very few rows now dominates the objective, which raises variance and the risk of fitting noise in those few rows. ## When resampling is the right lever Undersampling the majority is defensible when the majority is enormous and training time or memory is the binding constraint — dropping 90% of an overwhelming negative class can turn an overnight fit into a ten-minute one at little cost in signal, because the marginal negative row carries almost no information. It shifts the prior, so the same calibration caveat applies, and it discards data you paid for. Oversampling by duplication is the weakest of the three: for most losses it is equivalent to weighting, but it inflates row counts, corrupts cross-validation if applied before splitting, and interacts badly with anything that subsamples rows. If duplication is what you mean, weight instead. Synthesising new minority points is a different mechanism with its own tradeoffs and belongs to that discussion, not this one. ## Stacking, and how to compare fairly Applying all three at once is the classic mistake. Weighting 49:1 *and* balancing the classes by resampling *and* dropping the cut-off compounds three tilts toward the rare class, and the result flags almost everything. Pick one primary lever, and treat any second one as a deliberate, measured addition. Whatever you pick, the comparison protocol has to be honest. Resampling and reweighting are training-time choices, so they must be applied inside the training fold only, never to validation or test data. Evaluation always happens on data with the real class balance — otherwise you are measuring performance on a population that does not exist. And report the comparison at a matched operating point or across the full range of them, because a reweighted model compared at a fixed 0.5 cut-off against an unweighted one is not a comparison of models at all; it is a comparison of cut-offs. ## The answer that lands "Move the operating point first, because it is free, reversible and keeps the scores meaningful. Reweight when the fit itself is broken or when I know the cost ratio. Resample when the data volume forces my hand. And I never do more than one of them without measuring what the second one added."

  • Under what condition can no choice of operating point rescue the model, so you have to change the fit?
    When the ranking itself is uninformative — high-scoring rows contain no more positives than low-scoring ones. A cut-off can only slice an existing ordering, so if the ordering carries no signal, every cut-off trades recall for precision along a useless line. That is the signature of a degenerate fit and the case where reweighting or better features are the only options.
  • Is it ever reasonable to apply both class weights and undersampling of the majority?
    Yes, but deliberately. A common pattern is undersampling for tractability on an enormous majority class, then a modest weight to encode a cost ratio the sampling did not already deliver. The trap is applying the balanced ratio on top of already-balanced data, which tilts twice. Compute the effective end-to-end ratio and check it against the cost ratio you actually intend.
  • How should the comparison between the three options be measured?
    On a validation set with the untouched natural class balance, never on resampled or reweighted validation data. Compare at matched operating points or across the whole range of them, since a reweighted model judged at a fixed 0.5 cut-off is being penalised for a cut-off nobody chose. Report the error counts the business cares about, not a single blended score.

saying these in an interview costs you the question

  • Reaches for resampling before checking the cut-off
  • Applies weighting, resampling and a lower cut-off together
  • Evaluates on rebalanced validation data
  • Says imbalance always requires a training-time fix
  • Compares a weighted and unweighted model at the same 0.5 cut-off

context