skip to content

A model trained on 1:1 undersampled data outputs 0.5 when the live positive rate is 4%. How do you fix the probabilities?

level: seniorimportance: should knowfreq 40%

answer

  1. only the prior changed, not the shapes
  2. work on the odds, not the probability
  3. one constant shift in log-odds
  4. ratio of true odds to training odds
  5. 0.5 should land on the base rate

basics

~20 s

Balancing the training data shifted the log-odds by a constant. Undo it: multiply the predicted odds by the ratio of true prior odds to training prior odds. A 0.5 score becomes 0.04 at a 4% base rate.

solid answer

~50 s

Undersampling the majority class changes the class prior but leaves the class-conditional feature distributions alone, so the model's log-odds are off by one constant everywhere. The fix is a prior correction: take the predicted odds `p/(1-p)`, multiply by `(pi_true/(1-pi_true)) * ((1-pi_train)/pi_train)`, and convert back to a probability. With `pi_train = 0.5` and `pi_true = 0.04` the factor is 0.04/0.96, so 0.5 maps to 0.04 and 0.9 maps to about 0.27 - the base rate is restored at the middle of the range. It is a strictly increasing map, so ranking and AUC are untouched; only the numbers move. I would then verify on a held-out sample drawn at the natural 4% prevalence with a reliability diagram, and fall back to fitting a calibrator on that natural-prevalence sample if the resampling did more than change proportions.

code

python · 16 lines
python
def prior_correct(p_model, train_rate, true_rate):
    # Undersampling shifted every log-odds by one constant.
    # Multiply the predicted odds by the ratio of the two prior odds.
    odds = p_model / (1.0 - p_model)
    factor = (true_rate / (1.0 - true_rate)) * ((1.0 - train_rate) / train_rate)
    fixed = odds * factor
    return fixed / (1.0 + fixed)


# model trained at 50% positives, deployed where the true rate is 4%
for score in (0.2, 0.5, 0.9):
    print(score, '->', round(prior_correct(score, 0.5, 0.04), 4))

# 0.2 -> 0.0103
# 0.5 -> 0.04     (an unremarkable row lands on the base rate)
# 0.9 -> 0.2727

go deeper

for a junior

Recall the symptom and its cause: training on an artificially balanced sample makes scores hover near 0.5 even when positives are rare, because the model was taught the wrong base rate.

for a middle

Derive the fix. Odds factor into a prior term times a likelihood ratio, resampling touched only the prior, so multiplying the predicted odds by the ratio of the two prior odds restores the scale.

for a senior

Show the operational discipline: verify on a natural-prevalence holdout rather than the balanced one, re-derive any cutoff on the corrected scale, and recognise when the single-constant assumption no longer holds.

for a principal

Own the upstream call. Argue when rebalancing is worth its cost at all given that probabilities are usually the deliverable, and set whether the team corrects analytically or reserves natural-prevalence data for a fitted calibrator.

## Why balancing breaks the probabilities By Bayes' rule, the odds a row is positive factor into a prior term and a likelihood-ratio term: `odds(x) = [pi / (1 - pi)] * [f1(x) / f0(x)]` where `pi` is the class prior and `f1`, `f0` are the feature distributions within the positive and negative classes. Randomly dropping majority rows to reach a 1:1 ratio changes `pi` from 0.04 to 0.5 and, in expectation, leaves `f0` and `f1` untouched - a uniform random subsample of the negatives has the same shape as the full set of negatives. So the model trained on the balanced data estimates the right likelihood ratio and the wrong prior. Its odds are the correct odds multiplied by a constant, or equivalently its log-odds are the correct log-odds plus a constant shift. This is not a modelling failure and it is not noise. It is a deterministic offset with a closed-form inverse. The symptom is unmistakable: scores pile up around the middle of the range, the mean predicted probability across the population is roughly 0.5 while only 4% of rows are positive, and the reliability diagram sits far below the diagonal everywhere. ## The correction Let `pi_train` be the positive rate the model was trained at and `pi_true` the rate it will meet. For a predicted probability `p`: ``` odds = p / (1 - p) factor = (pi_true / (1 - pi_true)) * ((1 - pi_train) / pi_train) odds_fixed = odds * factor p_fixed = odds_fixed / (1 + odds_fixed) ``` With a 1:1 training set the second bracket is 1 and the factor collapses to the true prior odds, `0.04 / 0.96 = 0.0417`. Then: - `p = 0.5` gives odds 1, corrected odds 0.0417, `p_fixed = 0.04` - the base rate, exactly as it should be for a row the model finds unremarkable. - `p = 0.9` gives odds 9, corrected odds 0.375, `p_fixed = 0.27`. - `p = 0.2` gives odds 0.25, corrected odds 0.0104, `p_fixed = 0.010`. For a linear model on the log-odds scale this is simply subtracting a constant from the intercept; the coefficients are already right. ## What the correction does and does not touch - **Ranking is untouched.** The map is strictly increasing in `p`, so every pairwise comparison survives and AUC is unchanged to the decimal. Do not expect the correction to improve discrimination - it cannot. - **Decision cutoffs move.** A rule written against balanced scores no longer means what it did once the scale is restored; the cut has to be re-derived on the corrected scale. - **Aggregate quantities become usable.** Summing corrected probabilities over a population now estimates the expected number of positives; summing the balanced ones would have predicted half the book. ## When the closed form is not enough The derivation rests on one assumption: the resampling changed only the class proportions. That holds for uniform random removal of majority rows. It does **not** hold when the sampling was stratified or non-uniform, when the minority side was expanded with invented rows whose feature distribution differs from real positives, or when the training snapshot was drawn from a different period than the scoring population. In those cases the shift is no longer a single constant, and forcing the closed form gives a curve that is right at one point on the range and wrong elsewhere. The robust alternative, and a good belt-and-braces step even when the assumption holds, is to fit a monotone recalibration map on a **held-out sample drawn at the natural prevalence** - not on the balanced data. That sample has to be large enough to contain a workable number of positives, which is the real constraint at a 4% rate: ten thousand rows carry only about four hundred positives, and the top bins of the reliability diagram will still be thin. ## Verifying Evaluate on a natural-prevalence holdout, never on the balanced one, because the balanced set will make uncorrected scores look fine and corrected ones look terrible. Plot the reliability diagram on that holdout and check three things: the mean corrected probability across all rows should land near 0.04; the diagram should track the diagonal in the score region where decisions are made; and the per-bin counts should show whether the high-score bins hold enough positives to say anything at all. ## The prior question worth asking If probabilities are the deliverable, the cleanest option is often not to rebalance at all. Training on the natural distribution with the full data produces scores that need no correction, and modern learners handle a 4% positive rate without help. Undersampling buys training speed and sometimes better behaviour from a particular learner; it costs data and it costs the meaning of the output. Prior correction is how you buy the meaning back, and it is worth knowing precisely so the choice to rebalance is made with eyes open.

  • Does the prior correction change the model's AUC?
    No. Multiplying every odds by the same positive constant is a strictly increasing transform of the score, so every pairwise ordering is preserved and AUC is identical to the decimal. That is the whole point: undersampling did not damage the ranking, only the scale, and the correction repairs only the scale. Anyone expecting a discrimination gain has misunderstood what was broken.
  • When does the closed-form correction fail, and what do you do instead?
    It assumes only the class proportions changed. Non-uniform sampling, invented minority rows whose feature distribution differs from real positives, or a training snapshot from a different period all break the single-constant shift. Then fit a monotone calibration map on a held-out sample drawn at the natural prevalence instead, and confirm with a reliability diagram on that same natural-rate data.
  • How do you verify the correction actually worked?
    On a holdout drawn at the real 4% prevalence, never the balanced set. Check that the mean corrected probability across all rows lands near 0.04, that the reliability diagram tracks the diagonal in the score band where decisions are taken, and that the high-score bins hold enough positives to be worth reading at all.

The model learned how much more suspicious one row is than another, but was told the crowd is half criminals. Correcting the prior tells it the crowd is 4% criminals; every relative judgement it made stays intact.

saying these in an interview costs you the question

  • Says balancing improves the probabilities rather than distorting them
  • Expects the correction to raise AUC
  • Applies the correction to the probability instead of the odds
  • Measures calibration on the balanced holdout
  • Assumes any resampling scheme leaves a single constant shift
  • Keeps a decision cutoff tuned on the balanced scale

context