skip to content

Your class-weighted model outputs scores averaging 0.5 when only 2% of rows are positive — why?

level: seniorimportance: should knowfreq 46%

answer

  1. weighting imitates a different class mix
  2. the imagined base rate is not the real one
  3. a constant offset on the log-odds
  4. ranking metrics hide it completely
  5. log of the weight ratio, about 3.9 at 49:1

basics

~20 s

The weights changed the objective, so the fit targets the reweighted class mix rather than the real one. Balanced weights make both classes equally heavy, pulling the optimum toward 0.5. The scores still rank correctly but are no longer probabilities.

solid answer

~50 s

Weighting a loss is the same as pretending the training population has a different class mix. With balanced weights on a 98/2 split, each class carries equal total weight, so the model behaves as if the base rate were 50%, and its scores land near 0.5 on average. In log-odds terms the whole score surface is offset by roughly `log(w_pos / w_neg)` — about 3.9 for a 49:1 ratio. Ranking metrics are invariant to a monotone shift, so dashboards look unchanged and nobody notices until something downstream consumes the number as a probability: an expected-value calculation, a volume forecast, a cut-off copied from another model, or an average blended with another model's scores. Fix it by subtracting the offset, by calibrating on an unweighted holdout with the real class mix, or by not weighting and moving the operating point instead.

go deeper

for a junior

Know that once you weight the classes, the numbers coming out are ranking scores rather than real probabilities, and that a score near 0.5 on a rare event is expected, not broken.

for a middle

Be able to explain that weighting imitates a different class mix, that the effect on a logistic model is a constant log-odds offset of log(w_pos / w_neg), and roughly how large that is at 49:1.

for a senior

Show the diagnosis and the remedy: recognise that ranking metrics hide the problem, name the downstream consumers that break, and calibrate on a holdout with the untouched class balance rather than on rebalanced data.

for a principal

Own the contract. Decide whether the model ships a score or a probability, make that explicit to every consumer, and require a calibration check on natural-balance data before any output is multiplied by money.

## Weighting is a change of prior in disguise Multiplying every positive row's loss by `w_pos` and every negative row's by `w_neg` is, for most standard losses, indistinguishable from having sampled a training set in which positives were `w_pos / w_neg` times more common. The fitted model therefore answers a different question from the one you think you asked: not "what fraction of rows like this one are positive in the real world?" but "what fraction would be positive in a world where positives were 49 times more frequent?" With the balanced rule on a 98/2 split, the two classes end up with identical total weight, so that imagined world has a 50% base rate. An intercept-only fit lands at exactly 0.5, and a real model's scores spread around a far higher centre than the 2% you would expect. ## The size of the distortion For a logistic model the distortion is a constant offset on the log-odds scale. The weighted fit satisfies ``` logit_weighted(x) = logit_true(x) + log(w_pos / w_neg) ``` With `w_pos = 25` and `w_neg = 0.51`, the ratio is 49 and the offset is `log(49) = 3.89`. A row whose true probability is 2% has true log-odds of about `-3.89`; add the offset and it comes out at 0, that is, a score of 0.5. The arithmetic is exact for a linear model with an intercept and approximate for a flexible model, but the direction and rough magnitude carry over: every score is inflated, and small true probabilities are inflated the most in absolute terms. ## Why nobody notices for weeks This is the part interviewers reward. A constant log-odds offset is a strictly monotone transformation of the scores, so it does not change the ordering of any two rows. Every ranking-based number — area under the ROC curve, any metric computed by sweeping a cut-off — is unchanged. The offline evaluation looks healthy. The distortion only surfaces where the absolute value of the score is used: - **Expected-value decisions.** "Act when score times loss-avoided exceeds the cost of acting" silently over-triggers when the score is inflated fivefold. - **Volume forecasts.** Summing scores to predict how many events will occur is a standard trick and it now produces wild over-estimates. Summing 0.5 across a thousand rows predicts 500 events where 20 will occur. - **Transferred cut-offs.** A 0.8 cut-off tuned on the previous unweighted model means something completely different on the weighted one. - **Blending.** Averaging this model's scores with a calibrated model's scores mixes two different scales. - **Human trust.** An analyst reading "this patient has a 0.62 probability of the event" against a 0.3% prevalence will either panic or, worse, learn to ignore the number. A hospital sepsis-alert model fitted at 0.3% prevalence with a 1:200 weighted log loss is the sharp case: clinically the alert may be tuned fine, but the printed probability is meaningless, and clinicians calibrate their own trust against numbers they can sanity-check. ## The three fixes **Undo the offset analytically.** For a model with a linear score and an intercept, subtract `log(w_pos / w_neg)` from the log-odds before converting back to a probability. Cheap and exact when the assumption holds; unreliable for a flexible model where the weighting has changed more than the intercept. **Calibrate on a held-out set with the real class mix.** Fit a one-dimensional mapping from the model's score to an observed frequency on data that was never reweighted or resampled — a logistic recalibration of the score, or an isotonic fit if you have enough positives and want a shape-free mapping. This is the general-purpose answer and it works regardless of what the weighting did. The catch at 2% prevalence is that you need enough positives in the calibration set for the mapping to be stable, which usually means a larger holdout than you would otherwise keep. **Do not weight at all.** If the only reason for the weights was to make the model flag more cases, and the ranking was already good, fit unweighted and move the operating point instead. You keep honest probabilities and lose nothing. ## The rule to state out loud A weighted model produces a *score*, not a *probability*, until something restores the scale. Decide up front which of the two the consumer needs. If any downstream step multiplies the number by money, sums it, thresholds it against an externally chosen value, or shows it to a human, calibration is not optional — and it must be measured on data with the untouched class balance.

  • Which offline metrics would have caught this, and which would not?
    Anything computed from the ordering alone would not: area under the ROC curve and any cut-off sweep are invariant to a monotone shift, so they look identical before and after weighting. What catches it is a metric that reads the absolute score — log loss or a Brier score computed on a holdout with the real class mix, or simply comparing the mean predicted score against the observed positive rate.
  • How would you restore usable probabilities without giving up the weighted fit?
    Hold out a set that was never reweighted or resampled and fit a one-dimensional mapping from score to observed frequency — a logistic recalibration, or an isotonic fit if the positive count supports it. At a 2% base rate the constraint is positives, not rows, so size the calibration set by how many positives it contains, not by percentage.
  • Does the same distortion appear if you undersample the majority class instead of weighting?
    Yes, and for the same reason. Undersampling changes the class mix the model is fitted on, so the fitted scores target that altered prior. Any prior-changing intervention — weighting, undersampling, oversampling — has to be undone before the output is treated as a probability.

It is like a bathroom scale nudged up by ten kilos: everyone's relative order is intact, so weigh-off contests still work, but the moment someone doses a drug by body weight the error matters.

saying these in an interview costs you the question

  • Calls the inflated scores a bug in the training code
  • Says area under the ROC curve would have revealed it
  • Calibrates on rebalanced or reweighted holdout data
  • Rescales scores by dividing them by a constant
  • Believes only badly fitted models are miscalibrated

context