skip to content

Class Weights and Thresholds

Cost-sensitive class weights in the loss, per-row sample weights, and when a rare positive is a threshold problem rather than a sampling one. Interviewers ask which remedy you reach for first.

on this pageshow

questions

4

What does a class weight do to a classifier's loss when the positive class is 2% of the rows?

level: middleimportance: must knowfreq 72%

answer

  1. changes the objective, not the data
  2. each rare row counts as many rows
  3. gradient contribution scaled by the weight
  4. the loss optimum moves off the base rate

basics

~20 s

A class weight multiplies every loss term from that class by a constant, so each rare positive counts like many rows. The optimiser trades more errors on the common class for fewer on the rare one, and fitted probabilities rise.

solid answer

~50 s

Training minimises a sum of per-row losses. A class weight turns that into a weighted sum: `L = sum_i w(y_i) * loss_i`, with one constant for positives and another for negatives. No rows are added, removed or invented — only the objective changes. For log loss with a logistic link the per-row gradient with respect to the score is `(p_i - y_i)`, so the weight simply scales that pull to `w * (p_i - y_i)`. At a 2% base rate an unweighted intercept-only fit is already optimal at `p = 0.02`, because 2 positives pulling down and 98 negatives pulling up cancel. Weight the positives 49 times more and the optimum moves to `p = 0.5`. The practical effect is a decision region that reaches further into the rare class: recall rises, precision falls, and the predicted scores are no longer on the natural probability scale.

code

python · 16 lines
python
n_pos, n_neg = 2, 98          # 2% positive rate
n, k = n_pos + n_neg, 2
w_pos = n / (k * n_pos)       # 25.0   balanced rule n / (k * n_c)
w_neg = n / (k * n_neg)       # 0.5102

def pull(p, wp, wn):
    # gradient of the (weighted) log loss wrt the intercept's log-odds
    return wp * n_pos * (p - 1) + wn * n_neg * p

for p in (0.02, 0.5):
    print("p =", p,
          "| unweighted pull:", round(pull(p, 1, 1), 3),
          "| weighted pull:", round(pull(p, w_pos, w_neg), 3))

# p = 0.02 | unweighted pull: 0.0   | weighted pull: -48.0
# p = 0.5  | unweighted pull: 48.0  | weighted pull: 0.0

go deeper

for a junior

Be ready to say in one sentence that a class weight makes errors on the rare class cost more during training, without changing the rows. Know that recall usually goes up and precision usually goes down.

for a middle

You are expected to write the weighted loss, explain that the weight scales each row's gradient contribution, and state the balanced rule n / (k * n_c) with the numbers for a 98/2 split.

for a senior

Show that you know the consequences: distorted output probabilities, unchanged ranking, higher variance on a class that still has only a few real examples, and the quiet interaction with regularisation when the total weight mass changes.

for a principal

Own the framing that the weight ratio is a cost statement. Push the team to derive it from the business cost of the two error types rather than from class frequencies, and to record that ratio somewhere reviewable.

## What is actually being changed A supervised model is fitted by minimising a loss summed over training rows. For binary log loss the per-row term is ``` loss_i = -[ y_i * log(p_i) + (1 - y_i) * log(1 - p_i) ] ``` and the training objective is `L = sum_i loss_i`. Class weighting replaces that with ``` L = sum_i w(y_i) * loss_i ``` where `w(1)` is one constant applied to every positive row and `w(0)` another applied to every negative row. This is important to say out loud in an interview: **the data set is untouched**. Nothing is duplicated, deleted or synthesised. The only thing that changes is how much each mistake costs the optimiser. That is why the technique is usually described as *cost-sensitive learning* rather than a sampling technique. ## Why the rare class gets ignored without it Take a 98/2 split and the simplest possible model, an intercept only, so every row gets the same predicted probability `p`. For log loss with a logistic link, the derivative of a row's loss with respect to the row's score (the log-odds `z`) is exactly `p - y`. Summing over rows, the unweighted gradient is ``` 2 * (p - 1) + 98 * (p - 0) ``` which is zero at `p = 0.02`. The unweighted optimum *is* the base rate; the model is behaving correctly, it simply has almost no reason to predict anything large. Now apply weights and the gradient becomes ``` w_pos * 2 * (p - 1) + w_neg * 98 * p ``` which is zero when `p = (w_pos * 2) / (w_pos * 2 + w_neg * 98)`. If the two classes carry equal total weight, that is `p = 0.5`. In log-odds terms the whole score surface is shifted by `log(w_pos / w_neg)`. With a richer model the shift is not a clean constant, but the direction is the same: the fitted boundary moves so that the rare class occupies more of the input space. ## The balanced rule The common automatic rule sets `w_c = n / (k * n_c)`, where `n` is the number of rows, `k` the number of classes and `n_c` the count of class `c`. On 98 negatives and 2 positives this gives `w_pos = 100 / (2 * 2) = 25` and `w_neg = 100 / (2 * 98) = 0.51`, a ratio of 49:1, which is exactly `n_neg / n_pos`. Each class ends up carrying total weight `n / k = 50`, and the total weight over the whole data set still sums to `n`. That last property matters when a penalty term is involved: because the data term keeps its original scale, the strength of L1 or L2 regularisation relative to the data is unchanged. If instead you hand-set weights of 1 and 49, the total weight mass grows to roughly twice `n`, and the same penalty is now relatively weaker — a silent change in regularisation that people rarely account for. Balanced is a default, not an answer. It targets a 50/50 effective prior, which is the right target only if the cost of a false negative really is `n_neg / n_pos` times the cost of a false positive. When you actually know the costs, encode them; when you do not, the ratio is a knob to tune, not a formula to trust. ## Relation to duplicating rows For an integer weight `w`, adding a minority row `w` times reproduces the weighted loss exactly, so the two are equivalent for a plain full-batch fit. They diverge as soon as anything counts rows rather than weight: bootstrap draws in a bagged ensemble, row subsampling in a booster, a minimum-samples-per-leaf constraint in a tree, mini-batch composition, and cross-validation splits, where duplicated rows leak the same observation into training and validation folds. Weighting is the safer expression of the same intent. ## What it buys, and what it costs It buys a decision region tilted toward the rare class — higher recall, lower precision — and it can genuinely change the fit for learners whose structure depends on class mass, such as a tree whose split criterion is computed on weighted counts, or a margin-based learner deciding where to place the boundary. What it usually does **not** buy is new information: the ordering of scores often barely improves, so a ranking metric may look nearly identical before and after. It costs you two things. First, the effective sample for the rare class is still those few rows, now shouting louder, so variance and overfitting risk on that class rise. Second, the outputs are no longer calibrated to the real base rate, which breaks anything downstream that treats a score as a probability.

  • How are balanced class weights computed from the class counts, and when is that ratio the wrong choice?
    The rule is `w_c = n / (k * n_c)`, so each class ends up with equal total weight — on a 98/2 split that is 25 for positives and about 0.51 for negatives. It is wrong whenever the real cost ratio is not the inverse frequency ratio. If a false negative costs five times a false positive, the target ratio is 5, not 49, and balanced will over-flag badly.
  • Is class weighting the same thing as duplicating every minority row?
    For a plain full-batch loss with an integer weight, yes — the objectives are identical. They differ anywhere rows are counted rather than weighted: bootstrap resampling, row subsampling in a booster, minimum-leaf-size constraints, and cross-validation, where duplicates put the same observation in both the training and validation fold and inflate the score.
  • Would you expect class weights to improve a ranking metric such as area under the ROC curve?
    Usually only marginally. Weighting mostly relocates the decision region rather than teaching the model to separate the classes better, and a ranking metric ignores where the boundary sits. Expect recall up, precision down, ordering roughly unchanged. If the ranking metric moves a lot, the unweighted fit was probably degenerate — for example a tree that never split the minority off at all.

It is a scoring change, not a roster change: the team still plays the same eleven, but goals in one half of the pitch are suddenly worth twenty-five points each, so everyone repositions.

saying these in an interview costs you the question

  • Says class weights add or duplicate minority data
  • Claims weighting always raises area under the ROC curve
  • Assumes weighted output probabilities are still calibrated
  • Treats the balanced rule as the correct ratio by default
  • Thinks weighting fixes having only a handful of positives

context

open as a page

Your class-weighted model outputs scores averaging 0.5 when only 2% of rows are positive — why?

level: seniorimportance: should knowfreq 46%

basics

~20 s

The weights changed the objective, so the fit targets the reweighted class mix rather than the real one. Balanced weights make both classes equally heavy, pulling the optimum toward 0.5. The scores still rank correctly but are no longer probabilities.

open as a page

With a 2% positive rate, how do you decide between class weights, resampling, and moving the operating point?

level: principalimportance: should knowfreq 58%

basics

~20 s

Decide by what is actually broken. If the model already ranks cases well and only the cut-off is wrong, move the operating point - it is free and reversible. Reweight when the fit itself ignores the rare class.

open as a page

When would you give each training row its own weight rather than one weight per class?

level: middleimportance: nice to knowfreq 32%

basics

~20 s

When the cost of an error varies row by row, not just class by class. In insurance-claim triage, weighting each claim by the euro value at risk makes the model spend its capacity where the money is.

open as a page