When would you give each training row its own weight rather than one weight per class?
answer
- cost varies within a class, not just between
- class weight is the constant-within-class case
- also used to undo a sampling design
- watch a few huge weights dominate
- effective sample size from squared weights
basics
~20 sWhen the cost of an error varies row by row, not just class by class. In insurance-claim triage, weighting each claim by the euro value at risk makes the model spend its capacity where the money is.
solid answer
~50 sA class weight assumes every error on a class costs the same. Often it does not: in claim triage a missed high-value claim can cost a thousand times a missed small one, so the natural weight is per row — the amount at risk. A class weight is just the special case where the weight is constant within a class. The second common use is correcting a known sampling bias, where the weight is the inverse of the row's probability of having been sampled, undoing a stratified draw. The main hazard is heavy-tailed weights: if a handful of rows carry most of the mass, the effective sample size collapses and the fit rides on those few rows. Check `n_eff = (sum w)^2 / sum(w^2)`, and consider capping or log-scaling the weights. Whatever weights you train on, evaluate with them too.
go deeper
Know that weights can be set per row and not only per class, and that the usual reason is that some rows are more expensive to get wrong than others.
Explain that class weighting is the special case of constant-within-class row weights, name at least two legitimate uses, and know that the weighted objective must be the one you evaluate.
Show the operational care: check effective sample size against heavy-tailed weights, cap or compress them, keep the weight free of any outcome information, and normalise the scale so a penalty term still bites.
Decide whether the cost belongs in the model at all. Separating a probability model from a severity model is often cleaner than one weighted fit, and the choice determines what the team can audit later.
## The generalisation Weighted training minimises `L = sum_i w_i * loss_i` with one weight per row. Class weighting is the special case where `w_i` depends only on `y_i`. Once you see it that way, the interesting question is what else the weight could encode. ## Use one: per-row cost The canonical case is a cost that varies within a class. An insurance-claim triage model predicts whether a claim needs manual investigation. Every missed fraudulent claim is a false negative, but one is worth 400 euro and another 400,000. Class weighting treats them identically, so the model happily learns to catch the abundant cheap cases and misses the rare expensive ones — technically respectable, financially useless. Setting `w_i` to the euro value at risk makes the objective approximate expected monetary loss rather than an error count, and the fitted model reallocates its capacity toward the rows where being wrong is expensive. The same shape appears wherever the loss unit is not the row: a churn model where customers have very different lifetime values, a demand forecast where some stores dominate revenue, a triage model where clinical severity varies by patient. ## Use two: undoing a known sampling design If the training data was not drawn uniformly — a stratified sample, a survey with unequal selection probabilities, a log that kept all positives but only 1 in 20 negatives — the fitted model reflects the sampled mix, not the real one. Weighting each row by the inverse of its selection probability restores the original population inside the objective. The negatives kept at 1 in 20 get a weight of 20. This is a correctness fix, not a tilt: you are recovering the distribution you meant to fit, and it is the honest way to train on a down-sampled log. ## Use three: confidence and recency Weights can also encode how much you trust a row. Labels from expert adjudication might carry weight 1 and labels from a heuristic rule weight 0.3. In a drifting environment, weights decaying with age make recent rows count more without hard-cutting the history. Both are legitimate, and both are judgement calls you should be able to defend rather than tune blindly. ## The failure mode: effective sample size Heavy-tailed weights are the thing to worry about. Monetary values are typically log-normal or worse, so a few enormous claims can carry most of the total weight. The fit then depends on a handful of rows, variance explodes, and cross-validation scores swing wildly between folds depending on which giant landed where. The standard diagnostic is the effective sample size ``` n_eff = (sum_i w_i)^2 / sum_i (w_i^2) ``` which equals `n` when all weights are equal and collapses toward 1 as one weight dominates. If 200,000 rows give `n_eff` of 300, you are not training on 200,000 rows in any meaningful sense. Remedies: cap weights at a high quantile (winsorise), use a compressive transform such as `log(1 + amount)` instead of the raw amount, or split the problem — model the probability of the event and the size of the loss separately, then combine them at decision time. That last option is often the cleanest, because it stops one model from having to learn two different things. ## Other things to get right **The weight must be known before the outcome.** The value at risk of a claim is known when the claim is filed, so it is legitimate. The amount eventually paid out is not — it is a function of the label, and using it leaks the target into training in a way that is easy to miss because it enters through the weight rather than through a feature. **Evaluate with the same weights.** If training optimises euro-weighted loss, an unweighted error count on the test set measures something you did not ask for. Report the weighted metric, and usually the unweighted one alongside it, so you can see the tradeoff you bought. **Scale interacts with regularisation.** Raw monetary weights can sum to millions, dwarfing a penalty term calibrated for weights near 1. Normalise the weights so they average 1 — the ratios are what matter, not the absolute scale. **Weights are not features.** Putting the amount at risk in as a feature tells the model to predict differently for large claims; putting it in as a weight tells the model that being wrong on large claims is worse. They answer different questions, and sometimes you want both.
- How do you detect that a few rows dominate a weighted fit?Compute the effective sample size, `(sum w)^2 / sum(w^2)`. It equals the row count when all weights are equal and collapses toward one when a single weight dominates. If it is a tiny fraction of your data, the model is effectively trained on a handful of rows. Cap the weights at a high quantile or compress them with a log transform, then recheck.
- Should the amount at risk go in as a weight, as a feature, or both?They mean different things. As a feature it lets the model predict a different probability for large claims; as a weight it says being wrong on large claims costs more. Both are often right, and using both is fine as long as you can say why. What is never fine is a weight derived from the outcome itself, which leaks the label through the back door.
- Can per-row weights and class weights be combined?Yes — multiply them, since the objective takes one number per row. A common combination is a monetary weight for cost and a class factor for imbalance. The risk is losing track of the effective ratio between the two classes once both are applied, so compute the resulting total weight per class and check it against the tilt you intended.
saying these in an interview costs you the question
- Treats per-row weights as interchangeable with features
- Uses an outcome-derived quantity as the weight
- Ignores that a few huge weights dominate the fit
- Trains on weighted loss, evaluates on unweighted counts
- Feeds raw monetary weights alongside a fixed penalty term