skip to content

How does reweighing training rows shrink a group disparity before any model is fitted?

level: middleimportance: nice to knowfreq 30%

answer

  1. Cross-tabulate group against label
  2. Compare observed mass to independence
  3. Expected count over observed count
  4. Thin favourable cells get pushed up
  5. Group column not needed at scoring time

basics

~20 s

Each training row gets a weight equal to the count its (group, outcome) cell would have if group and label were independent, divided by the count actually observed. The weighted data carries no group-label association, and scoring never needs the protected column.

solid answer

~50 s

Cross-tabulate the training rows by protected group and by label. For a cell holding group `s` and label `y`, the weight is `w = P(S=s) * P(Y=y) / P(S=s, Y=y)` - the mass the cell would carry under independence, over the mass it actually carries. Cells that are under-represented, typically the disadvantaged group's favourable outcomes, get weights above one; over-represented cells get weights below one. Fitting on the weighted rows removes the group-label association from what the learner sees, which pushes the model towards equal selection rates. It requires a learner that accepts per-row weights, and it only touches the association between group and label - features that carry the same information indirectly are untouched. Two costs: the effective sample size falls, so variance rises, and the residual gap must still be measured on held-out data.

code

python · 14 lines
python
from collections import Counter

# (protected group, label) for 1000 training rows
rows = ([("A", 1)] * 400 + [("A", 0)] * 100
        + [("B", 1)] * 100 + [("B", 0)] * 400)

n = len(rows)
cell = Counter(rows)
by_group = Counter(g for g, _ in rows)
by_label = Counter(y for _, y in rows)

for (g, y), observed in sorted(cell.items()):
    expected = by_group[g] * by_label[y] / n
    print(g, y, "observed", observed, "weight", round(expected / observed, 3))

go deeper

for a junior

Be ready to say that reweighing changes how much each training row counts rather than adding or deleting rows, and that the weight depends on both the group and the label, not the group alone.

for a middle

Derive the weight as expected mass under independence divided by observed mass, say which cells go above one, and explain why the fitted model no longer needs the protected column at scoring time.

for a senior

Show the operational checks: cell counts before fitting, effective sample size and variance afterwards, evaluation on the unweighted held-out population, and the residual gap that survives the repair.

for a principal

Weigh reweighing against the alternatives as a default policy - it keeps every row, is auditable as a small table, and ships a group-blind model, which is often worth more to legal review than a marginally larger gap reduction.

## The idea in one line Instead of deleting rows, duplicating rows, or editing labels, reweighing leaves the training set exactly as it is and changes only how much each row *counts*, so that in the weighted data the protected group and the outcome look independent. ## Building the weights Start with the two-way table of training rows: protected group `S` down the side, label `Y` across the top. For each cell: - **Observed mass**: `P(S=s, Y=y)`, the fraction of training rows in that cell. - **Expected mass under independence**: `P(S=s) * P(Y=y)`, what the cell would hold if group told you nothing about the outcome. - **Weight**: `w(s, y) = P(S=s) * P(Y=y) / P(S=s, Y=y)`, equivalently expected count over observed count. Every row in that cell carries that weight. A cell that is thinner than independence predicts - usually the disadvantaged group's favourable outcomes - gets a weight above one; a cell that is fatter gets a weight below one. Weights are strictly positive and no row is discarded, which is the main attraction over dropping or duplicating rows. Worked shape: with 1000 rows split evenly between groups A and B and evenly between positive and negative labels, but with 400 of A's 500 rows positive and only 100 of B's 500, every cell expects 250 rows. The (A, positive) and (B, negative) cells hold 400 and get weight 0.625; the (A, negative) and (B, positive) cells hold 100 and get weight 2.5. ## What the model sees A learner that accepts per-row weights multiplies each row's contribution to the loss by its weight. Under the reweighed distribution, knowing the group tells the learner nothing about the label's marginal rate, so the fitted model has no incentive to reproduce the historical difference in selection rates. This is a repair aimed at independence between group and outcome - it does not by itself equalise error rates conditioned on the truth, and you should say so rather than claim it fixes every parity notion at once. ## Why practitioners like it - **No data is invented or thrown away.** Every original row survives with its original features and its original label. - **The protected attribute is a training-time input only.** The shipped model reads ordinary features; nothing at decision time needs to know the group. In deployments where reading the attribute at decision time is forbidden, this matters enormously. - **It is model-agnostic** - any learner that supports row weights can consume it, and the weights are a plain, auditable table you can show a reviewer. ## The costs and failure modes - **Effective sample size drops.** Concentrating mass on fewer rows raises the variance of everything downstream - the fitted parameters, and any metric estimated on weighted data. Roughly, the effective count behaves like `(sum of weights)^2 / sum of squared weights`, which is smaller than the row count whenever the weights are unequal. - **Extreme weights on thin cells.** If the disadvantaged group's favourable cell holds a handful of rows, its weight is large and those few rows dominate the fit. Check the cell counts before trusting the weights, and be prepared to say the data cannot support the repair. - **It repairs an association, not a mechanism.** Reweighing balances group against label. If the model can reconstruct the same distinction from other information in the features, part of the disparity comes back, and the residual gap is exactly what your held-out audit is for. - **Weights must not leak into evaluation.** Fit on weighted rows; measure the disparity and the accuracy on the unweighted held-out population, because that is the population the model will actually face. - **Multiple attributes multiply cells.** Two protected attributes and a binary label give eight cells, and the thinnest one governs how noisy the whole scheme is. ## How to talk about it in an interview Give the formula, say which direction the weights move, name the one-line advantage (group-blind at inference), and name the one-line cost (variance up, thin cells dangerous). Then say what you would check afterwards: the cell counts before fitting, and the gap plus the accuracy on held-out data after.

  • What does reweighing do to the variance of the fitted model?
    It raises it. Unequal weights concentrate the loss on fewer effective rows, so the effective sample size - roughly the squared sum of weights over the sum of squared weights - falls below the row count. The thinner the cell carrying the large weights, the more a handful of rows steers the fit, and the noisier both the parameters and any weighted estimate become. Check cell counts before trusting the scheme.
  • Your disadvantaged group's favourable-outcome cell holds twelve rows. What do you do?
    Do not ship the weights. Twelve rows given a large multiplier means the model is being steered by a dozen individuals, and the resulting disparity number has a confidence interval wide enough to cover almost anything. Options: collect more data for that cell, widen the group definition if that is defensible, move to a mitigation that does not depend on a thin cell, or report that the data cannot support a credible repair.
  • Should the weights also be used when evaluating the model?
    No. Fit on the weighted rows, but measure accuracy and the residual group gap on the unweighted held-out data, because that reflects the population the model will actually score. Evaluating on the weighted distribution measures performance on a population you invented for training purposes and will systematically flatter the fix.

saying these in an interview costs you the question

  • Says reweighing just upweights the smaller group
  • Ignores the label when forming the weights
  • Thinks it needs the protected column at scoring time
  • Claims it satisfies every fairness criterion at once
  • Never checks how many rows are in each cell
  • Evaluates the model on the weighted distribution

context