skip to content

Mitigating Group Disparity

After an audit shows a gap you can reweight the training rows, constrain the objective, or move one group's threshold. Interviewers ask which stage to repair at and what accuracy it costs.

on this pageshow

questions

4

Where can you intervene to shrink a model's between-group disparity: before, during, or after fitting?

level: middleimportance: must knowfreq 58%

answer

  1. Three places in one pipeline
  2. Before the fit: change the data's mass
  3. During the fit: a term in the objective
  4. After the fit: change the decision rule
  5. Who still needs the group at scoring time

basics

~20 s

Three points. Pre-processing reweights or repairs the training data, in-processing adds a fairness penalty or constraint to the training objective, and post-processing changes the decision rule per group after fitting. Each needs different access and pays a different accuracy cost.

solid answer

~50 s

Mitigations sort into three families by where they touch the pipeline. **Pre-processing** changes the data before fitting: reweighing rows so each (group, outcome) cell carries the mass it would have if group and outcome were independent, for example in a hospital no-show intervention model, so no protected column is needed at inference. **In-processing** changes the fit itself: an equalized-odds penalty added to the training objective of a job-ad delivery model, or an adversarial head trained to predict a customer's dialect from the model's representation while the model is updated to make that head fail. **Post-processing** leaves the model alone and changes decisions: per-group score thresholds, or a reject-option band where borderline disadvantaged cases are flipped or sent to a human. The practical selector is access and permission: post-processing needs the protected attribute at decision time, the other two need it only at training time.

go deeper

for a junior

Be ready to name the three families - before, during, and after fitting - and give one concrete example of each. Knowing that reweighing happens before training and thresholding after is most of the credit here.

for a middle

Explain the mechanics: how a penalty term enters the objective, what an adversarial head is optimising against, and why pre- and in-processing leave you with a group-blind scorer while post-processing does not.

for a senior

Show that you pick by constraint, not by taste - what you control, whether the protected attribute is present and permitted at decision time, whether the labels are trustworthy - and that you re-measure and monitor the residual gap after shipping.

for a principal

Own the policy: which family your organisation defaults to, what evidence is required before a mitigation ships, and how the accuracy cost and residual gap get documented and reviewed as populations drift.

## The problem being solved A disparity here means the model's behaviour differs between groups defined by a protected attribute such as sex, race, age or disability status - one group sees a higher false positive rate, a lower selection rate, or a worse recall than another. Suppose the audit is done and the gap is real and material. The question is what you can actually turn. Every mitigation touches the pipeline in exactly one of three places, and the taxonomy is worth memorising because it maps one-to-one onto what access you have and what you are allowed to do. ## Pre-processing: change the data You modify the training set before any model sees it, so the disparity is smaller in the signal the learner is given. - **Reweighing** attaches a weight to each training row so that, in the weighted data, group membership and the label are statistically independent. Rows in under-represented (group, favourable outcome) cells get weights above one; over-represented cells get weights below one. - **Relabelling** or *massaging* flips the labels of a small number of borderline rows near the decision boundary to balance the outcome rates. It is more invasive - you are editing ground truth - and needs a strong argument that the original labels were themselves biased. - **Representation repair** transforms the features so the protected attribute becomes hard to recover from them. The defining property: the protected attribute is consumed at training time and never again. The shipped scorer is group-blind. That is a large practical advantage. ## In-processing: change the fit You modify the objective the learner optimises. - **A penalty term.** Add a differentiable surrogate of the disparity to the loss: `loss = task_loss + lambda * disparity`. Turning up `lambda` buys a smaller gap and gives back accuracy. On a job-ad delivery model, an equalized-odds penalty that closed a 9-point false positive rate gap cost about 1.5 AUC points - a real and typical order of magnitude. - **A hard constraint.** Solve the same task subject to the disparity staying under a bound, rather than trading it off softly. - **Adversarial debiasing.** Attach a second head - an adversary - that tries to predict the protected attribute from the main model's representation or output. The adversary is trained to succeed; the main model is trained to do its own job *and* to make the adversary fail. On a customer-service escalation model an adversary might try to read the speaker's dialect off the representation. At convergence the representation carries little group information. The catch is that min-max training is unstable, and the guarantee only covers what that particular adversary is capable of detecting. Again the protected attribute is only needed while training. ## Post-processing: change the decision The fitted model is untouched; you change how its scores become decisions. - **Per-group thresholds.** Accept above 0.52 for one group and above 0.61 for another. Mechanically the most direct lever on a measured gap, because it acts on exactly the quantity being measured. - **Reject-option rules.** Define a band of scores near the boundary where the model is least confident, and inside that band favour the disadvantaged group or route the case to a human reviewer. The defining property is the mirror image of the other two: these rules read the protected attribute *at decision time*, which many deployments cannot do and many regulators treat very differently from using it during training. ## Choosing between them Ask, in order: (1) Do I control training, or only a vendor's scores? If only scores, post-processing is your only option. (2) Will the protected attribute be present, accurate and permitted at decision time? If not, post-processing is out. (3) Do I believe the labels? If the labels themselves encode the historical bias, a mitigation applied downstream of them is polishing a corrupted target, and data work comes first. ## The accuracy cost A constraint shrinks the feasible set, so the best training accuracy achievable under a fairness constraint can only match or fall below the unconstrained optimum - a fairness fix that appears free in-sample usually means the constraint was not binding. Out-of-sample, a constrained model occasionally scores *better*; read that as evidence the disparity came from fitting noise in a thin slice, and go check, rather than as a free lunch. Whichever family you pick, the fix is not done until the gap is re-measured on held-out data, the residual gap is written down, and monitoring is in place - mitigations decay as base rates and populations drift.

  • Your scoring service is not permitted to read the protected attribute at decision time. Which family survives?
    Pre-processing and in-processing. Reweighing and fairness penalties consume the protected attribute only while training, so the deployed scorer is group-blind and the mitigation is baked into the weights. Per-group thresholds and reject-option rules need the attribute at the moment of decision, so they are unavailable - and in regulated domains that restriction is often the point, not an accident. If you are handed only a vendor's scores and cannot use the attribute at decision time, you have no mitigation lever at all and must escalate.
  • How does adversarial debiasing work, and why is it the least popular of these in practice?
    An adversary head is trained to predict the protected attribute from the main model's representation, while the main model is updated to do its task and simultaneously raise the adversary's loss. At convergence the representation is close to group-uninformative. It is unpopular because min-max training is unstable and hard to tune, and because the guarantee is only as strong as the adversary: information a weak adversary misses is still in the representation for a downstream model to exploit.
  • Can a fairness constraint ever leave accuracy unchanged?
    Yes, when the constraint is not binding - the unconstrained optimum already satisfies it, so nothing moves. That is common when the measured gap was small or came from a handful of rows. If you add a constraint you believe is binding and see no accuracy change at all, suspect the penalty weight is too small, the surrogate is not tracking the real disparity, or the gap you were chasing was noise.

Fixing a badly balanced weighing scale: recalibrate the weights you put on it, rebuild the mechanism, or correct every reading after the fact. All three work; only the last needs the operator to know which item is on the pan.

saying these in an interview costs you the question

  • Assumes one mitigation is usable at every pipeline stage
  • Treats per-group thresholds as always deployable
  • Claims a fairness fix costs no accuracy at all
  • Never re-measures the gap on held-out data
  • Applies a fix on top of labels known to be biased
  • Uses the protected attribute at scoring time without checking it is permitted

context

open as a page

When is a different score threshold per protected group defensible in a deployed model?

level: seniorimportance: should knowfreq 48%

basics

~20 s

Rarely, and never as an engineering decision alone. Per-group thresholds act directly on the measured gap but read the protected attribute at decision time, which is explicit differential treatment. Ship one only with counsel's sign-off, a recorded group attribute, and monitoring.

open as a page

A fairness fix closes a 9-point between-group false positive rate gap but costs 1.5 AUC points. Do you ship it?

level: principalimportance: should knowfreq 34%

basics

~20 s

Not from those two numbers alone. Convert both sides into decision consequences and money, check the tradeoff curve for a cheaper knee, confirm the gap is real and not an artefact of biased labels, and put the call to a named accountable group rather than deciding it as an engineer.

open as a page

How does reweighing training rows shrink a group disparity before any model is fitted?

level: middleimportance: nice to knowfreq 30%

basics

~20 s

Each training row gets a weight equal to the count its (group, outcome) cell would have if group and label were independent, divided by the count actually observed. The weighted data carries no group-label association, and scoring never needs the protected column.

open as a page