skip to content

When is a different score threshold per protected group defensible in a deployed model?

level: seniorimportance: should knowfreq 48%

answer

  1. Acts directly on the measured gap
  2. The group is an input to the decision
  3. Ask who owns the legal call
  4. Is the attribute recorded and accurate
  5. Route to a human instead of deciding differently

basics

~20 s

Rarely, and never as an engineering decision alone. Per-group thresholds act directly on the measured gap but read the protected attribute at decision time, which is explicit differential treatment. Ship one only with counsel's sign-off, a recorded group attribute, and monitoring.

solid answer

~50 s

Per-group thresholds are the most direct lever available: accept a student-support model's cases above 0.52 for one group and above 0.61 for another, and the selection or error gap closes almost by construction, with the fitted model untouched. That directness is why they are dangerous. The rule reads the protected attribute at the moment of decision, so the decision differs by group by design - which in regulated domains such as credit, employment and housing is treated very differently from a group-blind model trained with fairness in mind. Three things must survive before it ships: a legal argument owned by counsel, not by me; an attribute that is actually recorded, accurate and complete at decision time, since self-reported or missing values silently break the calibration; and enough rows per group that the two numbers are not fitted to noise. If any fails, move the repair upstream, or use a reject-option band that routes borderline cases to a human instead of deciding differently.

go deeper

for a junior

Be ready to say what a per-group threshold is - one cutoff per protected group applied to the same model's scores - and that it needs to know the group at the moment of decision.

for a middle

Explain why it closes a measured gap so effectively, where the accuracy cost comes from when both groups must meet at a common operating point, and how a reject-option band differs from flipping the decision.

for a senior

Demonstrate deployment judgement: escalate the legal question rather than answering it, check that the attribute is present and accurate at decision time, put intervals on the cutoffs, define the missing-group branch, and schedule re-measurement.

for a principal

Own the position: whether your organisation permits decision-time use of protected attributes at all, what evidence and sign-offs a dial like this requires, and what the standing alternative is when it is refused.

## What the technique is The model is fitted once and left alone. It emits a score; the decision rule turns the score into an accept or reject by comparing it with a cutoff. A per-group threshold rule uses a different cutoff depending on the protected group the case belongs to - 0.52 for one group, 0.61 for another on a scholarship-award or student-support model. The reason it works so well is that the quantity being audited and the quantity being adjusted are the same quantity. If the gap you are repairing is stated in terms of selection rates or of error rates at the operating point, then moving the operating point per group can close it essentially exactly, and the optimal per-group cutoffs can be read off each group's error-rate curve directly. Group-specific cutoffs can be derived so that the groups meet at a common point on their respective curves; that shared point generally sits below where the stronger group alone could operate, which is where the accuracy cost comes from, and in some cases the exact match requires randomising decisions inside a band rather than a clean cutoff. ## Why it is nevertheless the hardest one to ship **1. It uses the protected attribute at decision time.** Pre- and in-processing consume the attribute during training and hand you a group-blind scorer. A per-group threshold cannot be group-blind - the group *is* an input to the decision. Two applicants with identical scores get different outcomes because of a protected characteristic. In domains covered by anti-discrimination law that is a materially different legal object from a model trained to be fair, and the analysis is jurisdiction- and domain-specific. The correct engineering behaviour is to bring the option to counsel with the numbers attached, not to decide it in a notebook and not to assert what is lawful. **2. The attribute has to exist, at decision time, and be right.** Protected attributes are frequently self-reported, optional, missing for a large share of records, or inferred. A rule that switches on a field which is blank a third of the time is not the rule you evaluated. Decide up front what happens to a case with no recorded group - almost always, apply the stricter or the default threshold and record that this happened - and measure how often it happens. **3. The two numbers are estimates.** Each cutoff is fitted on that group's held-out data. If the smaller group contributes a few hundred rows, the cutoff carries a wide interval, and a threshold tuned to the last decimal is tuned to sampling noise. Report the cutoffs with intervals, and prefer round, defensible numbers over precise ones. **4. It drifts.** Base rates and populations move. A pair of cutoffs that equalised something last quarter will not this quarter. Per-group thresholds are a standing commitment to re-measure and re-fit on a schedule, with an owner. ## The reject-option alternative Instead of deciding differently, decide *less*. Define a band around the decision boundary - the region where the model is least confident and therefore where its errors concentrate. Inside the band, route disadvantaged-group cases to a human reviewer rather than auto-rejecting them; outside the band, the model's decision stands for everyone. This keeps a human in the loop exactly where the model is weakest, and it is often easier to explain and to defend than a numeric cutoff difference, because the machine is not making a different decision - it is declining to make one. The cost is concrete and budgetable: band width times case volume equals extra reviews per week, at a known cost and turnaround per review. Widen the band and you close more of the gap and buy more review hours; that is the whole tuning problem, and it is a staffing conversation as much as a modelling one. ## What good sounds like in an interview Say that the technique is mechanically the most effective, name the decision-time use of the protected attribute as the reason it is the hardest to deploy, insist that the legal call belongs to counsel, raise attribute availability and quality, raise the noise in the cutoffs for small groups, and offer the two escape routes: move the repair upstream into training so the shipped rule is group-blind, or use a reject-option band that routes rather than decides. Candidates who present the dial as a quick fix, and candidates who dismiss it as illegal everywhere without qualification, both miss.

  • How would you set a reject-option band's width in practice?
    Treat it as a capacity problem. Band width times case volume gives the number of extra human reviews per period; price that against the review team's throughput and cost per case. Then sweep the width and plot residual gap against review hours - the curve usually bends, closing most of the gap for the first tranche of reviews. Pick the point the review team can actually staff every week, not the point that minimises the gap on paper.
  • The protected attribute is missing for 30% of applicants at decision time. What now?
    The rule you evaluated is not the rule you would run. Define the missing-group policy explicitly - normally apply the default or stricter cutoff - and evaluate the whole system including that branch, because the missing set is rarely a random sample. If a third of cases bypass the mitigation, the residual gap in production will be far larger than your offline number, which is an argument for moving the repair into training instead.
  • Why can't you just tune both cutoffs to whatever equalises the metric on your test set?
    Because both numbers are estimates from finite samples, and the smaller group supplies the fewest rows. A cutoff tuned to equalise a metric exactly on one held-out split will not equalise it on the next, and the shortfall is largest where you care most. Fit them with intervals, prefer round defensible values, sanity-check on a second split, and commit to re-fitting on a schedule as base rates drift.

saying these in an interview costs you the question

  • Presents per-group cutoffs as a quick uncontroversial fix
  • Asserts what is or is not lawful without counsel
  • Assumes the protected attribute is always available at decision time
  • Tunes two cutoffs on a small group without intervals
  • Ships the dial with no re-measurement schedule
  • Says a reject-option band is free because the model is unchanged

context