A fairness fix closes a 9-point between-group false positive rate gap but costs 1.5 AUC points. Do you ship it?
answer
- The two numbers are in different units
- Turn percentages into people and money
- One point is not a frontier
- Check the labels before paying for the fix
- Name who signs the tradeoff
basics
~20 sNot from those two numbers alone. Convert both sides into decision consequences and money, check the tradeoff curve for a cheaper knee, confirm the gap is real and not an artefact of biased labels, and put the call to a named accountable group rather than deciding it as an engineer.
solid answer
~60 sThe two numbers are in incomparable units, so first make them comparable. Nine points of false positive rate gap is a number of real people per period wrongly flagged on one side; 1.5 AUC points is worth whatever it is worth in the decision metric that actually pays - approvals at fixed review capacity, revenue per decision, losses avoided. Second, do not accept the frontier as given: sweep the penalty strength and plot residual gap against the business metric, because that curve usually bends and most of the gap often closes for a fraction of the quoted cost. Third, interrogate the inputs - is the 9-point gap larger than its uncertainty, and are the labels themselves clean, since a fix layered on biased labels launders the bias rather than removing it. Then hand the residual tradeoff, with those numbers attached, to the group that owns it: legal, risk, product and the affected users' representation. My job is to make the choice legible and to refuse to smuggle it into a hyperparameter.
go deeper
Be ready to say that a fairness fix normally costs some predictive performance and that the two quantities are measured differently, so a decision cannot rest on the two headline numbers alone.
Show that you would sweep the mitigation strength to see the whole gap-versus-accuracy curve, and re-express an AUC change at the actual operating point rather than quoting it as a ranking number.
Demonstrate the diagnostic instinct: check the gap against its uncertainty, question the labels, verify stability across splits, quantify both sides in people and money, and state the residual gap that remains after the fix.
Own the governance: publish the frontier rather than a point, refuse to bury the tradeoff in a single weight, place the call with a named accountable group including legal and product, and ship the decision with a written rationale and a reopening trigger.
## Why the question is not a modelling question Both quantities are real, both are measured on held-out data, and neither is denominated in anything a decision-maker can act on. A principal-level answer converts them, examines whether the tradeoff is actually as stated, and places the decision with whoever is accountable for it. ## Step 1: put both sides in the same units **The gap.** A 9-point difference in false positive rate is not an abstraction. Multiply it by the number of decisions the disadvantaged group receives per period and you get people - a count of individuals per week who are wrongly flagged and who would not have been had they belonged to the other group. Attach the consequence of one such error: a declined application, a withheld intervention, an escalation. That number, not the percentage, is what belongs in the decision memo. **The accuracy.** 1.5 AUC points is a change in ranking quality. Nobody is paid in AUC. Re-express it in the operating metric: at the review capacity you actually staff, how many fewer true positives are caught; at the current cutoff, how many more bad decisions per period; what that is worth in money or in lost outcomes. Sometimes 1.5 AUC points is a rounding error at the operating point; sometimes it is the whole margin. You cannot know which without recomputing at the operating point. ## Step 2: challenge the frontier A single pair of numbers describes one point on a curve. Sweep the strength of the mitigation - the penalty weight, the width of a band, the aggressiveness of the repair - and plot residual gap against the business metric. Fairness-accuracy frontiers are usually strongly bent: the first several points of gap are cheap and the last few are ruinous. If two thirds of the gap closes for 0.3 AUC points, the interesting decision is about the remaining third, and the original question was framed on the wrong point. Present the curve, not the point. ## Step 3: interrogate the inputs before paying for them - **Is the gap real?** Nine points measured on a slice with few rows carries an interval that may straddle much smaller values. Paying accuracy to chase a number that will not reproduce next quarter is a bad trade in both directions. - **Are the labels trustworthy?** If the recorded outcomes were themselves produced by a biased historical process, then the model is faithfully reproducing a bad target, and a mitigation bolted on downstream makes the metrics look better while the underlying harm survives. Fixing the label pipeline is slower, unglamorous and usually the higher-value intervention. - **Is the fix stable?** Re-run it on a different split or a different period. A mitigation whose effect swings between splits is not a fix, it is a fitted artefact. - **What is the residual?** No fix closes a gap to zero. State the number that remains, because that is what the organisation is actually accepting. ## Step 4: place the decision This is the part that separates a principal answer. The tradeoff is a values question with legal, reputational and commercial exposure, and it is not the modeller's to make alone - nor is it made by scalarising fairness and accuracy into one objective and letting an optimiser pick, which merely hides the choice in a coefficient nobody signed off on. What to bring: the curve, the human counts on both sides, the residual gap, the uncertainty, the alternatives you rejected and why, and a recommendation. Who decides: a named accountable group - legal, risk, product, and, where it can be arranged, some representation of the affected population. What ships with it: the decision written down with its rationale, the metric and cadence for monitoring the gap in production, and a trigger that reopens the decision if the gap or the population moves. ## The two failure modes The first is the engineer who ships the fix silently because it felt right, leaving an unowned business cost and an undocumented judgement. The second is the engineer who refuses to ship anything because the fix is imperfect and the accuracy cost is real, which preserves the full 9-point gap while looking rigorous. Both avoid the actual work, which is making an uncomfortable tradeoff explicit, quantified and owned.
- The frontier turns out to be flat - closing any of the gap costs proportionally. What changes?Then there is no cheap win to take and the decision is genuinely about values, so it moves entirely to the accountable group with my recommendation attached. I would also treat a flat frontier as a signal to look further upstream: flat usually means the disparity is baked into the features or the labels, and data collection, label review or a redesigned target may buy far more than any amount of post-hoc constraint strength.
- How do you monitor a shipped mitigation?Track the group gap and the business metric on the same dashboard and cadence, because a mitigation that decays shows up as the gap creeping back while accuracy quietly recovers. Include the volume and base rate per group, since drift in either invalidates the tuning. Define in advance the gap value that triggers a review, and name the owner who receives the alert - an unowned fairness dashboard is not monitoring.
- A stakeholder asks for a single fairness-accuracy score to optimise. How do you respond?I would push back. Scalarising the two into one objective does not remove the tradeoff, it buries it in a weight that nobody reviewed and that silently encodes how many wrongly flagged people one point of accuracy is worth. Better to publish the frontier, have the accountable group choose an operating point explicitly, and then, if convenient, express that choice as a coefficient - decision first, coefficient second.
saying these in an interview costs you the question
- Compares AUC points against error-rate points as if commensurable
- Treats a single measured tradeoff as the whole frontier
- Ships or refuses the fix without escalating the decision
- Never asks whether the labels encode the bias
- Hides the tradeoff inside one combined objective weight
- Reports the gap closed without stating the residual