skip to content

How can a risk score be calibrated within every group yet still show unequal false-positive rates?

level: seniorimportance: should knowfreq 44%

answer

  1. the base rates are not equal
  2. one criterion has to break
  3. recall, predictive value and FPR are linked
  4. false-positive rate follows base-rate odds
  5. COMPAS: both sides were right

basics

~20 s

When two groups have different base rates, a score meaning the same thing in both must produce different error rates. Equal predictive value and equal false-positive rates cannot both hold, so the two criteria genuinely conflict.

solid answer

~50 s

Calibration within group means a score of 0.7 corresponds to a 70 percent chance of the outcome in every group, so the number means the same thing to whoever reads it. Error-rate balance instead asks that false-positive and false-negative rates match across groups. These pull apart whenever base rates differ, and the arithmetic forces it: for any group, `FPR = (p / (1 - p)) * ((1 - PPV) / PPV) * TPR`, where `p` is that group's base rate. Hold predictive value and recall equal across two groups and the group with the higher base rate must carry the higher false-positive rate. That is exactly the COMPAS dispute: ProPublica showed the recidivism score wrongly flagged non-reoffending black defendants at close to twice the rate of white defendants, while Northpointe replied that predictive value was near-equal. Both readings were correct on the same confusion matrices.

code

python · 20 lines
python
groups = {
    "A": {"tp": 300, "fn": 300, "fp": 100, "tn": 500},
    "B": {"tp": 150, "fn": 150, "fp": 50, "tn": 850},
}

for name, c in groups.items():
    pos = c["tp"] + c["fn"]
    neg = c["fp"] + c["tn"]
    base = pos / (pos + neg)
    tpr = c["tp"] / pos
    fpr = c["fp"] / neg
    ppv = c["tp"] / (c["tp"] + c["fp"])
    print(name,
          "base rate", round(base, 2),
          "recall", round(tpr, 2),
          "predictive value", round(ppv, 2),
          "false-positive rate", round(fpr, 3))

# A base rate 0.5  recall 0.5  predictive value 0.75  false-positive rate 0.167
# B base rate 0.25 recall 0.5  predictive value 0.75  false-positive rate 0.056

go deeper

for a junior

Know that fairness criteria can genuinely conflict, and that a score meaning the same thing in every group does not imply the groups suffer errors at the same rate.

for a middle

Be able to reproduce the arithmetic: with predictive value and recall held equal across groups, the group with the higher base rate must carry the higher false-positive rate.

for a senior

Expect to be handed two per-group confusion matrices and asked which criterion holds, which breaks, and whether the base-rate gap is real or an artefact of how the labels were collected.

for a principal

Own the decision the impossibility forces. State publicly which criterion the system holds, what is surrendered, and who reviews that call when the population or the label source shifts.

## The two claims that collide **Calibration within group** (sometimes called sufficiency) says: for every score value `s` and every group `a`, `P(Y = 1 | S = s, A = a) = s`. A 0.7 means a 70 percent chance of the outcome whether the person scored belongs to one group or the other. This is what makes a score usable by a human decision-maker: the number carries the same meaning regardless of who is being scored. **Predictive parity** is its thresholded cousin — equal positive predictive value (precision) across groups. **Error-rate balance** says the opposite conditioning: equal false-positive rate and equal false-negative rate across groups, which is equalized odds. Both sound like minimum decency. They cannot both be had. ## The arithmetic that forces the conflict For one group with base rate `p = P(Y = 1)`, the definitions alone give `PPV = p*TPR / (p*TPR + (1 - p)*FPR)` and rearranging that produces the identity behind the whole argument (Chouldechova, 2017): `FPR = (p / (1 - p)) * ((1 - PPV) / PPV) * TPR` Read it slowly. The false-positive rate is not a free dial. Once you fix a group's base rate, its recall and its predictive value, the false-positive rate is **determined**. So if you insist that two groups share the same `PPV` and the same `TPR`, and their base rates `p` differ, their false-positive rates must differ — and they differ in a predictable direction, since `p / (1 - p)` increases with the base rate. The group with more positives in it absorbs the higher false-positive rate. Kleinberg, Mullainathan and Raghavan (2016) proved the score-level version: calibration within group, balance for the positive class and balance for the negative class can hold simultaneously only in two degenerate cases — the base rates are equal, or the prediction is perfect. Outside those two corners, at least one criterion has to be surrendered. ## A worked example Take two groups of 1200 people each. Group A has 600 positives (base rate 0.50); group B has 300 (base rate 0.25). Now insist on equal recall of 0.50 and equal predictive value of 0.75 in both. - Group A: 300 true positives, 300 false negatives, 100 false positives, 500 true negatives. False-positive rate = 100 / 600 = **0.167**. - Group B: 150 true positives, 150 false negatives, 50 false positives, 850 true negatives. False-positive rate = 50 / 900 = **0.056**. Same recall, same predictive value, a three-fold gap in false-positive rate — the same three-fold ratio as `p / (1 - p)` between the groups. No bug, no biased feature, no bad training run produced this. Definitions and unequal base rates produced it. ## COMPAS: two correct answers The 2016 argument over the COMPAS recidivism score is the canonical instance. ProPublica compared defendants who did **not** go on to reoffend and found that black defendants in that set were labelled high-risk at close to twice the rate of white defendants — a false-positive gap, an equalized-odds violation. Northpointe replied that among people the tool labelled high-risk, the share who actually reoffended was close to equal across groups — predictive parity holding, roughly 63 versus 59 percent. Neither side made an arithmetic error. They audited different conditional probabilities on the same per-group confusion matrices, and the measured base rates in that data differed. The identity above says both findings *must* coexist. The disagreement was never about the numbers; it was about which criterion the public is owed. ## What to do with this in practice 1. **Stop looking for the model that satisfies everything.** More data, a better architecture and more careful feature work do not dissolve an identity. They shrink the uncertainty in your measured rates; they do not reconcile the criteria. 2. **Interrogate the base-rate gap before accepting it.** If the outcome was recorded through a process that attended to the groups unevenly, then `p` is measuring the process, not the world, and every criterion defined relative to `Y` inherits that flaw. This is the question that separates a strong senior answer from a textbook one. 3. **Choose, then publish per-group numbers.** Report recall, false-positive rate, predictive value and the base rate for each group, side by side, and state which criterion the system is built to hold. A single reassuring fairness number is a way of hiding the trade, not of resolving it. 4. **Expect the choice to be contested.** Because both readings are mathematically valid, the decision is a policy decision wearing a statistical costume. Owning that openly is more defensible than pretending the model settled it.

  • Does collecting more training data resolve the conflict?
    No. The impossibility is arithmetic, not statistical: it follows from the definitions of the rates plus a base-rate gap, so it survives any sample size. More data narrows the confidence around each measured rate, which is worth having, but it cannot make calibration within group and error-rate balance compatible while the base rates differ.
  • What if the base-rate gap is itself an artefact of how the labels were produced?
    Then the impossibility still binds on the labels you hold, but the criteria conditioned on the label lose their claim to describe reality. Investigate provenance first: an outcome recorded through uneven follow-up or uneven enforcement measures the process rather than the phenomenon, and any fairness argument built on it inherits that flaw.
  • What should the model's evaluation report actually contain, given the conflict?
    All of the competing numbers. Publish per-group recall, false-positive rate, predictive value and base rate together, and name the criterion the system is designed to hold. The reader then sees the trade instead of one flattering statistic, and a later reviewer can tell whether the trade is still the one the team intended.

Two honest thermometers can be perfectly accurate in two cities and still set off the frost alarm far more often in the colder one. The instruments are not biased; the climates differ.

saying these in an interview costs you the question

  • Says one side of the COMPAS dispute simply computed its metrics wrong
  • Claims an unbiased model can satisfy every fairness criterion at once
  • Thinks the conflict disappears with more data or a better model
  • Believes calibrating each group separately closes the error-rate gap
  • Quotes one fairness number without the per-group base rates

context