skip to content

How do demographic parity and equalized odds differ as group fairness criteria?

level: middleimportance: must knowfreq 60%

answer

  1. one ignores the true label
  2. selection rate versus error rate
  3. equal recall and equal false-positive rate
  4. equal opportunity constrains recall only
  5. a perfect model can break parity

basics

~20 s

Demographic parity requires the same positive-prediction rate in every group, ignoring the true outcome. Equalized odds requires equal true-positive and false-positive rates in every group, comparing errors only among people who share a true label.

solid answer

~50 s

Both compare a classifier across groups defined by a protected attribute, but they condition on different things. Demographic parity, also called statistical parity, asks that the rate of positive predictions be equal in every group: `P(prediction = 1 | group)` is the same everywhere. It never mentions the true label, so it can be satisfied by a model that picks the right people in one group and near-random people in another. Equalized odds conditions on the truth: equal true-positive rates and equal false-positive rates across groups, so among people who genuinely qualify the model finds the same share in each group, and among those who do not it wrongly flags the same share. Equal opportunity is the relaxation that constrains the true-positive rate only. Which you want follows from the harm: unequal access, or unequal error.

go deeper

for a junior

Be ready to state each criterion in one sentence and to say which one looks at the true label. Knowing that demographic parity counts selections while equalized odds counts errors is the bar here.

for a middle

Expect to write both criteria as conditional probabilities, show why a perfect classifier can violate demographic parity when base rates differ, and explain exactly which half equal opportunity drops.

for a senior

Show that you pick the criterion from the harm. Name who bears a false positive and who bears a miss in the system under discussion, then defend the choice out loud rather than reciting definitions.

for a principal

Own the framing that no single criterion is fairness. Decide which one the organisation commits to, what that costs on the others, and how the trade is explained to people outside the team.

## What these criteria are actually comparing A group fairness criterion is a statistical statement about a classifier's behaviour on two or more groups defined by a protected attribute `A` (a gender, an age band, a region). Write `Y` for the true label, the outcome you care about, such as *the loan was repaid* or *the seller was fraudulent*; write `Yhat` for the model's binary decision after thresholding. Every criterion below says that some conditional probability is equal in every group. They differ in **what is conditioned on**, and that single choice is what makes them incompatible with each other. ## Demographic parity (the independence family) Demographic parity requires `P(Yhat = 1 | A = a)` to be the same for every group `a`. In words: the model selects the same share of each group. It is the only one of these criteria that never mentions `Y`, which means it can be audited from outside without access to outcomes — one big reason regulators, journalists and product managers reach for it first. Two consequences follow immediately. **It fights accuracy whenever base rates differ.** If `P(Y = 1 | A = a)` genuinely differs between groups, then even a classifier that predicts every label perfectly violates demographic parity, because a perfect model's selection rate in each group *equals* that group's base rate. There is no clever model that escapes this: the tension is definitional, not a modelling failure. **It constrains the count, not the choice.** Parity fixes how many people are selected in each group, not who. A model can satisfy it exactly by ranking one group carefully and choosing nearly at random in the other. The selection rates match; the second group is still being served a worse product. This is the standard objection to using demographic parity alone, and interviewers like to hear it stated. The gap is usually reported either as a difference of selection rates or as their ratio. ## Equalized odds (the separation family) Equalized odds requires `P(Yhat = 1 | Y = y, A = a)` to be the same across groups, for **both** `y = 1` and `y = 0`. That is two equalities at once: equal **true-positive rate** (recall) across groups, and equal **false-positive rate** across groups. The comparison now happens *within* a true label. Among people who genuinely qualify, the model finds the same share of them in each group. Among people who do not qualify, it wrongly flags the same share. Because it conditions on `Y`, unequal base rates do not by themselves put equalized odds in conflict with accuracy: a perfect classifier has a true-positive rate of 1 and a false-positive rate of 0 in every group, so it satisfies equalized odds trivially. The price is a different one — equalized odds takes the recorded label as ground truth. If `Y` was measured through an uneven process, the criterion faithfully preserves that unevenness while looking rigorous. ## Equal opportunity Equal opportunity is equalized odds relaxed to the positive label only: equal true-positive rate across groups, with the false-positive rate left unconstrained. Reach for it when the harm lives in the **missed positive**. A diabetic-retinopathy screening model that reaches recall 0.91 on one clinic's patient population and 0.74 on another is failing equal opportunity, and the failure is concrete: roughly a quarter of the sight-threatening cases at the second clinic never get referred, against under a tenth at the first. The asymmetry justifies dropping the other half — a false alarm here costs one unnecessary follow-up examination, while a miss costs vision. Flip the setting and the logic flips with it. Where the false positive *is* the injury — an account frozen, a person detained, a seller's livelihood stalled — the false-positive half of equalized odds is precisely the half you must not drop, and equal opportunity is the wrong criterion. ## Choosing between them The honest procedure is not to pick the criterion with the nicest name. It is: 1. **Name the decision and both error types.** Who bears a false positive, and how badly? Who bears a miss? 2. **Ask whether `Y` deserves to be conditioned on.** If the label came from a process that treated the groups differently, criteria defined relative to `Y` inherit that. Criteria that ignore `Y`, like demographic parity, become comparatively more defensible. 3. **Accept that satisfying one leaves the others free to drift**, and say which one you committed to. ## How you check them Split the evaluation set by group, build one confusion matrix per group, and read the rates off each: selection rate for demographic parity, true-positive and false-positive rate for equalized odds. Report them side by side with the base rates, never as a single aggregated fairness number — the aggregate is where the story disappears. A third family exists that conditions on the *prediction* rather than the truth — predictive parity and calibration within group — and it is the one that collides hardest with equalized odds when base rates differ.

  • What does equal opportunity relax compared with equalized odds, and when is that relaxation acceptable?
    Equal opportunity keeps only the true-positive-rate equality and drops the false-positive-rate equality. That is acceptable when a missed positive is the dominant harm and a false positive is cheap, as in a screening model whose worst false alarm is one unnecessary follow-up test. Where the false positive is itself the injury, dropping that half discards the constraint that mattered.
  • Can a model satisfy demographic parity and still treat one group badly?
    Yes. Parity fixes how many are selected, not who. A model can rank one group well and pick nearly at random in the other, hitting identical selection rates while the second group's selected members are far less likely to succeed. Parity is the start of an audit, not proof of fairness.
  • If base rates genuinely differ between groups, what does demographic parity cost?
    Accuracy, directly. An accurate model's selection rate tracks the base rate, so forcing equal rates means either rejecting qualified people from the higher-rate group or accepting unqualified ones from the lower-rate group. That can still be the right policy, but the cost is real and should be stated rather than buried.

Demographic parity checks that the same fraction of each group gets through the door. Equalized odds checks that, among people who should have gotten through, the same fraction did.

saying these in an interview costs you the question

  • Says equal accuracy within each group is demographic parity
  • Claims a perfectly accurate model always satisfies demographic parity
  • Treats equalized odds and equal opportunity as the same constraint
  • Assumes satisfying one criterion makes the model fair overall
  • Reports a single aggregate fairness score instead of per-group rates

context