How do you compute the four-fifths disparate-impact ratio for a hiring screen?
answer
- rate within group, not raw counts
- lowest rate over highest rate
- the 0.8 rule of thumb
- computed on decisions at a threshold
- a flag needing justification, not proof
basics
~20 sCompute each group's selection rate — the share of that group's applicants the screen passes — then divide the lowest by the highest. A ratio under 0.8 is the conventional flag for adverse impact and calls for justification.
solid answer
~50 sThe selection rate for a group is `passes / applicants` within that group. The disparate-impact ratio is the lowest group's selection rate divided by the highest group's. If a screen passes 50% of one group and 30% of another, the ratio is `0.30 / 0.50 = 0.6`, below the four-fifths (0.8) threshold used as a rule of thumb in US employment-selection guidance. Three things to say out loud. First, it is computed on *decisions*, so it moves whenever you move the threshold — a model unchanged in ranking can pass or fail depending on the cutoff. Second, it looks only at who gets selected, not at whether the selected people turned out qualified. Third, it flags rather than proves: a sub-0.8 ratio shifts the burden to justifying the screen, and a small applicant pool can push the ratio around on noise alone.
go deeper
Recall the arithmetic: selection rate is passes divided by applicants within a group, and the ratio is the lowest such rate over the highest. Remember 0.8 as the conventional line.
Explain what the ratio does and does not see. It reads decisions rather than errors, changes when the threshold changes, and says nothing about whether the selected candidates were qualified.
Demonstrate that you would compute it on realistic applicant flow at the shipped cutoff, report it with intervals and group counts, and recompute when the applicant mix drifts. Expect to explain how proxies produce a failing ratio with no attribute in training.
Own how the number is governed: who recomputes it, on what cadence, what threshold the business case justifies, and how you document a business-necessity argument instead of quietly cutoff-shopping until the ratio clears 0.8.
## The quantity For each group `g`, the **selection rate** is the fraction of that group's applicants who receive the favourable decision: ``` selection_rate(g) = selected(g) / applicants(g) ``` The **disparate-impact ratio** (also called the adverse-impact ratio) compares the least-selected group with the most-selected one: ``` DI = min_g selection_rate(g) / max_g selection_rate(g) ``` The **four-fifths rule** — a rule of thumb in the US federal Uniform Guidelines on Employee Selection Procedures — treats `DI < 0.8` as an indicator of adverse impact worth investigating. Worked example. A screen sees 400 applicants from group A and 120 from group B. It passes 200 of A (50%) and 36 of B (30%). `DI = 0.30 / 0.50 = 0.60`. That is well under 0.8, so the screen is flagged. ## What it is and is not **It is a rate of decisions, not a rate of errors.** The four-fifths check never looks at the label. It does not ask whether the passed candidates went on to succeed, or whether the screen's false-negative rate differs between groups. Two screens with identical accuracy profiles can have wildly different ratios if their thresholds sit at different points. **It is threshold-dependent.** Because it is computed on hard decisions, the same scoring model produces different ratios at different cutoffs. If the score distributions of two groups differ in shape, raising the bar can widen or narrow the ratio non-monotonically. Reporting a ratio without reporting the threshold it was computed at is incomplete. **It is a flag, not a verdict.** In its original legal setting a sub-0.8 ratio establishes a prima facie case; the employer may then argue the screen is job-related and consistent with business necessity, and a challenger may argue a less discriminatory alternative exists. A candidate who says "0.65, therefore illegal" has overclaimed. The engineering translation is: this number opens an investigation and a documentation obligation, it does not close one. **It is unstable on small pools.** The ratio is a quotient of two proportions. If one group contributed 40 applicants, its selection rate has a confidence interval tens of points wide, and so does the ratio. Report the ratio with an interval, or at least with the group counts, so a reader can see how much weight it bears. ## The connection to proxies The reason this check matters in a machine-learning setting is that **it can fail badly on a model whose training table never contained the protected attribute**. Selection rates are computed after the fact, from decisions the deployed model made, joined against attribute data held for audit. A hiring screen built on features like commute distance, university, or years of continuous employment can produce a 0.6 ratio purely through proxies — no attribute, no explicit rule, and yet unequal selection rates. That is also why the check has to be a *reporting* step rather than an assumption. Nothing in training tells you the ratio; you have to compute it on realistic applicant flow, at the threshold you actually ship, and recompute it whenever the threshold or the applicant mix changes. A screen that passed at launch can fail six months later because the composition of the applicant pool moved, with the model untouched. ## Practicalities that come up - **Which groups.** Compute the ratio for every group you have data for, not only the pair you expect to be a problem. Include an "unknown/declined to state" bucket in the counts so nobody can move the ratio by pushing people into it. - **Multi-stage funnels.** If the model is one gate among several, compute the ratio at your gate *and* end to end. A gate that looks acceptable in isolation can dominate the cumulative ratio when three such gates compose. - **Which group is the denominator.** The convention is the most-selected group, which makes the ratio at most 1. Some teams instead compare against the overall rate; either is defensible provided you say which you used. - **Rates, not counts.** A group can be a small share of hires and still have the highest selection rate. Comparing raw hire counts instead of rates is the most common arithmetic mistake on this question. ## What good sounds like Give the formula, give a two-line worked example, state the 0.8 threshold and where it comes from, then immediately qualify it: decision rates not error rates, threshold-dependent, a flag rather than proof, and shaky when a group's applicant count is small. Close on the proxy point — that a model with no protected attribute in it can still produce a failing ratio, which is precisely why the ratio has to be measured rather than reasoned about.
- Your ratio is 0.78 at the shipped threshold but 0.86 one point higher. What do you report?Both, with the threshold attached to each. The honest report is a curve of the ratio against the cutoff plus the cutoff you actually ship, because a single number invites cutoff-shopping. Then say which threshold the business case requires and what the ratio costs there, rather than quietly choosing the flattering one.
- One group has 40 applicants. How does that change how you present the ratio?Present it with an interval and the raw counts. With 40 applicants the group's selection rate carries an interval tens of points wide, so the ratio can swing across 0.8 on a handful of decisions. Say plainly that the estimate is underpowered and give the sample size needed before the number is actionable.
- Why can a model that never saw the protected attribute still fail this check?Because the ratio is computed on outcomes, not on inputs. Features like commute distance, institution, or employment continuity act as proxies, so the decisions distribute unequally by group even though no attribute was in the training table. The absence of the column is not a defence against the measurement.
saying these in an interview costs you the question
- Compares raw hire counts instead of within-group rates
- Treats a sub-0.8 ratio as automatic proof of discrimination
- Reports the ratio without saying at which threshold
- Claims the check cannot apply without the attribute in training
- Confuses selection rate with model accuracy for the group