A protected-group slice of 38 applicants shows a 12-point pass-rate gap. Is the gap real?
answer
- 38 people buys very little precision
- interval on the gap, not the rate
- Wilson beats the normal approximation here
- not significant is not parity
- twenty slices, one false flag expected
basics
~20 sNot established. With 38 people, a 95% interval on that group's pass rate spans roughly 30 percentage points, and the interval on the 12-point gap itself runs from about minus 4 to plus 29 points. Report intervals, not point estimates.
solid answer
~50 sTake 15 of 38 passing (39.5%) against 208 of 402 elsewhere (51.7%) — a 12.3-point gap. A Wilson 95% interval for the small slice runs roughly 0.26 to 0.55, and the interval on the *difference* runs about -0.04 to +0.29, so zero is inside it. The gap is not distinguishable from sampling noise. Two mistakes to avoid in opposite directions. Acting on it — rebuilding features or reweighing — means chasing noise and burning credibility when the number moves next month. Declaring parity is equally wrong: an underpowered test that fails to detect a gap has not shown there isn't one, and a real 12-point gap would go undetected here most of the time. The right answers are to report the interval alongside the point estimate, state the minimum gap this sample could detect, prespecify which slices you monitor so you are not fishing across twenty of them, and pool more data before deciding.
code
python · 21 linesimport math
def wilson(successes, n, z=1.96):
p = successes / n
denom = 1 + z * z / n
centre = (p + z * z / (2 * n)) / denom
half = z * math.sqrt(p * (1 - p) / n + z * z / (4 * n * n)) / denom
return centre - half, centre + half
small = (15, 38) # protected-group slice: 15 passes of 38
rest = (208, 402) # everyone else in the funnel
for label, (k, n) in (("slice", small), ("rest", rest)):
lo, hi = wilson(k, n)
print(label, round(k / n, 3), "95% CI", round(lo, 3), round(hi, 3))
p1, n1 = small[0] / small[1], small[1]
p2, n2 = rest[0] / rest[1], rest[1]
se = math.sqrt(p1 * (1 - p1) / n1 + p2 * (1 - p2) / n2)
gap = p2 - p1
print("gap", round(gap, 3), "95% CI", round(gap - 1.96 * se, 3), round(gap + 1.96 * se, 3))go deeper
Recall that a percentage from a few dozen people carries a wide margin of error, so a gap that size can appear with no underlying difference. Ask for the counts before reacting to any rate.
Explain how to put an interval on a proportion and, more importantly, on the difference between two proportions, and why the Wilson form is preferred to the plain normal approximation at small counts.
Show the reporting judgment: intervals beside every rate, a stated minimum detectable effect, prespecified slices to avoid fishing, and a pooling plan. Be equally firm that non-significance on 38 people does not license a parity claim.
Own the policy: which slices are monitored, what evidence threshold triggers action, how multiplicity is handled across the whole dashboard, and how you keep small groups visible in reporting rather than dropped for being unmeasurable.
## The trap Slice the funnel finely enough and some protected-group slice will always look alarming. The question is whether a 12-point gap on 38 people is a finding or an artefact — and the answer, almost always, is that on that sample size you cannot tell. ## The arithmetic Suppose the slice has 38 applicants of whom 15 pass (39.5%), against 402 applicants elsewhere of whom 208 pass (51.7%). The observed gap is 12.3 points. A **Wilson score interval** is the standard interval for a proportion at small `n` — better behaved than the plain normal approximation, which misbehaves when `n` is small or `p` is near 0 or 1. For `p̂` successes-proportion out of `n`, with `z = 1.96` for 95%: ``` centre = (p + z^2 / (2n)) / (1 + z^2 / n) half = z * sqrt(p*(1-p)/n + z^2/(4n^2)) / (1 + z^2 / n) ``` On the small slice this gives roughly `[0.26, 0.55]` — an interval about 30 points wide. Practically, the true pass rate for that group is consistent with anything from one-in-four to a bit over half. Overlapping intervals are suggestive but not the correct test; the honest quantity is an interval on the **gap itself**: ``` se = sqrt(p1*(1-p1)/n1 + p2*(1-p2)/n2) gap = p2 - p1 -> gap +/- 1.96 * se ``` Here `se ≈ 0.083`, so the 95% interval on the 12.3-point gap runs from about **-4 to +29 points**. Zero is inside it. Nothing has been established. ## Both wrong conclusions **"The gap is real, escalate."** Rebuilding features or changing the model on a result this noisy is chasing sampling variation. Next month's 38 people will produce a different number, the change will look like it worked or backfired for no reason, and the team learns to distrust fairness reporting. **"Not significant, so the model is fair for that group."** Equally wrong, and more common among people who have just learned the first lesson. A non-significant result on an underpowered slice is an absence of evidence. Compute the **minimum detectable effect**: at `n = 38` against a large comparison group, roughly a 16-to-17-point gap is needed before a 95% interval clears zero. A genuine 10-point disparity would slip through most of the time. Saying "we could not detect a gap smaller than about 17 points" is honest; saying "there is no gap" is not. ## Multiplicity Slices multiply. Cut the funnel by four attributes with three or four levels each, cross a couple of them, and you have twenty or more comparisons. At a 5% level, one in twenty flags on chance alone, so a single significant slice out of twenty is exactly what you expect from a model with no disparity at all. Options: prespecify a small set of slices you always report; apply a multiplicity correction such as controlling the false discovery rate; or treat all slice findings as hypothesis-generating and require a fresh sample before acting. What you cannot do is scan twenty slices, report the worst one, and present its nominal p-value as if it were the only test you ran. ## What to actually do with a small slice - **Report the interval next to every rate.** A dashboard of bare point estimates on small slices manufactures alarm and false comfort in equal measure. - **State power up front.** "This slice can detect gaps of about 17 points or more" tells the reader what the number is worth before they read it. - **Pool.** Accumulate over a longer window, or aggregate to a coarser grouping, until the interval is narrow enough to act on. Say how long that will take. - **Watch the direction across periods.** A gap of the same sign in six consecutive underpowered windows is meaningful evidence even though no single window is significant; that is a different and much stronger argument than one flagged slice. - **Do not delete the slice from the report.** A group too small to measure is itself a finding about your data collection, and it is precisely the group least likely to be noticed if things go wrong. ## Talking to non-statisticians The sentence that lands is: "with this many people, the number would bounce by that much even if the model treated everyone identically." Then give the interval and the number of applicants you would need before the number means something. Handing a stakeholder a raw 12-point gap on 38 people with no interval is the failure the interviewer is probing for — not the statistics, the reporting judgment.
- Leadership wants a yes-or-no answer today. What do you tell them?That the honest answer is "we cannot tell yet", with the interval attached: the gap is somewhere between roughly -4 and +29 points. Then give a decision plan — the sample size or time window needed to resolve it, what you will monitor meanwhile, and the gap size that would trigger action. A fabricated yes or no is the only genuinely wrong answer.
- You cut the funnel twenty ways and exactly one slice is significant. How do you read that?As roughly what chance alone produces at a 5% level. Treat it as hypothesis-generating: prespecify that slice, gather a fresh sample, and see whether it reappears. Reporting the worst of twenty scans with its nominal p-value overstates the evidence considerably.
- The gap is small and non-significant every quarter, but always in the same direction. Does that matter?Yes, considerably. Consistent sign across independent periods is far stronger evidence than any single underpowered window, and it can be combined formally by pooling the counts or meta-analysing the estimates. Repeated non-significance is not repeated evidence of parity.
A 12-point lead in a poll of 38 people is not a lead; it is a margin of error wearing a headline.
saying these in an interview costs you the question
- Reports the 12-point gap with no interval attached
- Reads non-significance as proof the groups are treated equally
- Escalates a model rebuild off one small noisy slice
- Scans many slices and reports the worst nominal p-value
- Drops the small group from reporting because it is too small