What is the prosecutor's fallacy when a DNA database cold hit has a one-in-a-million match probability?
answer
- which way round is the conditional
- P(match | innocent) is not P(innocent | match)
- half a million profiles were searched
- expected coincidental matches: about 0.5
- the prior comes from the case
basics
~10 sIt is swapping P(match given innocent) for P(innocent given match). A one-in-a-million random-match probability is not a one-in-a-million chance of innocence: searching a 500,000-profile database expects about half an innocent match by chance alone.
solid answer
~50 sThe prosecutor's fallacy is the transposed conditional. The forensic statistic is a random match probability: `P(profile matches | this person is not the source) = 1/1,000,000`. Restating that as a one-in-a-million chance the defendant is innocent silently assumes the prior odds were 1:1, which they never are. Do it in odds instead. Suppose the true source is somewhere in a searched 500,000-profile database and every entry is equally suspect beforehand: prior odds are 1:499,999, the likelihood ratio for a match is 1,000,000, so posterior odds are about 2:1 -- roughly two-thirds, strong but nowhere near a million to one. The cold-hit structure is what opens the gap: half a million comparisons at one in a million each expect 0.5 innocent matches. Report the conditional in the direction the lab measured it, and let the prior come from the rest of the case.
go deeper
Recognise the phrase 'a one-in-a-million chance he is innocent' as a reversed conditional, and be able to say which probability the laboratory actually measured.
Show the mechanics: write both conditionals, apply the odds form with an explicit prior, and explain why the number of database comparisons matters to the answer.
Demonstrate judgment about how the evidence was produced -- cold hit versus confirmatory test, error rates that bound the likelihood ratio, and how you would phrase the statistic so a non-technical audience cannot transpose it.
Own the reporting standard. Decide that your organisation publishes likelihood ratios with the prior stated separately, and be ready to defend that convention against pressure for a single confident-sounding number.
## The two conditionals Every version of this error is the same move: reading `P(E | H)` as `P(H | E)`. In the forensic setting, - `P(match | innocent)` is the **random match probability** (RMP). It is what the laboratory can estimate from allele frequencies: how often a profile drawn from the population would coincidentally match the crime-scene sample. Say 1 in 1,000,000. - `P(innocent | match)` is what a court cares about. It cannot be computed from the RMP alone, because it depends on how likely the person was to be the source *before* the DNA evidence arrived. The prosecutor's fallacy asserts that the second equals the first. It is the same structural mistake as reading a test's sensitivity as the probability of disease given a positive: the direction of conditioning has been reversed and the base rate has been dropped. ## Why the gap is enormous in a cold hit A **cold hit** means no suspect existed beforehand; a database of profiles was searched and one matched. Suppose the database holds 500,000 profiles. The expected number of coincidental matches in that search is `500,000 x (1/1,000,000) = 0.5`. In other words, if the true source were not in the database at all, you would still expect roughly a coin-flip's worth of innocent matches per search. Finding one match is therefore not astonishing on its own. Run the odds form of Bayes' rule under a deliberately simplified assumption -- that the source is certainly one of the 500,000 and all are equally likely a priori: ``` prior odds (this person is the source) = 1 : 499,999 likelihood ratio for a match = 1 / (1/1,000,000) = 1,000,000 posterior odds = 1,000,000 : 499,999 ~ 2 : 1 posterior probability ~ 2/3 ``` About two-thirds. That is meaningful evidence and it is nothing like the '99.9999% certain' that the fallacy suggests. Change the assumptions -- a larger suspect population, real doubt about whether the source is in the database at all -- and the number moves again, which is precisely the point: the RMP alone does not determine it. ## The contrast with a confirmatory test The same RMP means something very different when a suspect was already identified by independent evidence -- opportunity, motive, a witness -- that narrows the plausible source population to, say, 20 people. Now the prior odds are 1:19, the likelihood ratio is still 1,000,000, and the posterior odds are about 52,000:1. Identical laboratory statistic, radically different conclusion, because the prior differs. This is the operational lesson: a match found by searching a large database is weaker evidence than the same match on a pre-identified suspect, and the difference lives entirely in the prior. ## The mirror-image error The **defence attorney's fallacy** runs the other way: in a country of 300,000,000 people, an RMP of 1 in 1,000,000 implies roughly 300 people would match, so the evidence supposedly proves nothing. This misuses the base rate in the opposite direction, by pretending the prior suspect pool is the entire population when other case evidence has already excluded almost all of it. Both fallacies are failures to handle the prior honestly; one sets it implicitly at 1:1, the other at 1:300,000,000. ## How the statistic should be stated The defensible statement keeps the conditioning where the measurement was made: *if this person were not the source, the chance their profile would match is about one in a million*. Better still, report the likelihood ratio -- 'this evidence is about a million times more likely if the defendant is the source than if he is not' -- and leave the prior explicitly to the trier of fact, who has the rest of the case. Any sentence of the form 'there is a one-in-a-million chance he is innocent' has crossed into the fallacy. Two further honesty requirements: the RMP is a population-genetics estimate that can be badly off for close relatives or for a subpopulation with different allele frequencies, and it excludes non-coincidental sources of a match such as sample mix-up, contamination or transfer. Those laboratory error rates are frequently larger than the RMP itself, which means the realistic likelihood ratio is capped well below 1,000,000 no matter what the allele arithmetic says. ## How to answer in the room Name the error as a transposed conditional, write both conditionals side by side, run the odds calculation to show the size of the gap, and then make the structural point: a cold database search multiplies the number of comparisons, which raises the chance of a coincidental hit and lowers the prior for any individual matched entry. Finish by stating the statistic correctly.
- What is the defence attorney's fallacy, and why is it also wrong?It argues that in a population of 300 million an RMP of one in a million implies about 300 matching people, so the evidence is worthless. That treats the entire population as the suspect pool, ignoring that opportunity, geography and other case evidence have already excluded almost all of it. Both fallacies are dishonest handling of the prior, in opposite directions.
- Why is the same match weaker evidence from a cold database hit than from a pre-identified suspect?The laboratory statistic is identical; the prior is not. A cold hit runs hundreds of thousands of comparisons, so each individual entry starts with tiny prior odds of being the source and coincidental matches are expected. A suspect already narrowed to a handful of people by independent evidence starts with prior odds thousands of times better, so the same likelihood ratio lands far higher.
- Why does laboratory error rate cap how strong a DNA match can be?The random match probability only covers coincidental profile agreement between different people. A match can also arise from sample swap, contamination or transfer, and those rates are typically far larger than one in a million. Since the overall probability of a match given innocence is at least the error rate, the realistic likelihood ratio is bounded by that error rate, not by the allele arithmetic.
Almost every lottery winner bought a ticket, but buying a ticket does not make you a probable winner. The prosecutor's fallacy reads the first sentence as the second.
saying these in an interview costs you the question
- Says a one-in-a-million RMP means the defendant is a million-to-one guilty
- Treats the random match probability as a probability of innocence
- Ignores how many profiles the database search compared
- Assumes prior odds of guilt are 1:1 without saying so
- Overlooks laboratory error as a source of matches