Why can a 6% word error rate still be unusable for a clinical dictation product?
answer
- one average over very unequal words
- rare terms barely move the aggregate
- a dropped 'no' costs one deletion
- normalization changes the number reported
- slice the metric, don't average it
basics
~20 sA headline word error rate averages over all words, so it is dominated by common speech. Drug names, dosages and negations are rare in the corpus but carry nearly all the risk, and errors there can run several times the headline rate while barely moving it.
solid answer
~50 sWord error rate is edit distance over the whole transcript divided by reference word count, which means every word counts the same. In dictation, filler words and function words vastly outnumber the terms that matter, so a 6% overall WER can hide a 30% error rate on drug names and dosages — and the aggregate barely moves either way. WER is also blind to *meaning*: dropping a single "no" from "no history of chest pain" is one deletion out of six words, roughly 17% locally and invisible in the average, yet it inverts the clinical claim. The fix is to stop reporting one number. Slice WER by term class, add entity-level precision and recall over the vocabulary you actually care about, add a targeted negation check, and gate releases on those slices rather than on the headline.
code
python · 16 linesdef wer(ref: str, hyp: str) -> float:
r, h = ref.split(), hyp.split()
d = [[0] * (len(h) + 1) for _ in range(len(r) + 1)]
for i in range(len(r) + 1):
d[i][0] = i
for j in range(len(h) + 1):
d[0][j] = j
for i in range(1, len(r) + 1):
for j in range(1, len(h) + 1):
cost = 0 if r[i - 1] == h[j - 1] else 1
d[i][j] = min(d[i - 1][j] + 1, d[i][j - 1] + 1, d[i - 1][j - 1] + cost)
return d[len(r)][len(h)] / len(r)
print(round(wer("patient reports no chest pain", "patient reports chest pain"), 3))
print(round(wer("give fifteen milligrams daily", "give fifty milligrams daily"), 3))go deeper
Know that word error rate counts substitutions, deletions and insertions against a reference divided by reference word length, and that it treats every word as equally important.
Be able to work the arithmetic showing how a small set of critical terms can be badly wrong while the aggregate stays low, and name normalization as a reason cross-vendor numbers are not comparable.
Demonstrate an evaluation design: sliced WER, entity precision and recall, negation and numeric checks, and confidence-triggered human review, with release gates on the slices.
Own the risk framing — which error classes are tolerable, what human review capacity that implies, and how the metric set feeds a safety case rather than a leaderboard position.
## What WER actually computes Word error rate aligns the hypothesis against a reference transcript and counts the minimum substitutions, deletions and insertions needed to turn one into the other, divided by the number of reference words: `WER = (S + D + I) / N`. It is a normalized edit distance. Two properties follow directly and cause almost every misuse. **Every word is weighted equally.** "Um" costs exactly what "milligrams" costs. **It is text-shaped, not meaning-shaped.** It knows nothing about which errors change what a sentence asserts. ## Blind spot one — the average hides the tail In a 12-minute dictation of maybe 1,800 words, perhaps 40 are drug names, dosages, units or anatomical terms. Suppose the recognizer gets 30% of those wrong — 12 errors. Against 1,800 reference words that contributes about 0.7 points of WER. The headline could read 6% whether that rare-term error rate is 5% or 40%. The number the product lives or dies by is arithmetically invisible in the number the vendor benchmark reports. This is worse than it sounds because rare-term errors are systematic, not random. A model that has never seen a brand-new drug name in training will fail on it *every* time, for every user, on every dictation — a correlated failure that lands on the same clinician repeatedly. ## Blind spot two — semantic weight "No history of chest pain" → "history of chest pain" is one deletion. Formally trivial; clinically it reverses the record. The same applies to numbers ("15 mg" → "50 mg": one substitution) and to units. WER cannot distinguish a dropped filler from a dropped negation, so any evaluation that stops at WER is silent on exactly the errors with the worst downstream consequences. ## Blind spot three — normalization games WER depends heavily on text normalization: casing, punctuation, numbers written as digits or words, contractions, hyphenation. A vendor that normalizes aggressively before scoring will report a lower WER on the same audio than one that does not. Never compare numbers computed by different harnesses; re-run every candidate through your own normalizer on your own audio. ## What to measure instead 1. **Sliced WER.** Report separate rates for general speech and for a curated domain lexicon. Ship the slice, not the average. 2. **Entity-level precision and recall.** Treat drug names, dosages, units, dates and patient identifiers as entities: did each one appear correctly, and did the transcript invent one that was not spoken? Insertions of plausible-looking entities are the dangerous class, because a reviewer has nothing to notice. 3. **Negation and polarity checks.** Programmatically compare negation markers in reference and hypothesis; a mismatch is a hard failure regardless of WER. 4. **Numeric exactness.** Any digit-string mismatch is scored as a failure independently. 5. **Confidence-triggered review.** Route low-confidence spans that overlap the critical lexicon to a human, and measure how much of the residual error that catches. ## Reducing the errors, not just the metric Most providers accept a hint — a vocabulary list, keyword boost, or a short priming prompt naming the expected terms — which materially improves rare-word recognition without retraining. Beyond that: constrain post-processing with a lexicon so near-miss spellings snap to a known term rather than to a plausible English word; and lock the audio path (sample rate, codec, microphone) because the fastest route to a bad WER is a lossy telephony codec, not a bad model. ## Related metrics worth knowing **Character error rate** is the same construction over characters, useful for languages without clean word boundaries and for scoring near-miss spellings. **Diarization error rate** measures speaker attribution and is orthogonal — a transcript can be word-perfect and still attribute every sentence to the wrong person. Reporting all three, sliced, is the honest picture; reporting one aggregate WER is how a product ships that tests well and fails in clinic.
- Two vendors report 5.1% and 6.4% WER on the same benchmark. What do you do with that?Very little. WER is highly sensitive to normalization — casing, punctuation, digit versus word forms, contractions — and to the benchmark's audio conditions, neither of which matches your product. Re-run both on your own recordings through your own normalizer, slice by domain vocabulary, and compare entity-level accuracy. A 1.3-point difference on someone else's corpus routinely reverses on yours.
- Which transcription error class worries you most in a clinical setting, and why?Confident insertions and substitutions that produce a plausible term — a real drug name that was not spoken, or 50 mg for 15 mg. Deletions and garbled output at least look wrong to a reviewer; a fluent, plausible wrong term reads as correct and survives review. That is why I score entity insertions separately and force digit strings to match exactly.
- How would you cut rare-term errors without training your own model?Feed the recognizer the vocabulary: most providers accept a keyword or phrase-boost list, or a short priming prompt naming expected terms, which sharply improves rare-word recall. Then constrain post-processing so near-miss spellings snap to the known lexicon instead of to a common English word. Finally, fix the audio path — sample rate, codec and microphone often cost more accuracy than the model choice does.
saying these in an interview costs you the question
- Treating a single headline WER as the product's quality bar
- Comparing WER across vendors without re-normalizing text
- Assuming all word errors carry equal consequence
- Believing WER captures negation or meaning changes
- Ignoring hallucinated entity insertions because they barely affect WER