Why do off-the-shelf PII detectors miss hospital MRNs, and what does over-redaction cost?
answer
- recognizers ship for standard formats only
- facility-assigned means no shared pattern
- measure recall per entity type
- masking numbers deletes the dosages
- reversible placeholders beat destruction
basics
~20 sShipped recognizers cover identifiers standardised nationally or industry-wide; a facility-assigned medical record number has no fixed format, so nothing matches it. Misses leak silently, while blanket masking strips the dosages, dates and lab values the answer depended on.
solid answer
~50 sA PII detector is a classifier, not a boundary: regex patterns, checksum validators, deny lists and a name-recognition model, each emitting scored spans. Recall is only as good as the recognizer set, and the recognizers ship for identifiers with a stable format — card numbers, email addresses, national identifiers. A hospital MRN is assigned per facility with its own width and prefix, so no stock pattern fires; the same is true of internal case, claim and accession numbers, which are exactly the identifiers that re-identify a document inside the organisation. In Presidio you close the gap by registering a custom `PatternRecognizer` with its own regex, a validation step and context words that raise confidence when nearby text says "MRN" or "patient". Then measure recall **per entity type** on labelled real documents. Push the other way and you pay too: mask every number and dosages and dates vanish, answers degrade, and clinicians route around the tool into something unmonitored.
code
python · 15 linesfrom presidio_analyzer import AnalyzerEngine, Pattern, PatternRecognizer
mrn = PatternRecognizer(
supported_entity="MEDICAL_RECORD_NUMBER",
patterns=[Pattern(name="mrn", regex=r"\bMRN[:\s-]?\d{6,9}\b", score=0.6)],
context=["mrn", "chart", "patient"],
)
analyzer = AnalyzerEngine()
analyzer.registry.add_recognizer(mrn)
note = "Discussed discharge with the family, MRN 0042731, follow-up Tuesday."
results = analyzer.analyze(text=note, entities=["MEDICAL_RECORD_NUMBER"], language="en")
for r in results:
print(r.entity_type, note[r.start:r.end], r.score)go deeper
Know that PII detection is pattern and model matching with imperfect recall, and that identifiers unique to your organisation need recognizers you write yourself.
Explain how a custom recognizer is built — pattern, validation, context words — and why recall and precision must be reported separately per entity type on real documents.
Demonstrate that you would gate the pipeline: labelled sample in CI, per-entity recall floors, an adversarial slice, and reversible placeholders so utility survives. Name the failure where an over-redacted tool drives users to unmonitored channels.
Own the utility-versus-exposure frontier as a product decision, and say who signs off on the residual leak-through rate. Redaction is one layer in a budget that also includes minimization and isolation; funding it alone buys less than it appears to.
## What a detector actually is PII detection combines several weak signals. Deterministic recognizers match a pattern and optionally validate it — a card number matches a digit shape and then passes a Luhn check; a national health number matches ten digits and then passes a modulus-11 check digit. Context words near the span raise the score. A named-entity model handles the things with no format at all: person names, organisations, locations. An anonymizer then applies an operator per entity type — replace, mask, hash, encrypt. Every part of that is probabilistic. The output has a recall figure (what fraction of real identifiers were caught) and a precision figure (what fraction of flagged spans were really identifiers), and neither is 100%. That framing matters more than any implementation detail, because it tells you redaction cannot be the thing that carries an access-control decision. ## Why domain identifiers are the ones that leak Shipped recognizers exist where a format is standardised across a country or an industry, because only then can one regex generalise. A medical record number has no such standard: each facility chooses its width, prefix and zero-padding, and the same digit string means something different at the hospital next door. Internal case numbers, claim numbers, accession numbers, device serials, staff badge numbers and account references behave the same way. These are the highest-value misses. Precisely because they are internal, anyone inside the organisation with access to the record system can resolve one back to a person — so a "de-identified" note that still carries an MRN is not de-identified at all for the population most likely to read the logs. Free text is where they hide. Structured fields can be handled by schema; a note reading "discussed with the family, chart 0042731, Tuesday" needs the detector to fire mid-sentence. Real clinical text also brings hyphenation, OCR noise from scanned documents, identifiers broken across line ends, and abbreviations that appear nowhere in the training data of a general-purpose name model. Person-name recall in that setting is materially worse than the benchmark numbers you will see quoted, which were measured on clean edited prose. ## Building recall you can defend Register custom recognizers for each identifier family your domain uses: a regex, a validation function where a check digit or range exists, and context words. Give each a score you can tune, and prefer a lower threshold with a validator over a high-confidence bare regex. Then measure, and measure the right thing. Label a sample of real production documents span by span, and report recall and precision **separately for each entity type**. Aggregate accuracy is close to useless here, because email addresses are easy and dominate the count while the MRN you actually worry about is rare. Track a leak-through rate — identifiers surviving per thousand documents — because that is the number a risk owner can reason about. Keep an adversarial slice: OCR artefacts, digits written as words, identifiers split by newlines, unusual and transliterated names. Re-run the suite whenever recognizers or the underlying model change, since a detection pipeline silently regresses like any other classifier. ## The other direction, and why it is not free Recall is easy to buy if you stop caring about precision: mask every digit run and every capitalised token. What that destroys is the content the task depended on. In a clinical note the digits are dosages, lab values, gestational ages, dates and intervals; strip them and the assistant produces confident, fluent summaries that are wrong. Mask all names and a note discussing two patients collapses into ambiguity that the model resolves by guessing. There is a second-order cost that decides most real programmes. If the redacted tool gives poor answers, clinicians stop using it and paste the same note into whatever consumer chatbot is on their phone — an unmonitored, unlogged, entirely uncontrolled path. Over-redaction can therefore produce a worse privacy outcome than the leak it prevented. Utility is a safety property here, not a competing concern. ## Reversible placeholders The usual escape is to replace rather than destroy: substitute stable placeholders (PATIENT_1, MRN_1, DATE_1) held in a server-side map for the life of the request, and rehydrate the model's output afterwards. This preserves referential structure — the model can still tell that the same person appears three times — and keeps relative dates usable, while the sensitive values never enter the prompt, the logs or the cache. The cost is holding the map, which is itself sensitive and short-lived by design, plus the rehydration step becoming a correctness dependency. ## Where redaction belongs in the stack Redaction reduces the blast radius of copies you cannot eliminate. It does not decide who may see what — isolation and access control do that below the model — and it does not stop a determined extraction. Order the layers: minimise what is fetched at all, redact what remains, isolate the corpus by tenant and role, and keep prompt bodies out of default logging. Redaction sitting alone as the only control is the configuration that ends up in an incident report.
- How would you decide the confidence threshold for a custom recognizer?Empirically, from a labelled sample, and per entity type. Look at the recall and precision curve as the threshold moves, and pick based on which error hurts more for that identifier: for an MRN, a false positive costs one masked token, while a miss is a re-identifiable leak, so you set it low and lean on a validator or context words to keep precision usable.
- Why is person-name detection usually the weakest part of the pipeline?Names have no format, so they rely on a named-entity model rather than a pattern, and that model was trained largely on clean edited text. Clinical and operational notes bring misspellings, abbreviations, transliterations, OCR noise and names that double as ordinary words. Expect measurably lower recall there, measure it separately, and do not let a single aggregate accuracy number hide it.
- What does reversible placeholder mapping cost you?You must hold the map — placeholder to real value — for the life of the request, which is itself sensitive state that needs a short lifetime and its own access control. Rehydration also becomes a correctness dependency: if the model invents a placeholder that was never issued, or mangles one, the output must fail closed rather than emit a broken identifier.
- How would you catch a silent regression in redaction quality after a model or library upgrade?Treat the labelled sample as a test suite and run it in CI, asserting per-entity recall floors and a leak-through rate ceiling rather than an aggregate score. Include the adversarial slice — OCR noise, split identifiers, unusual names. A detection pipeline degrades quietly, so without a gate the first signal is an incident.
saying these in an interview costs you the question
- Assumes a stock detector catches domain-specific identifiers
- Reports one aggregate accuracy number instead of per-entity recall
- Masks every digit and calls the loss of clinical detail acceptable
- Treats redaction as a guarantee rather than a risk reducer
- Ignores that an unusable tool pushes users to unmonitored channels