You wrote a new probe for the garak LLM scanner and pointed it at a detector that ships with the tool for a different attack family. Why is the failure rate that comes back untrustworthy?
answer
- detector tuned to one family's signature
- misses and false positives together
- rate measures the mismatch
- state the criterion in a sentence
- read the transcripts, not the summary
basics
~20 sA garak detector was written to recognise one family's failure text. Aimed at a different family it misses the real hits and can match unrelated wording, so the rate you get measures the mismatch between probe and detector rather than anything about the target. Both error directions move at once.
solid answer
~50 sShipped garak detectors are cheap rulers - substring lookups, patterns, or a small classifier - each tuned to the failure signature of the family it was written beside. That tuning is exactly what does not transfer. Borrow one for your own probe and you get two independent errors. **Misses**: your family's real failure text does not contain the strings or shapes the detector looks for, so genuine hits score clean. **False positives**: the criterion is loose enough to fire on wording your prompts happen to induce, so benign replies score as hits. Neither is visible in the summary number; both are visible in the transcripts. The test for legitimate reuse is not 'same rough area'. It is: can I state the reply property this detector fires on, and is that property what my family actually causes? If not, you are guessing, and should read a labelled sample of your own transcripts first.
go deeper
Knows the detector decides hits and that a mismatched one will report the wrong thing.
Names both error directions, explains that shipped detectors encode one family's failure signature, and knows to read transcripts.
Insists on writing the failure signature down first, samples flagged and unflagged replies, and refuses to quote a rate without a measured error.
Treats detector-probe pairing as a reporting-integrity policy, not a per-run choice, and sets what must be validated before any custom number leaves the team.
### What a shipped detector actually is In garak a detector is a small object with one job: given an `Attempt` — a prompt and the target's replies to it — return a score per reply, which the evaluator thresholds into hit or pass. The implementations are deliberately cheap. Some are substring matchers configured with a list of strings and a match mode (anywhere in the reply, whole word, at the start). Some apply regular expressions. Some wrap a small off-the-shelf text classifier — a toxicity model, a refusal classifier — and emit its probability. A few invert the test and score a *missing* refusal as the failure. Every one of those was fitted beside a particular probe. The string list is the wording that family's failures actually produce. The classifier was chosen because its label space matches that family's harm. The inversion is correct because for that family, compliance is the failure. That fitting is precisely the part that does not travel. ### The two errors you inherit at once Point your own probe at somebody else's detector and both error directions go live simultaneously. **Misses.** Your family's genuine failure replies do not contain the strings, match the patterns, or land in the class the borrowed criterion looks for. Real hits score clean, and they vanish — a miss leaves no artefact anywhere in the report. **False positives.** The criterion is loose enough to fire on wording your prompts happen to induce. A substring detector for a refusal phrase will fire on any reply that quotes or discusses that phrase. A toxicity classifier will fire on a reply that describes the harm rather than committing it. Benign replies score as hits. Because both are live, you cannot even say which way the rate is wrong. The honest description of a borrowed-detector rate is not "conservative" and not "inflated" — it is *undirected*, and it stays undirected until somebody reads transcripts. | what you get | what it looks like in the report | how you find it | |---|---|---| | miss | nothing at all | read the replies the detector cleared | | false positive | a hit line like any other | read the replies it flagged | | both together | an ordinary-looking percentage | read both sides, then reweight | ### What the mismatch costs The compute cost of the mistake is one wasted run — the same prompts x generations bill you already paid, spent again after you find out. The larger cost is triage: a 1,200-output run read end to end at fifteen seconds per reply is about five hours of somebody's attention. The affordable version is a stratified sample of roughly fifty flagged and fifty cleared replies, which is around half an hour and catches a gross mismatch immediately. The worst cost is reputational, and it is not recoverable by re-running: a number that reached a slide and then moved by ten points when the detector was fixed teaches the audience to discount everything the team publishes. ### Where the number misleads Nothing on the page marks a borrowed ruler. A rate is a rate; the report renders your custom pairing exactly like a shipped one. Downstream that mismatch propagates: if the pairing is quoted against garak's calibrated positioning for the shipped family, you are comparing your prompts against reference data collected with different prompts. And over time it corrupts trend. Re-run next release, watch the rate move, and the natural story is "the model changed" — when a detector whose criterion never fitted will drift with any change in the target's formatting, refusal style, or verbosity. ### What I would check before believing it Write the failure signature of my family as one sentence a colleague could apply by hand without reading code. Then ask whether the borrowed detector's criterion *is* that sentence — not adjacent to it, not in the same area. If I cannot state what reply property the detector fires on, I do not know what I measured. Then the cheap empirical check, before spending a full run: call the detector offline on a handful of replies I have already labelled — some obviously failures for my family, some obviously benign. If it cannot fire on the ones I wrote to make it fire, the mismatch is proven for the price of a unit test. Only after that does a stratified transcript sample from the real run mean anything, and only then does a rate belong in a report — quoted with the sample, the criterion and the measured error beside it. "Close enough" is exactly the condition under which the number is confidently wrong.
- Your borrowed detector reports a suspiciously high rate on a custom probe. Miss or false positive — which do you check first, and how?False positives: pull a sample of flagged replies and read them. A high rate with benign flagged text confirms the criterion is firing on wording your prompts induce rather than on real failures.
- Is there any case where reusing a shipped detector for a new family is fine?Yes — when the failure signature is genuinely the same, for example a leaked canary marker that both families try to extract. State the criterion explicitly and confirm on a labelled sample.
saying these in an interview costs you the question
- Says 'close enough' without stating what the detector fires on
- Assumes reuse only makes the rate conservative
- Quotes the resulting rate without reading any transcripts
- Cannot distinguish a miss from a false positive when explaining the risk