skip to content

Which failure modes of garak's detectors produce false hits, and how do you recognise each one from the logged attempt?

level: middleimportance: must knowfreq 66%

answer

  1. matcher: who authored the matched span
  2. refusal-absence: no harm present at all
  3. classifier: topical but not capable
  4. truncation flips verdicts both ways
  5. guard text is not model text

basics

~20 s

Detectors that match a substring, or that infer compliance from the absence of a refusal, fire on responses that never complied: the model quoted the request back, refused in unusual wording, returned an error string, or was cut off mid-sentence. Classifier detectors add their own misfires. Read the response text to tell them apart.

solid answer

~50 s

Group the artefacts by the detector family that causes them. **Keyword or pattern matchers** fire when the trigger text appears anywhere — inside the prompt the model echoed back, inside a warning it wrote about the request, or inside a quoted refusal. Recognise it by finding the matched span in a position that is not the model's own compliance. **Refusal-absence heuristics** treat "no recognised refusal phrase" as compliance. They fire on error strings, empty responses, non-English refusals and unusual phrasings. Recognise it by the response containing no harmful content at all. **Classifier detectors** misfire on style: a response *about* the topic, a fictional framing, or a safety explanation can score like the real thing. Recognise it by the response being topical but not capable. **Truncation and transport** cut across all three: a clipped response can flip either way, and a guardrail in front of the endpoint may have written the text rather than the model.

go deeper

for a junior

Knows detectors can be wrong and that you read the response to check; may name only substring matching.

for a middle

Names the detector families and gives a concrete recognition test for each, including prompt echo and error text.

for a senior

Adds truncation, retries and interposed guardrails as cross-cutting causes, and converts the families into a sampling policy over buckets.

for a principal

Argues about which detector families are acceptable evidence for which claims, and insists the report states the sampling rate behind each number.

**Why sorting by cause beats sorting by hunch.** A garak sweep emits far more flagged rows than anyone will read, so the credibility of your report rests on being able to say *which mechanism* produced each false hit. "It's a false positive" is an unfalsifiable shrug; "this is a surface matcher firing on prompt text the model echoed back" is a claim someone can check against the transcript and either confirm or overturn. Interviewers listen for the second shape. **Family 1 — surface matchers (substring and regex detectors).** These ask only whether a trigger string occurs anywhere in the generator's returned text. They misfire whenever the trigger appears in text the model did not author as compliance: the model restates the request verbatim before declining it; the model writes a cautionary paragraph containing the very words the matcher hunts; the model quotes a policy or a safety notice that contains the term. *Diagnostic:* locate the matched span in the untruncated output and ask **who authored it**. If the span is the prompt reflected back, or sits inside a warning about the request, the match is real and the hit is not. **Family 2 — refusal-absence heuristics.** These invert the question: no recognised refusal phrase was found, therefore the model complied. The consequence is that everything which is *neither* a refusal *nor* a compliance is scored as a hit — provider error bodies rendered as text, empty completions, timeouts, content-filter stubs, refusals written in a language whose phrasing is not in the list, and terse conversational deflections. *Diagnostic:* the output contains no harmful content at all. Nothing to verify means nothing to report. This family is simultaneously the largest source of inflated counts and, because it buries real hits in noise, a common reason genuine findings get skipped during a rushed triage. **Family 3 — classifier detectors.** These generalise better and fail more subtly, because they score *topic and register* rather than capability. A model explaining why a request is dangerous is topically indistinguishable from one complying; a hypothetical, satirical or fictional framing can score high while conveying nothing actionable; a verbose safety lecture can outscore a terse real answer. *Diagnostic:* ask whether a reader could act on the text as written. Discussion of a subject is not capability in that subject. **Cross-cutting causes that ride on all three.** - **Truncation.** A clipped output flips the verdict in both directions: a refusal cut off after its restatement can read as compliance; a compliance cut off before the substance reads as clean. It is the nastiest artefact because it leaves no signature in the counters. - **Transport and retries.** Rate-limited or retried calls log placeholders and partials that a detector scores as ordinary text. - **Interposed guardrails.** If a filter sits in front of the model, the logged output may be the guard's canned message. That tells you the filter caught this prompt shape — a fact about the deployed stack, not about the model — and it moves the remediation from model behaviour to filter coverage. **What it costs.** Each classification decision costs one untruncated transcript read: two to five minutes each, honestly. That is the only meaningful budget line in triage, and it is why you sample by family rather than sweeping. A sweep of a few hundred hits is more than a working day; the discipline that survives contact with the calendar is *bucket, sample a fixed n per bucket, classify each read hit, and let the family base rate characterise the remainder*. **Where the number misleads.** Two ways, both common. First, the artefact families are not evenly distributed across probes: a probe scored by a refusal-absence detector can carry an artefact rate far above one scored by a classifier, so comparing raw hit counts between probes silently compares detector eagerness rather than model weakness. Second, artefacts are not merely noise you can subtract — they are *asymmetric*. Families 1 and 2 inflate; family 3 both inflates (discussion scored as capability) and deflates (an unusual real compliance the classifier never learned). Reporting "we removed the false positives" therefore does not leave you with a clean count; it leaves you with a count whose remaining error you have not measured. **What you check before signing.** Read a spread across each bucket rather than the first rows, since attempts are usually ordered by prompt variant and the head of a bucket is one variant repeated. Confirm the response is the model's, not a guard's. Confirm nothing was clipped. Where the sampled hits of a bucket are all one artefact family, do not quietly delete the bucket — say in the report that the behaviour is **untested** and, if it matters, rescore it with a stronger detector.

  • Why is truncation more dangerous than the other artefacts?
    It flips the verdict in both directions and leaves no trace in the counters — a clipped refusal can look like compliance and a clipped compliance can look clean.
  • The transcript shows the guardrail's canned block message rather than a model answer. What have you learned?
    That the deployed filter caught this prompt shape. It says nothing about the underlying model, and the remediation lives with the filter's coverage rather than with model behaviour.
  • How do you handle a probe where the sampled hits are all artefacts of one family?
    Treat the whole bucket as unreliable, say so explicitly in the report rather than quietly deleting it, and either re-score with a better detector or mark that behaviour as untested.

saying these in an interview costs you the question

  • Calling every false hit 'a false positive' without naming the detector behaviour that caused it.
  • Assuming a detector understands meaning rather than surface features.
  • Judging from a truncated snippet in a summary view.
  • Not noticing that an interposed filter, not the model, produced the response.

context