skip to content

garak's detectors range from plain substring matching, through pattern matching, up to small classifier models run locally. What does that range change about a scan's cost and about the errors you inherit?

level: middleimportance: must knowfreq 62%

answer

  1. literal vs pattern vs small classifier
  2. text matchers: free, brittle, legible
  3. classifier: download, inference, opaque
  4. misses cluster on rewording and language
  5. triage cost differs by matcher class

basics

~20 s

String and pattern detectors are nearly free and deterministic, but they only see the exact wording they were written for, so paraphrase, translation or odd formatting slips past. A classifier detector generalises further, costs local compute and a model download, and adds probabilistic errors you did not tune and cannot inspect per item.

solid answer

~60 s

The trade runs along two axes: cost and error shape. **Cost.** String and pattern detectors are trivial to run; they add essentially nothing to wall-clock time, which for most scans is dominated by the target endpoint's latency and rate limit anyway. A classifier detector has to be fetched and loaded, and then runs inference over every captured reply, so it adds local compute, memory and a first-run download. On a large sweep that can turn a scan that was purely network-bound into one that also queues on your own hardware. **Error shape.** A string or pattern detector is brittle but legible: you can read the rule and predict what it will and will not catch, and its misses cluster on rewording, other languages, encodings and formatting. A classifier is more tolerant of rewording but its boundary is opaque; it will produce confident wrong verdicts on inputs unlike its training data, and you cannot argue with any single verdict by reading a rule. Either way you inherit the error. Nothing in the report separates a verified failure from a matcher artefact.

go deeper

for a junior

Knows detectors are not all the same: some look for text, some are small models, and the text ones are cheap but easy to slip past.

for a middle

Explains both axes - runtime and memory cost versus brittleness and legibility - and can name the concrete escape routes from a literal match: rewording, another language, encoding, formatting.

for a senior

Plans the run around it: checks which detector classes the selected probes pull in, budgets local inference and download, and hand-scores a sample before trusting a classifier detector on an unfamiliar output style.

for a principal

Decides the house standard for which matcher class may back a published number, and requires the matcher class to be stated whenever a rate is quoted.

### What each matcher class can physically see garak ships detectors in roughly three families, and the difference between them is not quality - it is what kind of evidence each one is capable of noticing. A **substring detector** (garak's `StringDetector` and the many detectors built on it) holds a list of literals and asks whether any of them appears in the output. It exposes a `matchtype` - matching anywhere in the string, or only on word boundaries - and returns 1.0 or 0.0 per output. `TriggerListDetector` is the same idea with the literals supplied by the probe itself: the probe plants a trigger string in the attempt's notes and the detector checks whether the model repeated it. These are deterministic, auditable line by line, and you can predict their behaviour before running anything. A **pattern detector** widens that to regular expressions or structural shapes - a code fence, a key-looking string, a family of phrasings. More reach, but the rule is now something a reviewer has to read carefully to know what it accepts, and a pattern tuned against one model's house phrasing is often near-useless against another endpoint's. A **classifier detector** (garak's Hugging Face-backed detectors, e.g. the toxicity family) loads a small sequence-classification model, runs each captured output through it, and takes the probability of a named label as the score. That score is then thresholded - `detector_threshold` in the detector's params, 0.5 by default - to produce a hit. ### What each class costs | | wall-clock | setup | triage cost | |---|---|---|---| | substring / trigger | microseconds per output | none | read the literal list, minutes | | pattern | microseconds per output | none | read and reason about the regex | | classifier | 20-150 ms per output on CPU | model download of a few hundred MB, RAM or VRAM to hold it | hand-score a sample; there is no rule to read | Put numbers on the classifier row, because that is the one that surprises people. A modest run - say 300 prompts at `--generations 10` - is 3,000 outputs. Two classifier detectors over that is 6,000 forward passes; on CPU that is single-digit to low-double-digit minutes, plus the first-run download. Against a metered endpoint that throttles you to a few requests a second, that is still small next to the generation time, so the honest statement is not "classifiers make scans slow" - it is that they add a **local** resource requirement that a CI container or a laptop may not have, and they turn a purely network-bound job into one that also needs memory and disk. ### Where the number misleads, per class A brittle matcher's misses are **not random noise, they are a systematic bias in one direction**: content that arrives paraphrased, translated, encoded, expressed as code, or split across lines scores 0.0 and is silently a pass. So a substring-backed rate understates in a predictable way, and it understates *more* against models with an unusual output style. Inverted string detectors bias the other way. A mitigation-bypass style detector looks for the absence of refusal wording, so an empty reply, a truncated stream, or an error body captured as text reads as a bypass and inflates the rate. The classifier's failure is different again. Its scores are continuous and many of them sit near the 0.5 cut, and that cut was inherited, not chosen for your target. A change in the *style* of replies - a new branded preamble, longer answers, more structure, another language - shifts a whole cluster of scores across the line at once. The rate then jumps between two model releases with no change in the underlying behaviour, and nothing in the report signals low confidence. ### What you would check List what is actually attached (`garak --list_detectors`, and the probe's declared primary and extended detectors). For any classifier detector, note the model it pulls and the threshold in its params, and run one small scan first to measure the download and per-output latency on the hardware you intend to use. Then hand-score around 50 outputs from your own target against a written criterion: that agreement figure, on *your* target's output style, is the only evidence you will get about whether the detector's boundary transfers.

  • Your scan's wall-clock was previously set by the target's rate limit. What changes if the probes you add bring classifier-based detectors?
    You add local inference over every captured reply plus a first-run model download, so the run can become bound by your own hardware as well as by the endpoint.
  • Why is triaging a string-matcher's failures usually faster than triaging a classifier's?
    You can read the rule and immediately see why each reply matched. A classifier gives no rule, so the only way to judge it is to hand-score a sample against human judgement.

A substring detector is a printed list of banned words at the door: you can read exactly what it stops, and anyone who rephrases walks straight through. A classifier is a sniffer dog - it catches things the list never named, but nobody can ask it why it barked.

saying these in an interview costs you the question

  • Treats all garak detectors as interchangeable text matchers.
  • Assumes a classifier detector is simply more accurate, with no error of its own.
  • Ignores the local compute and download cost of classifier detectors when sizing a run.
  • Cannot name a single way a reply can carry the harmful content past a literal string match.

context