skip to content

In garak, what job does a detector do, and why is it the detector rather than the probe that decides whether an attempt is recorded as a failure?

level: juniorimportance: must knowfreq 72%

answer

  1. probe sends, detector judges
  2. verdict, not behaviour
  3. same replies, several detectors
  4. quoting the number = quoting the detector
  5. name the detector in the report

basics

~20 s

A garak probe only sends prompts. The detector reads each reply and rules whether the attack succeeded. Every pass and failure in the report comes from that ruling, so the number describes the detector's opinion of the outputs, not the model's behaviour directly. Different detectors over the same replies give different numbers.

solid answer

~50 s

In garak the work is split. A **probe** builds and sends the attack prompts; the **generator** carries them to the model under test and captures replies; the **detector** examines each captured reply and decides whether that attempt counts as a failure. The report's pass/failure rate, the absolute score and the calibration Z-score are all downstream of detector verdicts. That split matters because the detector is the only place where "did this work?" is defined. A probe can send a perfectly good attack and the run still reports zero failures because the attached detector never recognised success in the reply. The opposite happens too: an innocuous or refusing reply that happens to contain the wording the detector watches for is scored as a failure. So when you quote a garak number, you are quoting the detector. Name it, say what class of thing it is (string match, pattern, or a small classifier), and be ready to show the replies behind the count.

go deeper

for a junior

Says a probe sends the attack prompts and a detector decides whether each reply counts as a failure, and that the report is built from detector verdicts.

for a middle

Adds that several detectors can score the same captured replies without re-querying, and that changing detectors moves the reported rate with no change in the model.

for a senior

Treats the detector as an inherited error source: names it in the report, hand-verifies a sample of failing attempts, and refuses to compare rates across runs with different detector sets.

for a principal

Sets the house rule that no garak figure leaves the team without its probe set, detector set and a triaged sample, because downstream readers will otherwise treat it as an absolute safety rate.

### The four plugin families, and what each one owns A garak run is assembled from plugins, and the split between them is the whole answer to this question. A **probe** (`garak.probes.*`) owns the attack prompts: it decides what is sent and how many prompts there are. A **generator** (`garak.generators.*`) owns transport: it holds the endpoint, the auth headers and the knowledge of which field in the response body is the model's reply, and it returns that text. A **buff** (`garak.buffs.*`) optionally rewrites prompts on the way out. A **detector** (`garak.detectors.*`) owns adjudication: it reads what came back and rules on it. An **evaluator** then turns detector output into the pass/hit counts you read. Everything a run produces is stored per attempt. Each prompt becomes an `Attempt` record carrying the prompt text, the list of outputs (one per generation), and a `detector_results` map keyed by detector name. The report files (`*.report.jsonl`, the hit log, the HTML summary) are aggregations of that map. There is no other source for a pass or a failure in garak. ### A detector does not return "failed" This is the mechanism people usually miss. A garak detector's `detect(attempt)` returns **one float per captured output**, conventionally in 0.0-1.0, meaning roughly "how much does this output look like the thing I watch for". A separate evaluator applies a **threshold** - 0.5 by default in garak's threshold evaluator - to turn each score into a hit or a pass. So two knobs sit between the model's words and the number in your report: the scoring function inside the detector, and the cut applied to its score. The probe owns neither of them. Change either and the same captured replies produce a different report. ### Why adjudication is split out at all Because the expensive thing is the query, not the judgement. Probes declare a `primary_detector` (older releases call it `recommended_detector`) and often a list of `extended_detectors`, which garak will also run when asked; every one of them scores outputs that are **already captured**. Several harms can be looked for in the same reply - a bypassed mitigation, a leaked string, toxic content - without sending a single extra prompt. ### What a run costs, and where the cost sits Money and wall-clock live entirely on the generator side. garak repeats each prompt `--generations` times, so a probe family with a few hundred prompts is a few thousand metered calls, and the endpoint's rate limit, not your CPU, sets the wall-clock. Detection is close to free by comparison: a substring detector is microseconds per output, and even a classifier detector only adds a one-off model download plus local inference over text already on disk. The practical consequence is worth internalising: **re-adjudicating a stored run costs nothing; re-probing costs the entire run again.** Keep the report and hit-log artefacts, because they are the thing you can re-score for free when you doubt a detector. ### Where the number misleads "Model X failed 40% of attempts" reads like a property of the model. It is not. The numerator is what this detector set scored above the cut; the denominator is however many attempts this probe selection generated, multiplied by `--generations`. Both are configuration you chose. Two specific traps follow. First, some detectors are **inverted**: they score the *absence* of an expected refusal or mitigation as a hit, which means an empty reply, a truncated stream or an HTTP error body captured as text scores as a failure. Second, the report's absolute score and its calibration Z-score are computed from the same verdicts, so a flattering percentile does not launder a bad adjudicator - it just compares you against other runs judged by the same rule. ### What you would check before quoting it Open the run's `.report.jsonl` or hit log rather than the HTML summary; find the `detector_results` entries behind the failures; write down which detector name produced them and what class of matcher it is; hand-read a dozen raw prompt-and-reply pairs; and publish the probe selection, the detector set and the garak version next to any figure. A rate without those three is not reproducible by anyone, including you next quarter.

  • Two garak runs send byte-identical prompts to the same endpoint and get similar replies, yet report very different failure rates. What is the first thing you check?
    Which detectors ran. Adjudication is a separate plugin from the prompts, so a different detector set over comparable replies is the ordinary cause of a different rate.
  • Can one captured reply produce more than one verdict in a garak run?
    Yes. Detectors score already-captured outputs, so several can judge the same attempt for different kinds of success without any extra queries to the target.
  • Why is 'the model failed 40% of attempts' an incomplete sentence for a report?
    It omits the probe set that defined the attempts and the detector set that defined failure. Both must be stated for the number to be reproducible or comparable.

The probe is the person asking the awkward question; the detector is the referee who decides whether the answer broke a rule. Swap referees over the same recorded match and the scoreline changes without a single player moving.

saying these in an interview costs you the question

  • Says the probe decides whether the attack succeeded.
  • Quotes a failure rate without knowing which detector produced it.
  • Assumes a garak failure count is a count of real, exploitable findings.
  • Compares two runs' rates that used different detectors as if they measured the same thing.

context