skip to content

Scanner Plugins

A scan is built from the plugin roles you configure: the attack family, what rules a hit, the target connection, the transform, and what pairs probe to detector. Swap one and it proves something else.

on this pageshow

explore

questions

25

You want the garak LLM scanner to cover an attack family it does not ship a probe for. Beyond writing the prompt set itself, what else must you supply before a run can report a failure rate?

level: juniorimportance: must knowfreq 62%

answer

  1. prompts are the cheap half
  2. something must rule each reply
  3. plugin discovery must find it
  4. no judge, no rate
  5. silent zero = not loaded

basics

~20 s

Prompts alone produce no number. You must also supply, or point at, the thing that decides whether each reply counts as a hit, and make both discoverable to the scanner so the run loads them. Without that judging step garak collects replies but has nothing to score.

solid answer

~50 s

A garak run has two halves for every attack family: something that sends prompts to the target, and something that rules on the replies. Ship only the first half and the run produces transcripts, not a rate. So a new family means: the prompt-generating plugin, a stated pairing to whatever will judge its replies, and either an existing shipped judge that genuinely matches your failure signature or a new one you write. You also have to make the code visible to the scanner's plugin discovery and name it on the run, otherwise it is simply never loaded. The trap for a beginner is thinking the prompts are the deliverable. They are the cheap half. The judging half is where the number comes from, and it is the half you now own with no shipped baseline behind it.

go deeper

for a junior

Says a run needs both prompts and something that rules replies, and knows the scanner has to be able to load the new code.

for a middle

Adds that the judge must match the family's actual failure signature, and can debug a silent zero-hit result.

for a senior

Treats the judging half as the deliverable, and will not quote a rate from a custom family without saying how often the judge is wrong.

for a principal

Asks whether the family deserves a permanent plugin at all, given who will keep its judging criterion honest over time.

### The pipeline you are inserting into garak is an LLM vulnerability scanner, and it splits a scan into plugin roles that form a pipeline. A **generator** owns the connection to the target — a hosted model selected with `garak --model_type` and `garak --model_name`, or a class you write yourself when the target is a deployed application rather than a raw model. A **probe** owns one attack family: it holds or mints the prompts, and it declares which detectors should rule its replies. Optional **buffs** rewrite prompts before send. A **detector** receives an `Attempt` — one prompt plus the target's replies to it — and returns one score per reply. An evaluator thresholds those scores into hits and passes, and the harness writes a JSONL report, one line per attempt plus an eval line for each probe/detector pair. Read the pipeline and the answer falls out. Prompts drive sends; detectors produce scores; the evaluator turns scores into a rate. Supply only the prompts and the run produces attempt lines and no eval line — a pile of transcripts, not a failure rate. ### The three things that must be true 1. **The prompts exist, and you know their nature.** A fixed list gives you reruns that cover identical ground; prompts minted at run time do not, so a rate that moves between runs may be sampling rather than the target. 2. **A detector is bound to the probe.** In garak that binding is an attribute on the probe class naming its recommended detector, or a detector named explicitly on the command line with `garak --detectors`. Nothing is inferred from the probe's name or module; an unbound probe is scored by nothing. 3. **garak can import both.** Probes and detectors are resolved by dotted plugin name inside garak's own `garak.probes` and `garak.detectors` namespaces. Code sitting in a directory garak never imports is not broken — it is absent. `garak --list_probes` and `garak --list_detectors` are the one-command check that your class is visible at all. ### What the run costs Sends multiply out as prompts x `garak --generations` x (buffs + 1). A 120-prompt probe at `--generations 10` is 1,200 completions against one target for that one family. Read the `--generations` default off your own `garak --help` rather than trusting memory — it has changed across releases, and it multiplies every other number in the run. At a metered endpoint producing a few hundred output tokens per reply that is a visible invoice line, and at a 60-requests-per-minute limit it is roughly twenty minutes of wall clock; `garak --parallel_attempts` cannot outrun the endpoint's own limiter. The lopsided cost is engineer time: the prompt set is an afternoon, while a detector you can defend — plus the hand-labelled sample that says how often it is wrong — is days. ### Where the number misleads The characteristic failure is the **silent zero**. A probe garak never imported, a probe with no detector bound, and a detector whose criterion cannot fire on this family's failure text produce output that looks the same from the summary: nothing reported for your family. Worse, an absent eval row and a genuine 0% eval row read identically to anyone skimming a report, and both are cheerfully summarised as "we tested that and it was clean." A zero from a custom family is a claim about your plumbing until you have shown otherwise. The second trap is the denominator. garak's rate is computed over outputs, not prompts, so `--generations` sets the denominator. A run at 5 generations and a run at 10 are not comparable for the same probe, and "we re-ran it and the rate dropped" is very often just that flag having changed. ### What I would check Confirm the class appears in `garak --list_probes`. Run the probe against a cheap target with a single prompt and verbose output, and confirm attempt lines actually appear in the JSONL report with non-empty replies. Confirm an eval line exists for your probe/detector pair, with an attempt count matching prompts x generations. Then exercise the detector offline on two hand-written replies you have already labelled — one plainly a failure for your family, one plainly benign — and check that it fires on the first and not the second. A detector that cannot fire on a reply you wrote to make it fire will never fire on a real one.

  • A custom garak family runs to completion and reports zero hits. Name two causes that are not the model being safe.
    The plugin never loaded or was never named on the run, so nothing was sent; or the judge's matching criterion cannot fire on this family's failure text, so hits were sent and never counted.
  • Why does it matter whether your custom prompts are a fixed list or minted at run time?
    A fixed list gives comparable reruns and can be tuned against; generated prompts mean two runs cover different ground, so a change in the rate may be sampling, not the target.

The prompts are the exam paper and the detector is the answer key. Print the exam without a key and you get back a stack of completed papers and no grade — which is exactly what a garak run with an unbound probe hands you.

saying these in an interview costs you the question

  • Thinks writing the prompt list completes the job
  • Cannot say what turns replies into a score
  • Reads a zero-hit custom family as a clean result without checking the plugin loaded
  • Reuses whatever judge is default without checking it matches the family

context

open as a page

In garak, what job does a detector do, and why is it the detector rather than the probe that decides whether an attempt is recorded as a failure?

level: juniorimportance: must knowfreq 72%

basics

~20 s

A garak probe only sends prompts. The detector reads each reply and rules whether the attack succeeded. Every pass and failure in the report comes from that ruling, so the number describes the detector's opinion of the outputs, not the model's behaviour directly. Different detectors over the same replies give different numbers.

open as a page

In garak, what is a generator, and what do you have to own yourself when you wire one to a deployed chat application's authenticated HTTP endpoint instead of to a model you run locally?

level: juniorimportance: must knowfreq 70%

basics

~20 s

A garak generator is the adapter that sends each probe prompt to the system under test and returns the reply text. Aimed at a deployed app, you own the wiring: the endpoint and method, the auth and tenant headers, the request body shape, and which JSON field the reply is read from.

open as a page

In the garak LLM scanner, what does a probe supply to a scan, and what does a clean result across the probes you selected say about the probes you did not run?

level: juniorimportance: must knowfreq 70%

basics

~20 s

A garak probe is the plugin that supplies the prompts sent to the target for one attack family. A clean result covers only the probes you selected; it says nothing about families you never ran. Coverage means probes run out of probes available, and the default selection is not everything.

open as a page

You run a set of garak probes against an endpoint twice: once plain, once with a buff enabled that rewrites the probes' prompts before they are sent. The reported pass/fail rate is very different. Why can that difference come from garak's detectors rather than from the model's behaviour?

level: middleimportance: must knowfreq 60%

basics

~20 s

Detectors judge the reply text, and many are substring or pattern matchers. A buff changes the prompt, so the reply's wording, language or format changes too. A matcher tuned for the original wording then misses real hits or fires on harmless replies. The number moved because detection moved, not the model.

open as a page

You wrote a new probe for the garak LLM scanner and pointed it at a detector that ships with the tool for a different attack family. Why is the failure rate that comes back untrustworthy?

level: middleimportance: must knowfreq 58%

basics

~20 s

A garak detector was written to recognise one family's failure text. Aimed at a different family it misses the real hits and can match unrelated wording, so the rate you get measures the mismatch between probe and detector rather than anything about the target. Both error directions move at once.

open as a page

garak's detectors range from plain substring matching, through pattern matching, up to small classifier models run locally. What does that range change about a scan's cost and about the errors you inherit?

level: middleimportance: must knowfreq 62%

basics

~20 s

String and pattern detectors are nearly free and deterministic, but they only see the exact wording they were written for, so paraphrase, translation or odd formatting slips past. A classifier detector generalises further, costs local compute and a model download, and adds probabilistic errors you did not tune and cannot inspect per item.

open as a page

You point garak at a hosted application endpoint that allows 60 requests per minute per account. The planned run is 3,000 probe prompts with 5 generations each. What sets the wall-clock, and which levers actually shorten it?

level: middleimportance: must knowfreq 65%

basics

~20 s

Requests, not probes, set the clock: prompts times generations, divided by the requests per minute the endpoint allows. Here that is 15,000 requests, about 250 minutes at best. Parallelism above the cap only earns throttled responses and retries, which add requests. Cut generations, narrow the probe selection, or get a higher quota.

open as a page

In the garak LLM scanner, some probes send a fixed prompt list shipped inside the package while others mint their prompts during the run. How does that difference change what a passing result for a probe attests to?

level: middleimportance: must knowfreq 60%

basics

~20 s

A fixed-list probe sends the same readable strings every time, so a pass means only those exact prompts failed to elicit the behaviour, and anyone can tune against them. A prompt-generating probe sends different prompts each run, so a pass is a sample you cannot enumerate afterwards and a rerun tests different ground.

open as a page

You wrote both the probe and the detector for a custom family in the garak LLM scanner, and the run reports a 22% failure rate. What do you do before that number goes into a report?

level: seniorimportance: must knowfreq 55%

basics

~20 s

Measure your own detector's error. Sample flagged replies and read them by hand for false positives, and sample unflagged replies for misses. Label a fixed set, ideally with a second reader, and quote the 22% with that measured error attached rather than as a bare number nobody has checked.

open as a page

A garak scan against your deployed chat endpoint reports a high failure rate for one probe family. Before that number goes into a report, how do you establish that it reflects real model behaviour rather than detector error?

level: seniorimportance: must knowfreq 58%

basics

~20 s

Open the logged prompt-and-reply records behind those failures and hand-read a sample. Check what the detector actually matched: refusal text, a stock disclaimer, an echo of the prompt, an empty or error reply, or a safety lecture can all trip a matcher. Then report the verified count, the detector used and the sample size.

open as a page

In the garak LLM scanner, a buff is a plugin that rewrites a probe's prompts on the way out, before they reach the target endpoint. What does enabling a buff do to the amount of work a scan does?

level: juniorimportance: should knowfreq 45%

basics

~20 s

A buff sits between the probe and the target and transforms each prompt. Some buffs return one variant per prompt, others fan out into several, so the attempts sent multiply on top of the per-prompt generation count. Enabling buffs makes a scan longer and more expensive, never cheaper.

open as a page

A deployed assistant's HTTP endpoint answers with a JSON envelope rather than a bare string. What must garak's HTTP generator be told about that envelope, and what does a scan report look like when that setting is wrong?

level: middleimportance: should knowfreq 50%

basics

~20 s

It must be told where in the response body the assistant's text lives, because the endpoint returns an envelope, not a bare string. If that selector is wrong, garak stores the wrapper, an error blob or an empty string, and every attempt is judged against text the model never produced.

open as a page

You are asked to report whether rewriting prompts changes how often a target fails a chosen set of garak probes. How do you set up the buffed and unbuffed garak runs so the comparison actually means something, and what is the trap in the rate itself?

level: seniorimportance: should knowfreq 38%

basics

~20 s

Run the same probes twice against the same endpoint, changing only the buff. Hold generations, detectors and endpoint settings fixed, and run them close together. Then compare like for like: a fan-out buff sends several variants per original prompt, so decide whether your rate counts attempts or original prompts.

open as a page

Some of garak's buffs rewrite a probe's outbound prompts by calling a language model to paraphrase or translate them. What does that add to a scan you have to defend, and how do you handle it?

level: seniorimportance: should knowfreq 40%

basics

~20 s

It makes the run non-reproducible: a rerun paraphrases differently, so the prompts sent are not the same set. The paraphraser can also soften or refuse the payload, silently weakening the probe. Keep the unbuffed run as the baseline, read a sample of transformed prompts, and cite the logged prompts, not the buff name.

open as a page

A garak report can position a probe's result against previously measured reference models as well as giving an absolute score. Why does a probe and detector you wrote yourself not get that comparative reading, and what do you present instead?

level: seniorimportance: should knowfreq 40%

basics

~20 s

The comparative reading needs prior results for that exact probe and detector across reference models. Your plugin has never been run against them, so there is nothing to compare with. Present the absolute rate, your own baseline runs against two or three comparison targets, and the detector's measured error.

open as a page

How can a garak detector record a pass on a reply that was in fact a successful attack, and what does that imply about how you present the run's failure rate?

level: seniorimportance: should knowfreq 48%

basics

~20 s

It happens whenever the harmful content arrives outside what the detector looks for: another language, an encoding, code or pseudo-code, a summary, or wording the matcher was never written for. Those attempts score as passes. So the reported failure rate is a floor on the model's weakness, never a measure of it.

open as a page

garak treats each attempt as independent, but the deployed assistant you are scanning keeps conversation history server-side, keyed by a session identifier your generator sends in a header. What breaks, and how do you wire the generator so the run is sound?

level: seniorimportance: should knowfreq 45%

basics

~20 s

Reusing one session id makes attempt N see attempts 1 through N-1, so earlier prompts and refusals condition later replies and the run stops being reproducible or order-independent. Mint a fresh session or conversation id per attempt, or use a stateless mode of the endpoint, and confirm isolation with a memory probe.

open as a page

You rerun a garak probe that builds its prompts during the run against the same endpoint after shipping a mitigation, and the failure count drops. How do you establish whether the mitigation worked rather than the probe simply having sent different prompts?

level: seniorimportance: should knowfreq 40%

basics

~20 s

A probe that mints prompts each run does not send the same prompts twice, so one lower count is a different sample, not a result. Replay the specific saved prompts that failed before against the patched system, and rerun the generating probe several times on both builds before claiming an improvement.

open as a page

A vendor hands you a garak run in which their chat product passed every probe whose prompts come from a fixed, publicly readable list. What do you check before accepting that as evidence of robustness, and what would you ask them to run instead?

level: seniorimportance: should knowfreq 45%

basics

~20 s

Ask which probes ran and whether their prompts are a public fixed list. A public list can be filtered or tuned against, so a pass shows those strings are handled, not that the behaviour is absent. Ask for a rerun that includes probes minting prompts at run time, with the probe selection and logs attached.

open as a page

Your team wants a risk covered that the garak LLM scanner ships no probe for. When would you decline to write a custom probe and detector pair, and what would you do instead?

level: principalimportance: should knowfreq 32%

basics

~20 s

Decline when nobody will own the ruling criterion over time. A custom pair needs re-labelling as the target changes, and a drifting criterion quietly moves the number while looking stable. For a one-off question, test by hand or with a small held-out set; reserve custom plugins for risks you will re-run every release.

open as a page

Your team runs garak against every model release and publishes the failure rates to other teams. The detectors' error rates ship with the tool and you never measured them. How do you set policy for how those numbers may be used?

level: principalimportance: should knowfreq 38%

basics

~20 s

Treat the rates as a trend within one pinned probe and detector set, never as an absolute safety figure. Pin that configuration across releases, hand-audit a sample of failures and passes each cycle to estimate detector error, publish the audited numbers with sample sizes, and never let a rate alone gate a release.

open as a page

Before a garak engagement on a customer-facing assistant, you must choose where the generator points: the vendor model endpoint behind the app, a staging deployment of the app, or live production. How do you argue the choice, and what does each option's number actually mean?

level: principalimportance: should knowfreq 35%

basics

~20 s

Name what each choice measures. The bare model measures the model; a staging deployment measures the app's prompts, retrieval and filters; production measures the live system plus its abuse controls, at the cost of polluted logs, alerts and real spend. Default to a staging build that mirrors production, and state the choice in the report.

open as a page

Your team has a fixed number of calls it may spend on garak scans of a product endpoint each release. How do you decide whether any of that budget should go to buffed runs — where a buff rewrites probe prompts before sending — instead of more probes?

level: principalimportance: nice to knowfreq 28%

basics

~20 s

Spend the budget on breadth first: buffs add no new behaviour, they re-send prompts you already have. Buy buffed runs for a narrow slice where input normalisation or a filter is the thing under test, or to audit whether your own detectors are brittle. Treat them as an occasional audit, not a gate.

open as a page

You own a recurring garak scan attached to a release pipeline. How would you split the probe selection between probes carrying fixed, checked-in prompt lists and probes that generate their prompts each run, and what do you give up with each choice?

level: principalimportance: nice to knowfreq 28%

basics

~20 s

Gate the pipeline on a small pinned set of fixed-prompt-list probes, because only deterministic prompts give a comparable pass/fail signal per build. Run prompt-generating probes on a slower cadence for fresh coverage, reviewed by a human. Pinned probes go stale and invite tuning; generated ones produce noisy counts that cannot block a build.

open as a page