skip to content

Custom Probes and Detectors

Writing an attack family the scanner does not ship also means writing whatever decides it landed, and nothing calibrates that decision but you. Interviewers ask where such numbers come from.

on this pageshow

explore

questions

5

You want the garak LLM scanner to cover an attack family it does not ship a probe for. Beyond writing the prompt set itself, what else must you supply before a run can report a failure rate?

level: juniorimportance: must knowfreq 62%

answer

  1. prompts are the cheap half
  2. something must rule each reply
  3. plugin discovery must find it
  4. no judge, no rate
  5. silent zero = not loaded

basics

~20 s

Prompts alone produce no number. You must also supply, or point at, the thing that decides whether each reply counts as a hit, and make both discoverable to the scanner so the run loads them. Without that judging step garak collects replies but has nothing to score.

solid answer

~50 s

A garak run has two halves for every attack family: something that sends prompts to the target, and something that rules on the replies. Ship only the first half and the run produces transcripts, not a rate. So a new family means: the prompt-generating plugin, a stated pairing to whatever will judge its replies, and either an existing shipped judge that genuinely matches your failure signature or a new one you write. You also have to make the code visible to the scanner's plugin discovery and name it on the run, otherwise it is simply never loaded. The trap for a beginner is thinking the prompts are the deliverable. They are the cheap half. The judging half is where the number comes from, and it is the half you now own with no shipped baseline behind it.

go deeper

for a junior

Says a run needs both prompts and something that rules replies, and knows the scanner has to be able to load the new code.

for a middle

Adds that the judge must match the family's actual failure signature, and can debug a silent zero-hit result.

for a senior

Treats the judging half as the deliverable, and will not quote a rate from a custom family without saying how often the judge is wrong.

for a principal

Asks whether the family deserves a permanent plugin at all, given who will keep its judging criterion honest over time.

### The pipeline you are inserting into garak is an LLM vulnerability scanner, and it splits a scan into plugin roles that form a pipeline. A **generator** owns the connection to the target — a hosted model selected with `garak --model_type` and `garak --model_name`, or a class you write yourself when the target is a deployed application rather than a raw model. A **probe** owns one attack family: it holds or mints the prompts, and it declares which detectors should rule its replies. Optional **buffs** rewrite prompts before send. A **detector** receives an `Attempt` — one prompt plus the target's replies to it — and returns one score per reply. An evaluator thresholds those scores into hits and passes, and the harness writes a JSONL report, one line per attempt plus an eval line for each probe/detector pair. Read the pipeline and the answer falls out. Prompts drive sends; detectors produce scores; the evaluator turns scores into a rate. Supply only the prompts and the run produces attempt lines and no eval line — a pile of transcripts, not a failure rate. ### The three things that must be true 1. **The prompts exist, and you know their nature.** A fixed list gives you reruns that cover identical ground; prompts minted at run time do not, so a rate that moves between runs may be sampling rather than the target. 2. **A detector is bound to the probe.** In garak that binding is an attribute on the probe class naming its recommended detector, or a detector named explicitly on the command line with `garak --detectors`. Nothing is inferred from the probe's name or module; an unbound probe is scored by nothing. 3. **garak can import both.** Probes and detectors are resolved by dotted plugin name inside garak's own `garak.probes` and `garak.detectors` namespaces. Code sitting in a directory garak never imports is not broken — it is absent. `garak --list_probes` and `garak --list_detectors` are the one-command check that your class is visible at all. ### What the run costs Sends multiply out as prompts x `garak --generations` x (buffs + 1). A 120-prompt probe at `--generations 10` is 1,200 completions against one target for that one family. Read the `--generations` default off your own `garak --help` rather than trusting memory — it has changed across releases, and it multiplies every other number in the run. At a metered endpoint producing a few hundred output tokens per reply that is a visible invoice line, and at a 60-requests-per-minute limit it is roughly twenty minutes of wall clock; `garak --parallel_attempts` cannot outrun the endpoint's own limiter. The lopsided cost is engineer time: the prompt set is an afternoon, while a detector you can defend — plus the hand-labelled sample that says how often it is wrong — is days. ### Where the number misleads The characteristic failure is the **silent zero**. A probe garak never imported, a probe with no detector bound, and a detector whose criterion cannot fire on this family's failure text produce output that looks the same from the summary: nothing reported for your family. Worse, an absent eval row and a genuine 0% eval row read identically to anyone skimming a report, and both are cheerfully summarised as "we tested that and it was clean." A zero from a custom family is a claim about your plumbing until you have shown otherwise. The second trap is the denominator. garak's rate is computed over outputs, not prompts, so `--generations` sets the denominator. A run at 5 generations and a run at 10 are not comparable for the same probe, and "we re-ran it and the rate dropped" is very often just that flag having changed. ### What I would check Confirm the class appears in `garak --list_probes`. Run the probe against a cheap target with a single prompt and verbose output, and confirm attempt lines actually appear in the JSONL report with non-empty replies. Confirm an eval line exists for your probe/detector pair, with an attempt count matching prompts x generations. Then exercise the detector offline on two hand-written replies you have already labelled — one plainly a failure for your family, one plainly benign — and check that it fires on the first and not the second. A detector that cannot fire on a reply you wrote to make it fire will never fire on a real one.

  • A custom garak family runs to completion and reports zero hits. Name two causes that are not the model being safe.
    The plugin never loaded or was never named on the run, so nothing was sent; or the judge's matching criterion cannot fire on this family's failure text, so hits were sent and never counted.
  • Why does it matter whether your custom prompts are a fixed list or minted at run time?
    A fixed list gives comparable reruns and can be tuned against; generated prompts mean two runs cover different ground, so a change in the rate may be sampling, not the target.

The prompts are the exam paper and the detector is the answer key. Print the exam without a key and you get back a stack of completed papers and no grade — which is exactly what a garak run with an unbound probe hands you.

saying these in an interview costs you the question

  • Thinks writing the prompt list completes the job
  • Cannot say what turns replies into a score
  • Reads a zero-hit custom family as a clean result without checking the plugin loaded
  • Reuses whatever judge is default without checking it matches the family

context

open as a page

You wrote a new probe for the garak LLM scanner and pointed it at a detector that ships with the tool for a different attack family. Why is the failure rate that comes back untrustworthy?

level: middleimportance: must knowfreq 58%

basics

~20 s

A garak detector was written to recognise one family's failure text. Aimed at a different family it misses the real hits and can match unrelated wording, so the rate you get measures the mismatch between probe and detector rather than anything about the target. Both error directions move at once.

open as a page

You wrote both the probe and the detector for a custom family in the garak LLM scanner, and the run reports a 22% failure rate. What do you do before that number goes into a report?

level: seniorimportance: must knowfreq 55%

basics

~20 s

Measure your own detector's error. Sample flagged replies and read them by hand for false positives, and sample unflagged replies for misses. Label a fixed set, ideally with a second reader, and quote the 22% with that measured error attached rather than as a bare number nobody has checked.

open as a page

A garak report can position a probe's result against previously measured reference models as well as giving an absolute score. Why does a probe and detector you wrote yourself not get that comparative reading, and what do you present instead?

level: seniorimportance: should knowfreq 40%

basics

~20 s

The comparative reading needs prior results for that exact probe and detector across reference models. Your plugin has never been run against them, so there is nothing to compare with. Present the absolute rate, your own baseline runs against two or three comparison targets, and the detector's measured error.

open as a page

Your team wants a risk covered that the garak LLM scanner ships no probe for. When would you decline to write a custom probe and detector pair, and what would you do instead?

level: principalimportance: should knowfreq 32%

basics

~20 s

Decline when nobody will own the ruling criterion over time. A custom pair needs re-labelling as the target changes, and a drifting criterion quietly moves the number while looking stable. For a one-off question, test by hand or with a small held-out set; reserve custom plugins for risks you will re-run every release.

open as a page