skip to content

In the garak LLM scanner, what does a probe supply to a scan, and what does a clean result across the probes you selected say about the probes you did not run?

level: juniorimportance: must knowfreq 70%

answer

  1. probe = the prompt source
  2. selection is the denominator
  3. probes run / probes available
  4. per-probe lines, not one headline
  5. record probe list with the finding

basics

~20 s

A garak probe is the plugin that supplies the prompts sent to the target for one attack family. A clean result covers only the probes you selected; it says nothing about families you never ran. Coverage means probes run out of probes available, and the default selection is not everything.

solid answer

~40 s

A garak probe is the prompt source. Each probe holds or produces the prompts for one family of attacks, sends them through the configured generator, and hands the replies to a detector that decides which attempts count as hits. **You choose the probe selection, and that choice is the denominator of everything the run reports.** So "we ran garak and it passed" is not a claim until you name the probes. Two teams can both say it and have tested disjoint ground. The honest form is "these named probes, this many attempts each, against this endpoint, on this date". Note also that a probe pass is bounded by the prompts that probe actually sent — it is evidence about those attempts, not proof the underlying behaviour is unreachable by any prompt.

go deeper

for a junior

Says a probe carries the prompts for one attack family and that a scan only tests the probes you chose to run.

for a middle

Separates the three denominators — probes run, attempts per probe, attack surface — and knows garak can only measure the first.

for a senior

Insists any quoted result names the probe selection, attempt counts and package version, and keeps that selection under version control so reruns compare.

for a principal

Sets the org rule that a release claim cites probe selection and per-probe lines, and treats an unqualified 'garak was clean' as an unreviewable assertion.

**What a probe is, mechanically.** garak is an open-source scanner that probes an LLM endpoint for known failure families. Every run is assembled from four plugin roles: a **probe** supplies the prompts, a **generator** carries each prompt to the system under test and returns the reply, a **detector** rules whether a reply counts as a hit, and an optional **buff** transforms prompts before they are sent. The probe is where the content of the test lives. Concretely it is a Python class at `garak.probes.<module>.<ClassName>` — `garak.probes.dan.DanInTheWild`, say — carrying the prompt material for one attack family, a `goal` string describing what it is trying to make the target do, a recommended detector, and tags. You select probes with `garak --probes <module>` (every *active* class in that module) or `garak --probes <module>.<ClassName>` (one class). `garak --list_probes` prints the installed catalogue. **The arithmetic, because it bites first.** One run issues roughly ``` calls ≈ Σ over selected probes ( prompts in probe × garak --generations ) ``` `--generations` is how many completions garak requests per prompt, so a stochastic target gets sampled more than once; its default has moved between releases, so read `garak --help` on the version you actually installed rather than a remembered number. The consequence is that probe families are wildly unequal in cost. A family that machine-generates one payload variant per encoding scheme ships prompts by the thousand; a curated jailbreak family may ship a few dozen. Select both and the large one is ~99% of the bill, the wall-clock, and the aggregate score's weight. At roughly 800 tokens round-trip per call, 20,000 attempts is ~16M tokens — real money on a metered API, plus hours of wall-clock once the endpoint throttles you. `garak --parallel_attempts` raises concurrency, but the endpoint's rate limit, not the flag, sets the ceiling. **Three denominators people quietly swap.** | denominator | who can compute it | what it is worth | |---|---|---| | probes executed / probes installed | garak | the only one the tool measures — and it moves as the catalogue grows | | attempts inside one probe | you, from the report | one report line per probe whether it sent 12 prompts or 12,000 | | attack surface reached | nobody, from this run | garak sees one endpoint; the app is not the endpoint | Note also that `--probes all` is not literally all: probe classes carry an `active` flag, and inactive ones (slow, experimental, or dependency-heavy) are skipped unless you name them explicitly. **Where the number misleads.** A single headline pass rate is weighted by prompt counts nobody chose, so shifting one bulky family in or out moves the headline without anything about the target changing. A per-probe *pass* is the detector's silence, not the model's virtue — a detector that never fires on this product's reply shape produces the same green as a genuinely robust system. garak's calibration figure positions a result against other systems the project has measured; it is a relative position, so "better than the reference bag" is compatible with a high absolute failure rate. And every pass is bounded by the prompts that were actually sent: it is evidence about those attempts, never proof the behaviour is unreachable by some other phrasing. **What you check before quoting it.** Read the machine-readable report (`garak --report_prefix <name>` writes a `.report.jsonl` alongside the HTML). Its setup record carries the model, the probe specification and the generations count, so you can confirm what was requested rather than what someone remembers requesting. Then count the *distinct probe names that actually appear in the attempt records* — probes can fail to load (a missing dependency, an absent API key for a probe that needs a second model) and the run still completes and still prints a summary. Record the package version, keep the probe selection in a config file under version control rather than in shell history, and attach probe names, per-probe attempt counts, endpoint, version and date to any finding. The failure this prevents is ordinary: a team defends a release on "garak was clean", a reviewer later finds an untested family, and the tool turns out not to have lied — nobody wrote down the denominator.

  • Two engineers each report a clean garak scan of the same chat product but disagree about its safety. What is the first thing you ask for?
    The probe selection and attempt counts for each run. They very likely ran different probes, so the two clean results are about different ground.
  • Why is 'probes run out of probes available' still a weak coverage number even when you report it honestly?
    Because the denominator is the installed catalogue, not the system's attack surface, and it shifts when the package's probe set changes. It measures how much of the tool you used, not how much of the app you tested.

Saying a garak run was clean without naming the probes is like saying a student passed the exam without saying which questions were on the paper. The pass is real; the paper is what you actually learned about.

saying these in an interview costs you the question

  • Says 'we ran garak, it was clean' without naming which probes ran.
  • Treats the default probe selection as full coverage.
  • Quotes a single aggregate number and cannot say which probes contributed to it.
  • Claims a probe pass proves the behaviour cannot be elicited at all.
  • Confuses probes-run coverage with coverage of the application's attack surface.

context