skip to content

Probe Selection

Choosing between a full sweep and a handful of probes trades an hours-long token bill against families you will never be able to mention. Interviewers ask what you left out, and why.

on this pageshow

explore

questions

5

When you start a garak scan against a chat endpoint and do not name any probes, what set of attacks actually runs, and why is that neither nothing nor the whole probe catalogue?

level: juniorimportance: must knowfreq 70%

answer

  1. default set, not full catalogue
  2. module or single class
  3. some probes only run when named
  4. report enumerates what ran
  5. sample, not an assessment

basics

~20 s

garak runs a built-in default selection of its probes, not the entire catalogue. The full catalogue is far larger and would cost hours of metered calls. Some probes also only run when you name them. So a default run is a sample, and you must read the report to see which probes actually ran.

solid answer

~50 s

garak's probes are organised into modules, each holding one or more probe classes, and you can select at either granularity: a whole module, or one class inside it. Give no selection and you get a curated default set, which exists so the tool is useful in minutes rather than hours; it is a starting sample, not an assessment. Two things bite juniors here. First, the catalogue is much larger than the default, so "I ran garak" says almost nothing until you say *which* probes. Second, some probes are flagged so a broad sweep skips them unless you name them explicitly — a family can be absent from your results without any error appearing. The only reliable statement of what ran is the run's own report, which enumerates the probes attempted; list the catalogue first and diff it against that report before writing any coverage sentence.

go deeper

for a junior

Should know that garak selects probes by module or class, that leaving the selection empty runs a default subset rather than everything, and that the report is where you check what ran.

for a middle

Should add that prompt counts vary enormously per probe, that some probes are skipped unless named, and should be able to reconstruct the actual selection from the report.

for a senior

Frames selection as choosing the denominator for the coverage claim, and builds the catalogue-versus-report diff into the engagement's evidence trail.

for a principal

Sets the organisation's policy: which selections constitute a release-gating scan versus a smoke run, who approves a full sweep against a paid endpoint, and how untested families are recorded.

### What a probe is, and how garak's catalogue is shaped garak is an open-source LLM vulnerability scanner. Its unit of attack is a **probe**: a Python class that owns a fixed set of prompts embodying one attack idea, plus a declaration of which **detectors** — the classifiers that decide whether a reply counts as a hit — garak should run over the responses. Probe classes are grouped into modules by family (`garak.probes.<module>`), each module holding one or more classes. garak's selection flag, `--probes` (short form `-p`), takes a comma-separated list at **either** granularity: a bare module name runs the whole family, and `module.ClassName` runs exactly one probe. `garak --list_probes` prints the catalogue for the version you actually have installed, which is the only authoritative statement of what exists — the catalogue grows between releases, so a list you memorised last year is wrong. ### What runs when you name nothing Neither nothing nor everything. With no `--probes` given, garak falls back to a built-in default selection: a curated subset sized so that a first run finishes in minutes and shows you the report format, the per-probe pass and failure rates, and the scoring block. That default exists for **orientation**, not assurance. It is the tool's tutorial mode, and it is smaller than the full catalogue by a large factor. Two further filters sit underneath it and surprise people. First, a garak probe class carries an `active` attribute; a class flagged inactive is skipped when you select its module wholesale and runs only when you name that class explicitly. Second, some probes depend on something the machine may not have — a dataset download, a local support model, a key for a helper service — and will be skipped or will fail at load rather than announcing themselves loudly in the summary. ### What the two extremes cost Running the whole catalogue is the other pole, and it is expensive in a way the flag does not advertise. Prompt counts per probe span orders of magnitude: some probes carry a dozen handwritten prompts, others expand a public dataset into thousands. On top of that, garak's `--generations` flag re-sends every prompt N times and defaults to a small number greater than one, so the request count is already a multiple of the prompt count before you choose anything. Against a hosted, metered chat endpoint a full sweep is realistically tens to hundreds of thousands of requests — hours to overnight of wall-clock, and a bill that lands in real money rather than rounding error. It is a planned exercise with a written estimate, an owner's approval and a spend cap on the API key, never a default. ### Where the number misleads **Absence is silent.** A probe family that never ran does not appear in the report as a warning, an error, or a zero — it simply is not there, and a reader skimming the summary cannot distinguish "we tested encoding attacks and found nothing" from "we never sent an encoding prompt". This is the single most common way a garak result is over-read. **A zero-hit row is not proof of health either.** A probe whose calls all errored — rate-limited, timed out, or hitting a misconfigured generator — can still render a row that looks unalarming. The report shows what the detectors concluded about the attempts that came back, not that the attempts were real. So "we ran garak and it was clean" is not a finding. It is a sentence with no denominator, and it will be read as a clean bill of health by everyone downstream. ### What you would check 1. Run `garak --list_probes` and save that listing with the engagement notes — it pins the catalogue to the installed version. 2. Record the exact `--probes` string you passed, the `--generations` value, and the generator configuration the calls hit (`--model_type` / `--model_name`, the system prompt in force, decoding settings, any guard in front). 3. After the run, diff the probes present in the report against the selection you asked for, and explain every gap: inactive class, missing dependency, or typo in the module name. 4. Spot-read a handful of raw attempts behind one clean probe to confirm the responses arrived non-empty and non-errored. The statement you are then entitled to make is bounded and boring: *these probes, at this repeat count, against this configuration, on this date, produced no detector hits; everything else in the catalogue is untested.* Untested is a different word from safe, and the difference is the whole job.

  • You need to prove which garak probes ran. Where does that evidence come from?
    From the run's own report, which enumerates the probes attempted and their per-probe results — not from the command you believe you typed. Diff the report against the catalogue listing and against your intended selection.
  • Why is 'we ran garak on it' an unacceptable line in a security report?
    It names no denominator. Without the probe selection, the repeat count and the target configuration, the reader cannot tell whether a whole attack family was tested or skipped.
  • Is running the entire catalogue the safe default choice?
    No. Prompt counts per probe differ by orders of magnitude, so a full sweep against a metered endpoint is an hours-long, billable job. It is a planned exercise with an estimate attached, not a default.

saying these in an interview costs you the question

  • Believing the default run covers the whole probe catalogue.
  • Reporting 'garak found nothing' without stating which probes ran.
  • Assuming a missing probe family would have shown up as an error.
  • Treating a full sweep as free or as the obvious default against a hosted paid endpoint.

context

open as a page

You have been asked to scan a metered hosted chat endpoint with garak and to predict the spend before you launch. What multiplies out to the number of requests it will send, and which part of that product does your probe selection control?

level: middleimportance: must knowfreq 60%

basics

~20 s

Requests are roughly the sum, over each probe you selected, of that probe's prompt count times the repeat count per prompt. Your selection controls which probes and therefore the prompt counts, which differ by orders of magnitude between probes. Anything model-backed in the pipeline, such as a hosted detector, adds calls on top.

open as a page

A garak run over a hand-picked subset of probes finishes with nothing flagged. How do you write that up so it is not read as "the model is safe", and what denominator do you attach when you use the word coverage?

level: seniorimportance: must knowfreq 55%

basics

~20 s

Report it as: these named probes, at this repeat count, against this endpoint and configuration, on this date, produced no detector hits. Coverage means probes run divided by probes in the catalogue — not attack surface and not behaviours. Name the families you did not run as untested, never as passed.

open as a page

Your garak sweep against a rate-limited production endpoint keeps dying partway through — throttling responses, timeouts, an expiring credential. How do you restructure the probe selection so that a half-finished sweep still yields results you can report?

level: seniorimportance: should knowfreq 45%

basics

~20 s

Split the sweep into several small runs, each with a named probe selection and its own report, ordered so the families you must report on finish first. Keep the target configuration identical across runs so results stay comparable, log the selection per run, and treat a died-mid-probe result as incomplete, never as a pass.

open as a page

You have a fixed query budget for a garak engagement against a paid endpoint — enough for perhaps a fifth of the probe catalogue. How do you split it between breadth across many probe families and depth within a few, and what do you tell the stakeholder the result covers?

level: principalimportance: should knowfreq 35%

basics

~20 s

Spend a first slice broadly and shallowly across families to find where signal exists, then reinvest the rest depth-first on the families that hit and the ones the deployment's threat model cares about. Tell the stakeholder exactly which families ran, which were never run, and that unrun means untested.

open as a page