skip to content

In the garak LLM scanner, some probes send a fixed prompt list shipped inside the package while others mint their prompts during the run. How does that difference change what a passing result for a probe attests to?

level: middleimportance: must knowfreq 60%

answer

  1. public list = tunable, reproducible
  2. generated = fresh, not repeatable
  3. contamination arrives via the probe
  4. persist what was actually sent
  5. report the two families separately

basics

~20 s

A fixed-list probe sends the same readable strings every time, so a pass means only those exact prompts failed to elicit the behaviour, and anyone can tune against them. A prompt-generating probe sends different prompts each run, so a pass is a sample you cannot enumerate afterwards and a rerun tests different ground.

solid answer

~50 s

The two kinds attest to different things. **Fixed list.** The prompts live in the package, so they are public, enumerable and reproducible. That makes the result comparable release to release and easy to explain in a report. It also makes the list a target: anything that can be string-matched or fine-tuned away turns the probe green without changing the behaviour, and the list ages as defences absorb it. **Generated at run time.** The probe builds prompts during the scan, often by mutating or expanding a seed, so no one can pre-block them and each run explores fresh ground. The costs are the mirror image: results are not reproducible by default, counts wobble between runs, and "which prompt failed" is only answerable if you persist the run's log. The practical rule is to know which kind you sent before quoting the number, and to keep the generated prompts from any run that produced a finding.

go deeper

for a junior

Knows some garak probes carry prompts in the package and others build them during the run, and that only the first kind repeats exactly.

for a middle

Explains the tradeoff both ways: determinism and tunability versus freshness and variance, and says a pass means different things in each case.

for a senior

Reports the two families separately, treats a fixed-list pass as a floor, and requires generated-prompt hits to be persisted and replayed before they become findings.

for a principal

Owns the policy for which family gates a release and which produces exploratory findings, and expects industry-wide drift on public lists to be discounted.

**Why this is the first question to ask about any garak result.** A garak probe is a prompt source, and prompt sources come in two shapes with opposite failure modes. A *static* probe carries its prompts as data — a Python list on the class, or a resource file shipped inside the package that the class loads at start-up. A *generative* probe builds prompts while the scan runs: garak's `atkgen` drives a separate red-teaming model that composes the next prompt turn by turn from the target's own replies, and the tree-search style probes (the TAP family) run an attacker model plus a judge model over a branching search, keeping the promising branches and pruning the rest. Reading a report without knowing which shape you sent means you do not know what the number is evidence *of*. | | static list | generated at run time | |---|---|---| | prompts known before the run | yes, and readable by anyone | no | | two runs comparable | yes, same strings in the same order | no, two samples | | cost knowable in advance | yes: `prompts x --generations` | no: set by turn/width/depth parameters | | extra machinery needed | none | a second model (API spend or a local GPU) | | can be tuned against | yes — the strings are public | not in advance | **What each one costs.** A static probe's bill is arithmetic you can do before you start: prompt count times `garak --generations`, one generator call per attempt. A generative probe has no such number. Each conversational attempt spends several target calls *plus* an attack-model call per turn; a tree search spends an attacker call, a judge call and a target call at every node it expands, so the run's bill is on the order of three times the node count, and the node count is set by width and depth parameters rather than by anything you can count in advance. Practically this means a generative family can cost an order of magnitude more per *finding* than a static one, that its wall-clock is dominated by the slowest of three services, and that a rate-limited production endpoint can stretch a nominally small run across hours. Budget it as an open-ended job with a hard stop, not as a list you are working through. **Where the numbers mislead — in opposite directions.** *Static lists suffer contamination.* The strings are published in the package and mirrored all over the internet, so they end up in training corpora and in filter blocklists alike. A defence that memorises or pattern-matches those exact prompts scores a clean pass without generalising one paraphrase away, and nothing in the report distinguishes that from real robustness. The same effect drifts pass rates upward across the whole industry over time for reasons that have nothing to do with any one system — a public-list pass is a floor ("the well-known strings are handled"), never a robustness claim. *Generative probes suffer variance.* Two runs are two samples, so a count that moved may be pure noise. Worse, the sample *distribution* can shift under you without any local change: a different version of the attack model, a different judge, a changed temperature, and the instrument itself is now different. A hit is stronger evidence than a static hit — nobody could have tuned against a prompt that did not exist before the run started — but a pass is only "this sample found nothing". **What you check.** For a static probe, confirm the prompt file's provenance and treat the pass as known-ground regression. For a generative probe, persist what was actually sent: the report's attempt records carry the prompt text and the replies, and without them a hit cannot be reproduced, handed to the owning team, or re-tested after the fix. If your build exposes a seed setting, pin it — then verify it by diffing the prompts of two runs, and do not assume a seed makes an API-backed target deterministic, because the target's own sampling still moves. Report the two families separately rather than aggregating them into one score, and say explicitly in a finding when a prompt was generated, or a reader will assume a reproducibility the run never had.

  • Why can a system score a clean pass on a public, fixed-prompt garak probe without being any safer?
    The exact strings are public, so a filter or a fine-tune can neutralise them specifically. The probe measures those strings, and they are now handled; the underlying behaviour may be one paraphrase away.
  • What do you have to persist from a run of a prompt-generating probe, and why?
    The prompts that were actually sent and the replies. Without them a hit cannot be reproduced, cannot be handed to the owning team, and cannot be re-tested after the fix.

A fixed prompt list is a published driving-test route: the school drills it, everyone eventually passes, and passing stops meaning much. A generative probe is an examiner picking streets on the day — a fairer test, but two candidates' scores are not comparable, because they drove different streets.

saying these in an interview costs you the question

  • Treats every garak probe as reproducible by default.
  • Aggregates fixed-list and generated-prompt probes into one headline score without saying so.
  • Reads a fixed-list pass as evidence the behaviour cannot be elicited.
  • Files a finding from a generated prompt without saving the prompt that produced it.
  • Explains a moved count between two generated-prompt runs as a fix without rerunning.

context