You want the garak LLM scanner to cover an attack family it does not ship a probe for. Beyond writing the prompt set itself, what else must you supply before a run can report a failure rate?
answer
- prompts are the cheap half
- something must rule each reply
- plugin discovery must find it
- no judge, no rate
- silent zero = not loaded
basics
~20 sPrompts alone produce no number. You must also supply, or point at, the thing that decides whether each reply counts as a hit, and make both discoverable to the scanner so the run loads them. Without that judging step garak collects replies but has nothing to score.
solid answer
~50 sA garak run has two halves for every attack family: something that sends prompts to the target, and something that rules on the replies. Ship only the first half and the run produces transcripts, not a rate. So a new family means: the prompt-generating plugin, a stated pairing to whatever will judge its replies, and either an existing shipped judge that genuinely matches your failure signature or a new one you write. You also have to make the code visible to the scanner's plugin discovery and name it on the run, otherwise it is simply never loaded. The trap for a beginner is thinking the prompts are the deliverable. They are the cheap half. The judging half is where the number comes from, and it is the half you now own with no shipped baseline behind it.
go deeper
Says a run needs both prompts and something that rules replies, and knows the scanner has to be able to load the new code.
Adds that the judge must match the family's actual failure signature, and can debug a silent zero-hit result.
Treats the judging half as the deliverable, and will not quote a rate from a custom family without saying how often the judge is wrong.
Asks whether the family deserves a permanent plugin at all, given who will keep its judging criterion honest over time.
### The pipeline you are inserting into garak is an LLM vulnerability scanner, and it splits a scan into plugin roles that form a pipeline. A **generator** owns the connection to the target — a hosted model selected with `garak --model_type` and `garak --model_name`, or a class you write yourself when the target is a deployed application rather than a raw model. A **probe** owns one attack family: it holds or mints the prompts, and it declares which detectors should rule its replies. Optional **buffs** rewrite prompts before send. A **detector** receives an `Attempt` — one prompt plus the target's replies to it — and returns one score per reply. An evaluator thresholds those scores into hits and passes, and the harness writes a JSONL report, one line per attempt plus an eval line for each probe/detector pair. Read the pipeline and the answer falls out. Prompts drive sends; detectors produce scores; the evaluator turns scores into a rate. Supply only the prompts and the run produces attempt lines and no eval line — a pile of transcripts, not a failure rate. ### The three things that must be true 1. **The prompts exist, and you know their nature.** A fixed list gives you reruns that cover identical ground; prompts minted at run time do not, so a rate that moves between runs may be sampling rather than the target. 2. **A detector is bound to the probe.** In garak that binding is an attribute on the probe class naming its recommended detector, or a detector named explicitly on the command line with `garak --detectors`. Nothing is inferred from the probe's name or module; an unbound probe is scored by nothing. 3. **garak can import both.** Probes and detectors are resolved by dotted plugin name inside garak's own `garak.probes` and `garak.detectors` namespaces. Code sitting in a directory garak never imports is not broken — it is absent. `garak --list_probes` and `garak --list_detectors` are the one-command check that your class is visible at all. ### What the run costs Sends multiply out as prompts x `garak --generations` x (buffs + 1). A 120-prompt probe at `--generations 10` is 1,200 completions against one target for that one family. Read the `--generations` default off your own `garak --help` rather than trusting memory — it has changed across releases, and it multiplies every other number in the run. At a metered endpoint producing a few hundred output tokens per reply that is a visible invoice line, and at a 60-requests-per-minute limit it is roughly twenty minutes of wall clock; `garak --parallel_attempts` cannot outrun the endpoint's own limiter. The lopsided cost is engineer time: the prompt set is an afternoon, while a detector you can defend — plus the hand-labelled sample that says how often it is wrong — is days. ### Where the number misleads The characteristic failure is the **silent zero**. A probe garak never imported, a probe with no detector bound, and a detector whose criterion cannot fire on this family's failure text produce output that looks the same from the summary: nothing reported for your family. Worse, an absent eval row and a genuine 0% eval row read identically to anyone skimming a report, and both are cheerfully summarised as "we tested that and it was clean." A zero from a custom family is a claim about your plumbing until you have shown otherwise. The second trap is the denominator. garak's rate is computed over outputs, not prompts, so `--generations` sets the denominator. A run at 5 generations and a run at 10 are not comparable for the same probe, and "we re-ran it and the rate dropped" is very often just that flag having changed. ### What I would check Confirm the class appears in `garak --list_probes`. Run the probe against a cheap target with a single prompt and verbose output, and confirm attempt lines actually appear in the JSONL report with non-empty replies. Confirm an eval line exists for your probe/detector pair, with an attempt count matching prompts x generations. Then exercise the detector offline on two hand-written replies you have already labelled — one plainly a failure for your family, one plainly benign — and check that it fires on the first and not the second. A detector that cannot fire on a reply you wrote to make it fire will never fire on a real one.
- A custom garak family runs to completion and reports zero hits. Name two causes that are not the model being safe.The plugin never loaded or was never named on the run, so nothing was sent; or the judge's matching criterion cannot fire on this family's failure text, so hits were sent and never counted.
- Why does it matter whether your custom prompts are a fixed list or minted at run time?A fixed list gives comparable reruns and can be tuned against; generated prompts mean two runs cover different ground, so a change in the rate may be sampling, not the target.
The prompts are the exam paper and the detector is the answer key. Print the exam without a key and you get back a stack of completed papers and no grade — which is exactly what a garak run with an unbound probe hands you.
saying these in an interview costs you the question
- Thinks writing the prompt list completes the job
- Cannot say what turns replies into a score
- Reads a zero-hit custom family as a clean result without checking the plugin loaded
- Reuses whatever judge is default without checking it matches the family