In the garak LLM scanner, a buff is a plugin that rewrites a probe's prompts on the way out, before they reach the target endpoint. What does enabling a buff do to the amount of work a scan does?
answer
- outbound transform stage
- probes x fan-out x generations
- opt-in, not free coverage
- same prompts, different clothes
- triage cost beats token cost
basics
~20 sA buff sits between the probe and the target and transforms each prompt. Some buffs return one variant per prompt, others fan out into several, so the attempts sent multiply on top of the per-prompt generation count. Enabling buffs makes a scan longer and more expensive, never cheaper.
solid answer
~50 sA garak scan is a product of counts: probes selected, prompts per probe, generations per prompt. A buff inserts a transform stage on the outbound side, and the transformed prompts are what actually get sent. A one-to-one buff (case folding, a mechanical rewrite) leaves the attempt count alone but changes every prompt; a fan-out buff (paraphrasing, translating into several languages) turns one prompt into several attempts, so the whole run multiplies by that fan-out factor. The practical consequences: a default run applies no rewriting, so the baseline is the unbuffed run; adding buffs is opt-in spend, not free coverage. And you are not adding new probe material — you are re-sending prompts you already had, in different clothes. If a buff is itself model-backed, you also pay for the rewriting calls, not just the target calls.
go deeper
Says a buff rewrites prompts before they are sent and that turning one on makes the scan bigger and slower, not smaller.
Gives the multiplication explicitly (probes x variants per prompt x generations) and notes that a default run does no rewriting.
Adds that the dominant cost is triage of hits whose prompt no longer matches the probe's canonical text, and insists the run log carries the prompt as sent.
Frames it as budget allocation: the multiplier lands on every selected probe, so buffing is a deliberate slice of spend traded against breadth.
**Where the stage sits.** A garak run is a pipeline of four stages. A *probe* supplies prompt text. A *buff* may rewrite that text on the way out. A *generator* sends the result to the target endpoint and collects replies. A *detector* reads each reply and rules whether the attempt counts as a failure. The buff is the only stage that changes what the target actually received while leaving the probe's name on the report. You switch one on with garak's `--buffs` argument (or the equivalent `buffs` entry in a garak run config file), and `garak --list_buffs` prints what your installation ships. A default `garak` invocation enables none of them: the unbuffed run is the baseline, and buffing is opt-in spend. **The arithmetic.** Calls to the target are a product, not a sum: ``` attempts = (prompts contributed by the selected probes) x (variants each enabled buff emits per prompt) x (garak --generations) ``` `garak --generations` (short form `-g`) is the number of completions requested per prompt, and it already defaults to a value greater than one, so there is a multiplier in the run before any buff exists. Buffs come in two shapes. A **one-to-one** buff — garak's lowercase buff is the canonical example — rewrites every prompt and leaves the attempt count exactly where it was. A **fan-out** buff — the paraphrase buffs that emit several rewordings of each prompt, or a buff that renders each prompt into several other languages — multiplies the whole run by its fan-out factor. Two fan-out buffs enabled at once compound, they do not add. **What it costs.** Take a modest sweep: ten probes contributing roughly a hundred prompts each, with `garak --generations 5`. That is about 5,000 calls unbuffed. Enable a paraphrase buff that emits six variants and it is about 30,000. Money scales linearly with those calls and with reply length, so a hosted-model bill goes up sixfold. Wall clock scales too: against an endpoint that comfortably serves two requests a second, 30,000 calls is over four hours, and a rate-limited product endpoint — which is the normal case, since you are scanning a deployed app, not a raw model — can stretch that into a day. If the buff is model-backed you also pay for the rewriting itself: a local paraphraser wants a model download and GPU or CPU time of its own, a hosted one is a second bill and a second dependency. **The cost people underestimate is human.** Every hit still has to be triaged. A buffed hit costs strictly more per item than an unbuffed one, because the prompt that produced it is no longer the probe's canonical wording — you have to read the prompt *as sent* out of the run log before you can judge whether the reply is really a failure. A run that multiplies attempts by six can more than multiply review hours by six. **Where the number misleads.** Three specific readings go wrong. First, *attempts are not coverage*: a report showing 30,000 attempts looks more thorough than one showing 5,000, but the buffed run reached no behaviour the probes did not already carry — it re-sent the same ideas in different clothes. Coverage is counted in probes and behaviour classes, never in attempts. Second, *the per-probe rate quietly changes denominator*: the same probe now scores over buffed attempts, so its figure is not comparable to the same probe's figure from last release's unbuffed run. Third, *the fan-out factor you assumed may not be the one you got*: a paraphraser can return fewer usable variants than requested, or drop a prompt it declines to rewrite, so the multiplier is empirical, not documented. **What to check before launching the big run.** Measure the fan-out rather than assuming it: run one small probe with `garak --generations 1` and the buff enabled, then count attempt records in the run's `.report.jsonl` and divide by that probe's prompt count. Multiply the measured factor through the full sweep and price it — calls, money and hours — before you start. Confirm the run log records the prompt as sent, not merely the probe's original, or the buffed hits will be unreviewable. And keep a matching unbuffed run, because a buffed number is uninterpretable without it.
- Where in a garak run does the buff transform happen relative to the detector?Strictly before the send: the buff changes the outbound prompt, and the detector then reads whatever reply that transformed prompt produced. The detector never sees the original prompt's wording.
- If you enable two fan-out buffs at once, what happens to the run size?The factors compound, so the attempt count multiplies twice over. That is usually a reason to run them as separate arms rather than together, so you can attribute a result to one transform.
Enabling a fan-out buff is like sending a proofreader you pay by the page six translations of every page instead of one: nothing new has been said, but the invoice and the reading time both multiply by six.
saying these in an interview costs you the question
- Believing buffs add new attack material rather than re-sending existing prompts.
- Assuming a buff is one-to-one and quoting an attempt count that ignores fan-out.
- Enabling several buffs across a full probe sweep without estimating the multiplied call count.
- Not knowing that a default scan applies no prompt rewriting at all.