skip to content

Your team has a fixed number of calls it may spend on garak scans of a product endpoint each release. How do you decide whether any of that budget should go to buffed runs — where a buff rewrites probe prompts before sending — instead of more probes?

level: principalimportance: nice to knowfreq 28%

answer

  1. breadth buys new behaviour classes
  2. buffs buy robustness, not coverage
  3. normalisation layer in the path
  4. audit your own detector stability
  5. triage time is the scarce budget

basics

~20 s

Spend the budget on breadth first: buffs add no new behaviour, they re-send prompts you already have. Buy buffed runs for a narrow slice where input normalisation or a filter is the thing under test, or to audit whether your own detectors are brittle. Treat them as an occasional audit, not a gate.

solid answer

~50 s

Ask what each option can discover. More probes reach behaviour classes you have never tested — a category you do not cover is an unbounded blind spot. A buff reaches no new behaviour class; it tests whether a class you already cover survives a surface change to the prompt. So the default allocation is breadth first, and buffing is a targeted purchase. It earns its place in two cases. First, when a prompt-normalising or filtering layer sits in front of the model and is itself the thing under test — then the transform is the experiment, not decoration. Second, as an audit of your own instrument: if a modest rewrite swings your reported rate, your detectors are brittle and every number you have published rests on that brittleness. And budget the humans, not just the calls. Buffed attempts need the prompt-as-sent read before a hit can be triaged, so a run that multiplies attempts multiplies review time, which is usually the scarcer resource.

go deeper

for a junior

Recognises that buffs cost more calls and do not add new probe material, so breadth usually comes first.

for a middle

Argues coverage of untested behaviour classes over robustness of tested ones, and can price the multiplier.

for a senior

Names the specific cases that justify a buffed arm — a normalisation layer in the path, characterising a confirmed finding — and includes triage time in the budget.

for a principal

Sets the standing policy: unbuffed breadth is the gate, buffed arms are scheduled audits with a stated reference arm, fan-out factor and denominator, and never a headline number alone.

**Frame the choice by what each option can discover.** A fixed per-release call budget is spent on exactly two kinds of thing. Spending on *breadth* — more of garak's probes, selected with `garak --probes` — reaches behaviour classes you have never tested. An untested category is an unbounded blind spot: you do not know how large the risk is, because you have not looked. Spending on *buffs* reaches no new behaviour class at all; a buff re-sends prompts the probes already carried, in altered form, and what it buys is robustness evidence about classes you already test — at a multiplier applied to every one of them. In most programmes the untested categories are the larger unknown, so breadth wins the default allocation and buffing is a targeted purchase that has to argue for itself. **When buffing genuinely earns the money.** - *A normalisation or filtering layer sits in the product path.* If the deployed application lowercases, strips, translates, or screens input before the model sees it, that layer is under contract and plain probes leave it entirely unexercised. Here the transform *is* the experiment, and skipping it means you tested the model and reported on the product. - *You are auditing your own instrument.* A deliberate, scheduled buffed arm answers "how stable is our published number under a benign surface change". If a modest rewrite swings the rate, your detectors are wording-sensitive and every figure you have shipped rests on that brittleness. That is worth real money to learn once a quarter, and it is a finding about the programme, not the target. - *Characterising one confirmed finding.* Take a single failure you already believe and probe it under a transform. That tells you whether the weakness is a wording artefact or the behaviour class itself. It is a small, cheap, narrowly scoped run — the opposite of a default-on sweep. **When it is waste.** Default-on across a full sweep, because the multiplier lands on every selected probe and buys nothing new. As a per-release gate, especially with a stochastic rewriter, because you would be treating paraphrase variance as regression signal and chasing it. As a way to make a report look thorough, because the resulting numbers are harder to defend, not easier. **Price it honestly, including the parts that are not calls.** Suppose the release budget is 60,000 calls. An unbuffed sweep at `garak --generations 5` might use 25,000 of them across a wide probe selection. Adding a single fourfold buff to that same selection costs 100,000 — it does not fit, so in practice you buff a *slice*: pick the handful of probes that touch the layer under test, and the arm costs a few thousand. Then add what the calls do not cover. Triage scales with attempts and costs *more per item* on a buffed arm, because the prompt as sent must be read out of the log before a hit can be judged. A model-backed buff adds rewriting compute, a second dependency that can rate-limit or fail, and a second place your prompts land. On most teams the engineer-hours, not the endpoint bill, are the binding constraint, and a budget that counts only calls will overspend the scarce resource. **Where the number misleads at programme level.** The seductive metric is attempts run — a buffed sweep produces a much bigger figure and reads as a bigger scan. It is not coverage. Coverage is counted in probes and behaviour classes; a sixfold buffed run of the same probes has coverage identical to the unbuffed one. The second trap is release-over-release comparison: once buffs are on, this quarter's per-probe score is not the same statistic as last quarter's, and a trend line drawn through both is fiction. The third is that a *lower* buffed rate reads as robustness when it is at least as likely to be a blinded detector or a rewriter that defanged the prompt. **What I would check, and the policy I would set.** Track marginal discovery per thousand calls for each kind of spend — distinct triaged findings, not raw hits — and let that number, reviewed each quarter, move the split. Track triage hours alongside calls so the human budget is visible. Then: unbuffed breadth is the standing release scan and the gate. Buffed arms are scheduled, narrow, one transform at a time, aimed either at a defence layer genuinely in the path or at auditing detector stability, always reported next to their unbuffed reference arm with the measured fan-out factor and the chosen denominator stated. Nothing from a buffed arm is ever quoted as a headline number on its own.

  • What signal from a buffed audit arm would change how you report ordinary scan results?
    A large swing in the pass/fail rate under a benign rewrite. It says the detectors are wording-sensitive, so every published rate needs an error caveat or better detection before it is quoted.
  • When is spending on a buffed arm clearly better than spending on more probes?
    When a prompt-normalising or filtering layer is in the product path and is the thing you are contracted to test; unbuffed probes leave that layer entirely unexercised.

saying these in an interview costs you the question

  • Turning buffs on by default across a full sweep because it 'tests more'.
  • Making a buffed run the release gate, especially with a stochastic rewriter.
  • Counting buffed variants as additional coverage of behaviour categories.
  • Budgeting only endpoint calls and ignoring the multiplied human triage load.

context