skip to content

Buffs

A transform on the way out changes what the detector later sees, so a buffed run can score high or low for reasons that are not the model. Interviewers ask because it is the easiest number to fake.

on this pageshow

explore

questions

5

You run a set of garak probes against an endpoint twice: once plain, once with a buff enabled that rewrites the probes' prompts before they are sent. The reported pass/fail rate is very different. Why can that difference come from garak's detectors rather than from the model's behaviour?

level: middleimportance: must knowfreq 60%

answer

  1. detector sees only the reply
  2. buff moves the reply distribution
  3. hidden refusal = inflated rate
  4. hidden payload = deflated rate
  5. hand re-judge the changed verdicts

basics

~20 s

Detectors judge the reply text, and many are substring or pattern matchers. A buff changes the prompt, so the reply's wording, language or format changes too. A matcher tuned for the original wording then misses real hits or fires on harmless replies. The number moved because detection moved, not the model.

solid answer

~50 s

The buff is upstream of everything: it changes what the target receives, so it changes the shape of the reply, and the detector sees only the reply. Detectors for a given probe are written against the replies that probe's original prompts tend to produce — a refusal phrase list, a keyword for the payload, a small classifier trained on one language and register. So the error runs in both directions. A buff that shifts language or register can hide a refusal from a refusal-matching detector, and the attempt is then scored as a failure the model never actually committed — the rate goes **up** for nothing. The same shift can hide the payload from a keyword-matching detector, so real failures are scored clean and the rate goes **down** for nothing. Either way the delta is a property of the measurement, not the target. The only way to tell is to read a sample of buffed attempts and their replies by hand and re-judge them yourself.

go deeper

for a junior

Says the buff changed the prompt so the reply changed, and the detector may no longer recognise what it is looking for.

for a middle

Names both error directions and ties each to the kind of matcher involved: missed refusal inflates, missed payload deflates.

for a senior

Insists on a paired unbuffed arm and hand re-judgement of verdict-flipped attempts, and separates detector error from the buff mangling the payload.

for a principal

Treats it as a measurement-validity question: a transform that invalidates the detector makes the metric unreportable, and the team needs a rule for when a buffed number may be quoted at all.

**What the detector can and cannot see.** In garak, the pass/fail decision belongs to the *detector* attached to the probe, and the detector is handed only the reply text. It never sees the probe's original wording, and it has no idea a buff ran. Detectors are cheap on purpose so that a scan of tens of thousands of attempts finishes: many are substring or regular-expression matchers over the reply, and the heavier ones are small classifiers. Cheap matchers are implicitly tuned to a *distribution* of replies — the replies that this probe's prompts, as written, usually elicit, in the language and register they usually come back in. A buff deliberately moves the prompts off that distribution, and every judgement downstream inherits the move. **The two directions, concretely.** garak's detectors split into two logical polarities and the transform breaks each one differently. - Detectors that look for **refusal or mitigation language** score an attempt as a hit when they *fail* to find that language — garak's MitigationBypass detector is the family example. Rewrite the prompt so the reply comes back in another language, another register, or a different format, and a perfectly genuine refusal goes unrecognised. The measurement manufactures failures the model never committed, and the reported rate goes **up** for nothing. - Detectors that look for the **presence** of the thing the probe was fishing for — a keyword, a pattern, a classifier's positive class — score a hit when they *do* find it. Rewrite the prompt so a reply carries the same substance in different words or another language, and real failures slip past unrecognised. Failures are suppressed and the rate goes **down** for nothing. A single buffed run can do both at once on different probes, which is exactly why the aggregate delta can point either way and, on its own, tells you nothing about the target. **A third, quieter cause that is not detector error.** If the buff is model-backed, the rewriting model can soften, mangle or decline the request while paraphrasing. Those attempts left carrying less than the probe intended, came back clean, and landed in the same aggregate. It is a false negative that looks identical to a well-behaved target. **Where the report's own numbers mislead.** garak's report gives a per-probe absolute score and, alongside it, a calibration figure positioning the target against a bag of reference models. Those reference runs were scanned unbuffed. A buffed arm compared to that calibration is comparing two different measurements and calling the difference a property of the model; the Z-score is not portable across a transform. The same applies release over release: last quarter's unbuffed score for a probe and this quarter's buffed score for the same probe are not the same statistic, however identical the probe name looks in the table. **What it costs to settle.** Untangling this is human work, not compute. The minimum honest exercise is a paired unbuffed reference run (which doubles or worse the call budget, since the buffed arm was already multiplied by fan-out) plus a hand-judged sample. A stratified sample of, say, 150 to 300 verdict-flipped attempts read at roughly a minute each is half a day of an engineer who knows the domain. That cost is the reason people quote the delta instead — and it is the reason the delta is so often wrong. **What you actually check.** Take the attempts whose verdict differs between the two arms, sample them stratified by probe, and read the prompt *as sent* together with the reply out of the run's `.report.jsonl`. Re-judge each by hand and compare your labels with the detector's. You are estimating the detector's error rate separately in each arm; if those error rates differ materially, the arms are not comparable and no rate you compute from them means anything. Also look at whether the replies changed *form* at all — if the buffed replies come back in the same language and register, detector drift is a weaker explanation and a genuine behaviour change becomes more credible. When the sample shows the detector mis-ruling, the honest write-up is "detection is not valid under this transform", stated as a finding about the instrument, not a number about the model. When the sample holds up, you have a real result, and the transform is precisely what makes it interesting.

  • Which direction does a buff that shifts the reply's language usually push a refusal-matching detector?
    Towards scoring more failures: the refusal is present but unrecognised, so the attempt looks like the model complied when it did not.
  • What is the minimum evidence you would want before reporting that a buff changed the target's behaviour?
    A paired unbuffed run over the same probes and endpoint, plus hand re-judgement of a sample of the attempts whose verdict differs between the two arms.
  • Can a buff move the number without any detector error at all?
    Yes — a model-backed buff can soften or drop the payload while rewriting, so some attempts simply never carried the probe's behaviour.

It is like marking translated exam papers against an answer key still written in the original language: the marks move sharply, and none of the movement is about the students.

saying these in an interview costs you the question

  • Reporting a buffed pass/fail rate as evidence about the model with no unbuffed reference arm.
  • Assuming the error can only run one way (only false negatives, or only false positives).
  • Treating a lower buffed rate as the model being safer under obfuscation.
  • Never reading a single buffed prompt-and-reply pair before quoting the aggregate.

context

open as a page

In the garak LLM scanner, a buff is a plugin that rewrites a probe's prompts on the way out, before they reach the target endpoint. What does enabling a buff do to the amount of work a scan does?

level: juniorimportance: should knowfreq 45%

basics

~20 s

A buff sits between the probe and the target and transforms each prompt. Some buffs return one variant per prompt, others fan out into several, so the attempts sent multiply on top of the per-prompt generation count. Enabling buffs makes a scan longer and more expensive, never cheaper.

open as a page

You are asked to report whether rewriting prompts changes how often a target fails a chosen set of garak probes. How do you set up the buffed and unbuffed garak runs so the comparison actually means something, and what is the trap in the rate itself?

level: seniorimportance: should knowfreq 38%

basics

~20 s

Run the same probes twice against the same endpoint, changing only the buff. Hold generations, detectors and endpoint settings fixed, and run them close together. Then compare like for like: a fan-out buff sends several variants per original prompt, so decide whether your rate counts attempts or original prompts.

open as a page

Some of garak's buffs rewrite a probe's outbound prompts by calling a language model to paraphrase or translate them. What does that add to a scan you have to defend, and how do you handle it?

level: seniorimportance: should knowfreq 40%

basics

~20 s

It makes the run non-reproducible: a rerun paraphrases differently, so the prompts sent are not the same set. The paraphraser can also soften or refuse the payload, silently weakening the probe. Keep the unbuffed run as the baseline, read a sample of transformed prompts, and cite the logged prompts, not the buff name.

open as a page

Your team has a fixed number of calls it may spend on garak scans of a product endpoint each release. How do you decide whether any of that budget should go to buffed runs — where a buff rewrites probe prompts before sending — instead of more probes?

level: principalimportance: nice to knowfreq 28%

basics

~20 s

Spend the budget on breadth first: buffs add no new behaviour, they re-send prompts you already have. Buy buffed runs for a narrow slice where input normalisation or a filter is the thing under test, or to audit whether your own detectors are brittle. Treat them as an occasional audit, not a gate.

open as a page