skip to content

You run a set of garak probes against an endpoint twice: once plain, once with a buff enabled that rewrites the probes' prompts before they are sent. The reported pass/fail rate is very different. Why can that difference come from garak's detectors rather than from the model's behaviour?

level: middleimportance: must knowfreq 60%

answer

  1. detector sees only the reply
  2. buff moves the reply distribution
  3. hidden refusal = inflated rate
  4. hidden payload = deflated rate
  5. hand re-judge the changed verdicts

basics

~20 s

Detectors judge the reply text, and many are substring or pattern matchers. A buff changes the prompt, so the reply's wording, language or format changes too. A matcher tuned for the original wording then misses real hits or fires on harmless replies. The number moved because detection moved, not the model.

solid answer

~50 s

The buff is upstream of everything: it changes what the target receives, so it changes the shape of the reply, and the detector sees only the reply. Detectors for a given probe are written against the replies that probe's original prompts tend to produce — a refusal phrase list, a keyword for the payload, a small classifier trained on one language and register. So the error runs in both directions. A buff that shifts language or register can hide a refusal from a refusal-matching detector, and the attempt is then scored as a failure the model never actually committed — the rate goes **up** for nothing. The same shift can hide the payload from a keyword-matching detector, so real failures are scored clean and the rate goes **down** for nothing. Either way the delta is a property of the measurement, not the target. The only way to tell is to read a sample of buffed attempts and their replies by hand and re-judge them yourself.

go deeper

for a junior

Says the buff changed the prompt so the reply changed, and the detector may no longer recognise what it is looking for.

for a middle

Names both error directions and ties each to the kind of matcher involved: missed refusal inflates, missed payload deflates.

for a senior

Insists on a paired unbuffed arm and hand re-judgement of verdict-flipped attempts, and separates detector error from the buff mangling the payload.

for a principal

Treats it as a measurement-validity question: a transform that invalidates the detector makes the metric unreportable, and the team needs a rule for when a buffed number may be quoted at all.

**What the detector can and cannot see.** In garak, the pass/fail decision belongs to the *detector* attached to the probe, and the detector is handed only the reply text. It never sees the probe's original wording, and it has no idea a buff ran. Detectors are cheap on purpose so that a scan of tens of thousands of attempts finishes: many are substring or regular-expression matchers over the reply, and the heavier ones are small classifiers. Cheap matchers are implicitly tuned to a *distribution* of replies — the replies that this probe's prompts, as written, usually elicit, in the language and register they usually come back in. A buff deliberately moves the prompts off that distribution, and every judgement downstream inherits the move. **The two directions, concretely.** garak's detectors split into two logical polarities and the transform breaks each one differently. - Detectors that look for **refusal or mitigation language** score an attempt as a hit when they *fail* to find that language — garak's MitigationBypass detector is the family example. Rewrite the prompt so the reply comes back in another language, another register, or a different format, and a perfectly genuine refusal goes unrecognised. The measurement manufactures failures the model never committed, and the reported rate goes **up** for nothing. - Detectors that look for the **presence** of the thing the probe was fishing for — a keyword, a pattern, a classifier's positive class — score a hit when they *do* find it. Rewrite the prompt so a reply carries the same substance in different words or another language, and real failures slip past unrecognised. Failures are suppressed and the rate goes **down** for nothing. A single buffed run can do both at once on different probes, which is exactly why the aggregate delta can point either way and, on its own, tells you nothing about the target. **A third, quieter cause that is not detector error.** If the buff is model-backed, the rewriting model can soften, mangle or decline the request while paraphrasing. Those attempts left carrying less than the probe intended, came back clean, and landed in the same aggregate. It is a false negative that looks identical to a well-behaved target. **Where the report's own numbers mislead.** garak's report gives a per-probe absolute score and, alongside it, a calibration figure positioning the target against a bag of reference models. Those reference runs were scanned unbuffed. A buffed arm compared to that calibration is comparing two different measurements and calling the difference a property of the model; the Z-score is not portable across a transform. The same applies release over release: last quarter's unbuffed score for a probe and this quarter's buffed score for the same probe are not the same statistic, however identical the probe name looks in the table. **What it costs to settle.** Untangling this is human work, not compute. The minimum honest exercise is a paired unbuffed reference run (which doubles or worse the call budget, since the buffed arm was already multiplied by fan-out) plus a hand-judged sample. A stratified sample of, say, 150 to 300 verdict-flipped attempts read at roughly a minute each is half a day of an engineer who knows the domain. That cost is the reason people quote the delta instead — and it is the reason the delta is so often wrong. **What you actually check.** Take the attempts whose verdict differs between the two arms, sample them stratified by probe, and read the prompt *as sent* together with the reply out of the run's `.report.jsonl`. Re-judge each by hand and compare your labels with the detector's. You are estimating the detector's error rate separately in each arm; if those error rates differ materially, the arms are not comparable and no rate you compute from them means anything. Also look at whether the replies changed *form* at all — if the buffed replies come back in the same language and register, detector drift is a weaker explanation and a genuine behaviour change becomes more credible. When the sample shows the detector mis-ruling, the honest write-up is "detection is not valid under this transform", stated as a finding about the instrument, not a number about the model. When the sample holds up, you have a real result, and the transform is precisely what makes it interesting.

  • Which direction does a buff that shifts the reply's language usually push a refusal-matching detector?
    Towards scoring more failures: the refusal is present but unrecognised, so the attempt looks like the model complied when it did not.
  • What is the minimum evidence you would want before reporting that a buff changed the target's behaviour?
    A paired unbuffed run over the same probes and endpoint, plus hand re-judgement of a sample of the attempts whose verdict differs between the two arms.
  • Can a buff move the number without any detector error at all?
    Yes — a model-backed buff can soften or drop the payload while rewriting, so some attempts simply never carried the probe's behaviour.

It is like marking translated exam papers against an answer key still written in the original language: the marks move sharply, and none of the movement is about the students.

saying these in an interview costs you the question

  • Reporting a buffed pass/fail rate as evidence about the model with no unbuffed reference arm.
  • Assuming the error can only run one way (only false negatives, or only false positives).
  • Treating a lower buffed rate as the model being safer under obfuscation.
  • Never reading a single buffed prompt-and-reply pair before quoting the aggregate.

context