Some of garak's buffs rewrite a probe's outbound prompts by calling a language model to paraphrase or translate them. What does that add to a scan you have to defend, and how do you handle it?
answer
- stochastic rewriting stage
- same config, different prompts
- paraphraser sanitises the payload
- log the prompt as sent
- exploration tool, not a gate
basics
~20 sIt makes the run non-reproducible: a rerun paraphrases differently, so the prompts sent are not the same set. The paraphraser can also soften or refuse the payload, silently weakening the probe. Keep the unbuffed run as the baseline, read a sample of transformed prompts, and cite the logged prompts, not the buff name.
solid answer
~60 sA model-backed buff turns a deterministic pipeline into a stochastic one. Three things follow. **Reproducibility.** Two runs of "the same" configuration send different text, so a rate that moves between runs may be paraphrase variance rather than a change in the target. Only the run's attempt log — the prompt as actually sent — identifies what you tested. **Payload integrity.** The rewriting model is itself aligned. It can decline to paraphrase, or quietly sanitise, so some attempts leave carrying a weaker request than the probe intended. Those come back clean and look like the target holding, which is a false negative you will never see in the aggregate. **A second system to operate.** The paraphraser needs its own compute and can fail, hang or rate-limit independently of the target. The handling is the same in each case: treat the buffed run as an extra arm, never a replacement for the canonical unbuffed one, and read a sample of the transformed prompts before you believe any number that came out of it.
go deeper
Says the rewriting is done by another model, so the prompts differ between runs and the run is not exactly repeatable.
Adds that the rewriting model can refuse or soften the prompt, and that the log of prompts as sent is what identifies the experiment.
Separates variance, payload degradation, attribution loss and the second operational dependency, and pairs each with a concrete handling step.
Rules on where such a run may be used at all — exploration and audit yes, release gate no — and what a finding produced this way must carry to be actionable.
**What changes when the rewriter is itself a model.** In an unbuffed garak run the prompts are fixed artefacts of the probe: rerun the same command and the same text goes out. Some of garak's buffs are mechanical and keep that property — a lowercase buff is a pure function of its input. The paraphrase buffs are not: they load a sequence-to-sequence model and *sample* rewordings, and a translation buff calls a translation model or service. The run configuration therefore stops being a complete description of the experiment. "We ran probe set X against endpoint Y with buff Z" no longer says what was sent; only the attempt log does. **Failure modes worth naming.** - *Variance mistaken for signal.* A few points of movement in a pass/fail rate between two buffed runs of the identical configuration may be nothing but the paraphraser's sampling. You cannot tell without repeats, and repeats multiply a budget you already multiplied by fanning out. Even where garak's run seed is set, the seed may not reach into a buff's own decoder, so do not assume determinism you have not observed. - *Silent payload degradation.* The rewriting model has its own alignment and its own refusal behaviour. It can decline to paraphrase, or return a polite, defanged version of an adversarial prompt. The target then behaves well because it was asked nicely, the attempt is scored clean, and the report says the target resisted. This is the failure mode that most flatters a target, and no automated stage in the pipeline looks for it. - *Silent fall-through.* When the rewriter errors, times out or returns nothing usable, the pipeline may send the original prompt instead. Part of your "buffed" arm is then unbuffed, and the arm is a mixture whose proportions you do not know. - *Attribution loss.* A hit found under paraphrase has to be reproduced from the exact prompt in the log. Handing an owning team the probe's canonical text will not reproduce it, and the finding gets closed as unrepeatable. - *Operational coupling and data egress.* A local paraphraser wants its own weights on disk and its own GPU or CPU cycles, competing with nothing else you scheduled; a hosted one is a second dependency that can rate-limit or fail independently, and it is a second place your adversarial prompts land — which matters for the customer whose endpoint you are testing. **What it costs.** Beyond the fan-out multiplier on target calls, budget the rewriting compute (a local paraphraser is a model download in the hundreds of megabytes to a few gigabytes plus inference time on every prompt), the repeats you need to bound variance — three repeats of a fourfold fan-out is twelve times the unbuffed arm — and the reading time for the transformed-prompt sample, which is the only check that catches payload degradation. **Where the number misleads.** A buffed rate that comes back *lower* than the unbuffed baseline reads naturally as "the target is robust to rewriting". Two other explanations produce the identical number: the detector no longer recognises the reply, or the rewriter defanged the request. Both are properties of your instrument, and neither is visible in the aggregate. A buffed rate that comes back *unstable across reruns* reads naturally as flakiness in the target; more often it is the paraphraser's sampling. Both readings fail the same way — they attribute to the target something that happened before the target was ever reached. **How to run it responsibly.** Keep the unbuffed run as the reference arm and ship it every time. Run the buffed arm separately, one transform at a time, so any delta attributes to one transform. Pin the rewriting model's identity and revision, and record it beside the run, or "the same buff" will mean different things across quarters. Pull a stratified sample of transformed prompts out of the run's `.report.jsonl` and read them, checking that the probe's actual demand survived the rewrite; count how many did, and quote that proportion in the write-up as an explicit bound on the arm's validity. Repeat the buffed arm enough times to see its own variance before attributing any movement to the target. When you file a finding, quote the prompt as sent verbatim and note that it was model-generated, so the reader knows configuration alone will not reproduce it. And for anything that must act as a regression signal across releases, prefer a deterministic transform or a frozen set of variants generated once and checked in: a live paraphraser is an exploration and audit instrument, never a gate.
- Why can a model-backed buff produce false negatives that the run's report cannot show you?The rewriter may defang the request before it is sent. The attempt is scored clean and looks like the target holding, and only reading the transformed prompt reveals that the probe never arrived intact.
- What would you use instead if the buffed comparison has to run every release?A deterministic transform, or a frozen set of variants generated once and checked in, so the release-over-release signal is not confounded by fresh sampling.
A model-backed buff is a translator who is also a censor: the message that reaches the other side may already have had its teeth pulled, so the polite reply you get back is evidence about the translator, not about the recipient.
saying these in an interview costs you the question
- Treating a model-paraphrased run as reproducible because the configuration is identical.
- Never checking whether the rewriter kept the probe's payload intact.
- Using a model-backed buff as a per-release regression gate.
- Filing a finding that quotes the probe's canonical prompt rather than the prompt actually sent.