You are asked to report whether rewriting prompts changes how often a target fails a chosen set of garak probes. How do you set up the buffed and unbuffed garak runs so the comparison actually means something, and what is the trap in the rate itself?
answer
- two arms, one variable
- freeze probes, detectors, generations
- run arms adjacent in time
- attempts vs original prompts
- hand-judge the verdict flips first
basics
~20 sRun the same probes twice against the same endpoint, changing only the buff. Hold generations, detectors and endpoint settings fixed, and run them close together. Then compare like for like: a fan-out buff sends several variants per original prompt, so decide whether your rate counts attempts or original prompts.
solid answer
~50 sTreat it as a two-arm experiment with one variable. Same probe selection, same detectors, same generations per prompt, same endpoint and endpoint configuration, runs close together in time so a model or policy update behind the endpoint does not become the hidden variable. Enable one buff, not several, so the delta attributes to one transform. The trap is the denominator. A fan-out buff turns each original prompt into several attempts, and a target that fails on any one variant now has more chances to fail. A per-attempt rate and a per-original-prompt rate answer different questions: the first says "how often does a sent prompt succeed against it", the second says "how often does this probe's idea get through at least once". Quoting one while implying the other is the standard way this comparison misleads. State which denominator you used, and if you report both arms, use the same one for each.
go deeper
Says run the same probes against the same endpoint with only the buff changed, and compare the two reports.
Adds the frozen variables (generations, detectors, endpoint config), one buff at a time, and notes the attempt count is not equal between arms.
Names the denominator choice explicitly, runs the arms adjacent in time, adds repeats for a stochastic buff, and hand-judges verdict flips before quoting anything.
Decides what the organisation is allowed to publish from such a comparison and what evidence must accompany it, including cost-normalised framing for a metered target.
**The design: one variable, two arms.** You are running an experiment, so it obeys experimental discipline. Freeze everything except the transform: the same probe selection (`garak --probes`), the same completions per prompt (`garak --generations`), the same detectors — in practice the probes' own defaults, unchanged — the same generator and endpoint configuration (`garak --model_type` / `--model_name` plus whatever auth, system prompt, temperature and decoding settings that endpoint applies), and the same seed if the run supports one. Vary one thing: garak's `--buffs` setting, with one buff enabled, not several, so any delta attributes to a single transform. Run the two arms adjacent in time. A hosted endpoint can be updated under you — a model version, a system prompt, a moderation policy — and a week between arms silently promotes that change to your independent variable. If the buff is stochastic, add repeats *within* the buffed arm before comparing across arms, so you can see the transform's own variance first. **The denominator trap.** This is where an otherwise sound comparison misleads. Suppose a probe contributes 50 prompts and the buff emits 4 variants of each: the buffed arm sends 200 attempts against the unbuffed arm's 50. There are three defensible statistics, and they answer different questions. | statistic | denominator | what it answers | the catch | |---|---|---|---| | per-attempt failure rate | attempts sent | how often a sent prompt gets through | arms differ fourfold in effort | | failed at least once per original prompt | original prompts | does this probe's idea get through at all | flatters the buffed arm: four tries, four chances | | failures per fixed call budget | calls | which spend finds more, at equal cost | needs the fan-out factor to construct | None of the three is wrong. Quoting one while implying another is, and the at-least-once figure is the one that seduces, because it rises with effort alone even against a target with no particular weakness to the transform. State the denominator in the same sentence as the number, and use the same one for both arms. **The validity check that must come first.** Every rate comparison silently assumes the detector rules buffed replies as accurately as it rules unbuffed ones. That assumption fails often, because a transform moves replies off the distribution the detector was tuned for: a refusal-matching detector blinded by a language shift manufactures failures, a keyword matcher blinded by rewording suppresses them. So before computing any delta, pull the attempts whose verdict *differs* between arms, sample them stratified by probe, and read the prompt as sent and the reply from the run's `.report.jsonl`. Re-judge by hand and compare your labels with the detector's, separately per arm. If the detector's error rate differs materially between arms, the arms are not comparable and there is no delta to report — the finding is about your instrument. **What it costs.** The buffed arm already costs the fan-out multiplier; the reference arm adds the unbuffed run on top; repeats for a stochastic buff multiply again — three repeats of a fourfold fan-out plus one reference arm is thirteen times the unbuffed call count. Then the hand-judged sample: 200 verdict-flipped attempts at a minute or two each is half a day of a domain-literate engineer, and that half day is not optional, because it is the only thing standing between you and a number that is an artefact. On a metered endpoint, price all of that before you promise the comparison. **What the write-up carries.** Both arms' full configurations; the measured fan-out factor, not the documented one; which denominator you chose and why; the size, sampling scheme and outcome of the hand-judged sample, including the detector's estimated error rate in each arm; and, for anything you are calling a finding, the verbatim prompt as sent. A comparison reported without the fan-out factor and the denominator is not reproducible, and a reader who cannot reconstruct which of the three statistics you used cannot check your claim at all.
- Why does a fan-out buff flatter the buffed arm on a 'failed at least once' statistic?More variants means more independent chances for the target to slip, so the at-least-once figure rises with effort alone, independent of any real weakness to the transform.
- What single check comes before you compute any delta between the two arms?Hand-judging a sample of the attempts whose verdict differs, to confirm the detector is still ruling correctly on the transformed arm's replies.
saying these in an interview costs you the question
- Comparing arms run weeks apart against a hosted endpoint that may have changed.
- Enabling several buffs at once and attributing the delta to one of them.
- Quoting a per-attempt rate while making a per-prompt claim, or vice versa.
- Computing the delta before checking the detector still rules buffed replies correctly.