A finished garak run prints, for each probe it ran, a pass/failure rate alongside an absolute score. What is that rate a fraction of, and does a higher failure rate mean the tested model did better or worse?
answer
- denominator = that probe's attempts
- attempt = prompt x generations
- detector labels each attempt
- higher failure rate = worse
- never average across probes
basics
~20 sIt is a fraction of that one probe's own attempts, not of the whole run. An attempt is one prompt sent to the tested model; the detector attached to that probe labels it pass or fail. A higher failure rate means more attempts got through, so the model did worse on that probe.
solid answer
~50 sThe denominator is the attempts that **this probe** generated: its prompt list multiplied by however many generations you asked for per prompt. Each attempt is judged by the detector garak pairs with that probe, and the rate is simply how many of those attempts the detector called a failure. Higher failure rate = worse behaviour on that probe. The practical consequence is that rates are **per-probe, and not comparable across probes**. Two probes in the same run have different prompt counts, different prompts and different detectors, so an identical rate on two of them says nothing about relative severity, and averaging the column into one "garak score" is arithmetic on unlike denominators. The absolute score printed beside the rate is a graded summary of the same attempts for that probe alone; the calibration Z-score beside it is a different quantity again, comparing this rate against a reference set of models on the same probe.
go deeper
Says the rate is per-probe, that an attempt is one prompt sent to the model, and gets the direction right.
Adds that the denominator is prompts times generations, that the numerator is a detector's label, and that rates are not comparable across probes.
Frames the rate as an estimate with a sample size and a noisy numerator, and states how they would phrase it in a report so polarity and denominator are explicit.
Sets the house rule for how per-probe rates may be quoted and aggregated in team reporting, and why a single headline percentage is not offered.
### What the report actually counts A garak run is a cross-product of three things you set on the command line: the target (`garak --model_type` plus `--model_name`), which probes to run (`garak --probes`, or `all`), and how many completions to draw per prompt (`garak --generations`, whose default has changed between releases — read it back out of the run's own record rather than assuming a number). Every probe class ships a fixed list of prompts. garak expands that list by the generations setting into **attempts**, sends each attempt to the target, and hands the response to the **detector** the probe declares. A detector is a small classifier — a string or regex match, or a shipped model — that scores each attempt as a number; garak's evaluator thresholds that score into pass or fail and writes one `eval` record per probe-and-detector pair into the run's `report.jsonl`, carrying `passed` and `total`. The rate you read in the HTML report is those two integers. So the denominator is the attempts of **that one probe**, and the numerator is one classifier's thresholded opinion about them. Neither is a property of the run as a whole, and neither is ground truth. ### Direction, and why to write it out longhand More failures means more attempts the detector judged the model handled badly, so a higher failure rate is worse for the target. The trap is not the concept, it is the column heading: a report carries pass counts, failure counts and graded bands side by side, and "the probe scored 12%" is ambiguous in a way that survives into a slide and then into a decision. Write the sentence with its denominator attached — "the detector flagged 12 of 100 attempts on this probe" — and the polarity question cannot arise. ### What the run costs Attempts are the cost unit, and each attempt is one call to the target. attempts = prompts × generations, summed over the probes you selected. One probe with 60 prompts at 5 generations is 300 calls. A sweep across the full probe catalogue is comfortably tens of thousands of calls: hours of wall clock against a rate-limited hosted endpoint, and a bill that scales linearly with `garak --generations`. Doubling generations exactly doubles both money and time and buys you nothing but a firmer denominator. That arithmetic is why first-pass runs are usually done at a single generation, and it is also why so many rows in that first report have attempt counts far too small to quote as a percentage. ### Where the number misleads **Small denominators.** At ten attempts, one attempt is ten points of rate. A one-generation run over a short prompt list produces exactly that, and the difference between 10% and 20% on such a row is one coin flip. This is what the uncertainty interval beside the rate exists to say, and why it is left blank below a sample-size floor. **The numerator is a classifier.** A detector that keyword-matches will flag a refusal that merely quotes the forbidden word, and will miss a compliant answer phrased in a way it does not recognise. Both errors move the rate without the model's behaviour changing at all. The rate is therefore the input to triage, not a finding you can sign. **Cross-probe arithmetic.** Different probes carry different prompt counts, different prompts, different detectors and different real-world severities. Averaging the rate column into one "garak score" divides unlike numerators by unlike denominators and produces a number that no longer refers to anything. The same objection kills "12 of 40 probes had hits, so we are 30% vulnerable": that is a count over probes, a third denominator again. ### What to check before quoting a rate Read `total` for the row straight out of the run's `report.jsonl` rather than trusting a remembered `--generations` value. Confirm the interval printed at all. Open a handful of the flagged attempts in the hit log and read the model's actual output, because five minutes of that tells you whether the detector is measuring what you think. Say which quantity you are quoting — the probe's own rate, its absolute graded band, or the calibration comparison beside it — because a listener who hears only "the score" will supply whichever reading suits the conversation.
- You raised generations per prompt from 1 to 10 and the failure rate for a probe changed a lot. Did the model get worse?Not necessarily. You changed the denominator by a factor of ten, so the earlier rate was very coarse and could not have landed near the true value. The later number is better estimated, and only a comparison at the same attempt count says anything about the model.
- Can you compare a probe's failure rate in garak against an attack-success rate someone else measured with a different tool?No, not directly. The prompt sets differ, the number of attempts differs and, most importantly, the thing deciding what counts as a success differs. You can compare trends within one tool and one configuration; a cross-tool percentage comparison is not a like-for-like number.
saying these in an interview costs you the question
- Treating the rate as a run-wide number rather than one probe's number.
- Averaging per-probe rates into a single headline percentage.
- Assuming the numerator is ground truth rather than a detector's label.
- Cannot say how many attempts produced the rate they are quoting.
- Confusing the absolute score with the calibration Z-score printed beside it.