A garak run over a hand-picked subset of probes finishes with nothing flagged. How do you write that up so it is not read as "the model is safe", and what denominator do you attach when you use the word coverage?
answer
- three denominators for coverage
- probes run over catalogue, not surface
- record endpoint, system prompt, date
- a zero is the detector's verdict
- untested is not passed
basics
~20 sReport it as: these named probes, at this repeat count, against this endpoint and configuration, on this date, produced no detector hits. Coverage means probes run divided by probes in the catalogue — not attack surface and not behaviours. Name the families you did not run as untested, never as passed.
solid answer
~60 sThree separate denominators hide behind the word coverage, and only the first is something a garak run measures: probes run out of the catalogue available; attack surface reached (the app's tools, retrieval path, multi-turn flows); and real-world behaviours a user might elicit. A subset run gives you a fraction of the first and says nothing about the other two. So the write-up is a bounded sentence, not a verdict. It states the exact selection, the repeat count, the target configuration the calls hit — endpoint, system prompt, decoding settings, any guard in front — and the date. A clean result then means: for those probes, under that configuration, garak's detectors fired no hits. The second bound is the detector. A no-hit row is the detector's judgement, not the model's behaviour; a miss by a detector and a genuine refusal look identical in the report. And garak's own reporting frames results relatively — a pass or failure rate per probe, an absolute score, and a calibration Z-score against a comparison set — which is a position, not a safety certificate. Untested families go in an explicit list.
go deeper
Should at least say that a clean scan only covers the probes that ran and should not be called 'safe'.
Should give the probes-run-over-catalogue denominator and record the target configuration and repeat count alongside the result.
Separates the three meanings of coverage, bounds the zero by detector sensitivity, spot-checks raw attempts, and lists untested families explicitly.
Decides what a scan result is allowed to gate, ensures the write-up template forces the denominator, and makes sure nobody downstream converts a relative score into a safety claim.
The failure this question tests is a single leap: from *the tool printed no failures* to *the system is safe*. Every clause of the write-up exists to block that leap, and the discipline is to make each bound explicit rather than trusting the reader to infer it. ### Fix the denominator Three different things get called "coverage", and a garak run measures only the first: | claimed coverage | what it would mean | does the scan measure it? | |---|---|---| | probe coverage | probes run / probes in the catalogue you selected from | **yes** — this is the only one | | attack-surface coverage | share of the deployment's reachable paths exercised | no — garak talks only to the generator you wired up | | behavioural coverage | share of real-world eliciting behaviours reproduced | no — probes sample families, they do not enumerate them | So state the number with its units visible: *N of the M probes in the catalogue we selected from*, produced by diffing `garak --list_probes` against the probes present in the report. The moment that fraction is relabelled "80% coverage" in a slide, it has silently become a claim about the product rather than about the tool. Attack-surface is the substitution that hurts most in practice. If the deployment has retrieval, tool or function calling, multi-turn session state, or a front-end guard, and the scan pointed at a bare model endpoint, none of those paths were touched — and in real incidents that is where the failure lives. ### Fix the configuration A result is only meaningful against the configuration that produced it, and every part of that configuration drifts. Record the endpoint and model version, the system prompt in force (with its version), decoding settings, `--generations`, whether a guard sat in front, and the date. A clean scan against a raw model says nothing about the guarded production stack; a clean scan against the guarded stack says nothing about the model underneath. Those are two different results and they support two different decisions. ### Fix the epistemics of a zero Every zero in the report is a **detector's** decision, not an observation of the model. A detector that missed and a model that genuinely refused produce identical rows. So "no hits" is evidence bounded by detector sensitivity, and the bound is not small: detectors keyed to particular phrasings can miss a model that complies in a different register, and a probe whose calls all errored also reports as unalarming. The verification is cheap and you should budget for it: pull the raw attempts behind at least one clean probe out of the JSONL report and read a sample of the responses yourself. Thirty to sixty minutes per engagement catches the two failures that would otherwise invalidate the whole write-up — attempts that never reached the model, and a detector criterion that does not match how this model phrases compliance. Compare each probe's attempt count against its expected prompts x generations figure while you are in there. ### Use the tool's own framing, not a stronger one garak reports per-probe pass and failure rates, an absolute score, and a **calibration Z-score** positioning the target against a comparison set of models. A Z-score is a *relative position*: it says this target does better or worse than comparable models on that probe. It is not a threshold, not a margin of error, and not a pass mark — the comparison set could itself be uniformly weak. Reporting "scored above average" as though it were "meets bar" is how a relative statistic becomes a certification nobody meant to issue. ### The deliverable sentence > "garak <version>, probes X, Y, Z at N generations, against endpoint E with system prompt version V and guard G, on date D: no detector hits. Families A, B, C were not run and remain untested. Retrieval and tool-calling paths were out of scope. One clean probe was spot-checked at the raw-attempt level." Boring, bounded, and it survives a later incident. The cost of getting this wrong is not abstract: a write-up that reads as a clean bill of health gets used to gate a release, and when something surfaces afterwards the argument is about what you claimed, not about what you ran. Unrun families go in their own explicit list, under the heading *untested* — never under *passed*, and never omitted, because an omitted family reads as fine to every downstream reader.
- A stakeholder asks for a single coverage percentage. What do you give them?Probes run over probes in the catalogue you selected from, stated with those words. Refuse to let it be relabelled as attack-surface or behavioural coverage, and pair it with the explicit list of untested families.
- Why is a clean garak result against a raw model endpoint weak evidence about the deployed product?The scanner only exercises the generator it was wired to. Retrieval content, tool calls, session state and any front-end guard are paths it never traversed, and those are where most real deployment failures live.
- How do you sanity-check that a probe with zero hits really ran?Read the raw attempts behind it in the report log: confirm calls were made, responses came back non-empty and non-errored, and spot-read a few outputs against the detector's criterion.
Reporting 80% coverage after a subset scan is like a pollster announcing 80% turnout when the real number is 80% of the people who answered the phone. The percentage is arithmetically true; the denominator quietly changed from everyone to the ones we managed to reach.
saying these in an interview costs you the question
- Writing 'garak found no vulnerabilities' with no selection or configuration attached.
- Reporting a coverage percentage without saying what the denominator is.
- Treating a zero-hit probe as proof of model behaviour rather than a detector verdict.
- Presenting the scan's relative score as a pass mark or certification.
- Listing unrun probe families under 'passed' or omitting them entirely.