A vendor hands you a garak run in which their chat product passed every probe whose prompts come from a fixed, publicly readable list. What do you check before accepting that as evidence of robustness, and what would you ask them to run instead?
answer
- ask for the probe selection first
- public list = floor, not robustness
- raw model or shipping path?
- date and build of the run
- request generated-prompt rerun plus logs
basics
~20 sAsk which probes ran and whether their prompts are a public fixed list. A public list can be filtered or tuned against, so a pass shows those strings are handled, not that the behaviour is absent. Ask for a rerun that includes probes minting prompts at run time, with the probe selection and logs attached.
solid answer
~50 sTreat the artefact as a claim with a denominator, and go looking for the denominator. **Check first:** which probes were selected, how many attempts each sent, which endpoint the run pointed at (the production path with its guardrails, or a bare model behind them), and when it ran relative to the build you are buying. **Then the contamination question.** If every passing probe draws from a published list, the pass is consistent with two very different worlds: a genuinely robust system, and one with a filter tuned to those exact strings. The report cannot separate them, and neither can the vendor's summary. **What to ask for:** the same selection plus probes that generate prompts during the run, executed against the shipping configuration, with the run logs and the probe selection handed over rather than a screenshot of a score. If the vendor will only supply the public-list result, record that limitation in your own assessment rather than restating their conclusion.
go deeper
Notices that a passing report only covers the probes that were run and asks which ones those were.
Explains that a publicly readable prompt list can be defended against string by string, so a pass is a floor rather than proof.
Works the whole chain — selection, prompt provenance, which endpoint, build and date, what ruled hits — and asks for a rerun with generated prompts plus the raw artefacts.
Turns this into a procurement standard: what a supplier must ship with any scanner result before it counts as evidence in a risk decision.
**Frame.** You are not evaluating a model here. You are evaluating an artefact produced by an instrument somebody else drove, on a day of their choosing, against a target of their choosing. Everything you may conclude is bounded by configuration choices that live in the run's own records — and that a summary screenshot deliberately or accidentally omits. **What the artefact actually contains.** A garak run writes a machine-readable report (`garak --report_prefix <name>` controls the filename) as newline-delimited JSON alongside the human-facing HTML. The first records describe the run's setup: the model type and name behind the generator, the probe specification, the generations count, the package version. Then comes one record per attempt, carrying the prompt that was sent, the outputs that came back, and the detector's per-output results, and finally the per-probe evaluation lines the HTML renders as a score and a coarse severity grade. **The HTML is a rendering; the JSONL is the evidence.** Ask for the JSONL, and recompute the pass rates yourself rather than reading their summary — it takes minutes and it is the only way to know that the numbers on the slide came from the run attached to it. **The questions, roughly in order.** 1. *Which probes, and how many attempts each?* Without the selection there is no denominator, and a narrow selection with a clean sweep is easy to produce in good faith. Count distinct probe names in the attempt records rather than trusting the cover note; probes that failed to load leave no attempts but do not stop the run. 2. *Static list or generated?* If every passing probe draws from a published prompt file, the pass is equally consistent with a robust system and with a filter tuned to those exact strings. That is contamination arriving through the probe, and no field in the report separates the two worlds. 3. *Against what endpoint?* If the generator pointed at a raw model API rather than the deployed path, the numbers describe a system nobody ships. If it pointed at the deployed path, the guardrails are inside the measurement and you cannot attribute the pass to the model. Both are legitimate runs; only one answers your question, and the setup record tells you which you got. 4. *When, and on which build?* Results age. A scan predating the release you are buying is a scan of a different system. 5. *What ruled the hits?* A pass is a detector staying silent. You need not audit their detector choices to note in writing that the ruling is part of the claim. **What it costs to fix the gap.** Asking for a rerun is asking the vendor to spend: a full static selection is prompts times generations against a metered endpoint, and adding generative probes adds a second model's calls plus an open-ended search budget, so a serious rerun against a rate-limited production path is a days-long job rather than an afternoon. Running it yourself under a testing agreement moves that same bill onto you and adds the legal work. Ask for the artefacts *first* — they are free — and only then negotiate the rerun you actually need. **Where the vendor's number misleads.** The calibration figure is the most common trap: it positions their score against other systems garak's maintainers have measured, so a favourable value means "relatively less bad on these probes", which is entirely compatible with a high absolute failure rate if the reference systems all fail too. The severity grade in the HTML is a coarse bucket, not a measurement. And an aggregate over probes is weighted by prompt counts, so a selection heavy in one bulky family can move the headline without anything about the product changing. **How you write it up.** State the bounded claim — "the published prompt families in the named probes produced no detector-recognised hits against this endpoint, on this date, at this package version" — and, separately, what remains untested: generative families, the shipping path if they scanned the bare model, everything outside the one endpoint. If they will not release logs, record the limitation, downgrade the evidence to a vendor assertion, and price the residual uncertainty. Reviewers can act on that. Nobody can act on "their scan was clean".
- The vendor scanned the model endpoint directly, behind none of the product's guardrails. Is the result useless?No, but it answers a different question. It characterises the model in isolation; it says nothing about what ships. You need a run against the deployed path before it informs a purchase.
- The vendor refuses to hand over run logs, citing confidentiality. What do you do?Record the limitation explicitly, downgrade the evidence to a vendor assertion, and either run your own scan under a testing agreement or price the residual uncertainty into the decision.
saying these in an interview costs you the question
- Accepts a summary score without asking which probes ran.
- Cannot say why a public prompt list weakens a passing result.
- Does not ask whether the scan hit the shipping path or a bare model endpoint.
- Restates the vendor's conclusion in their own assessment instead of the bounded claim.
- Treats a calibration figure as an absolute statement of safety.