In a garak run, what does it actually mean when an attempt is marked a hit, and why is that not yet a confirmed vulnerability?
answer
- probe elicits, detector judges
- hit = detector opinion, not evidence
- read the untruncated transcript
- echo, error text, truncation
- you write the finding, not the tool
basics
~20 sIt means garak's detector, not a human, judged that one model response matched its rule for failure. The detector may be a string match or a small classifier, so it can be wrong. Open the logged attempt, read the prompt and the full response, and confirm the model really did the thing.
solid answer
~50 sA garak run has two independent halves. A **probe** generates prompts and sends them to the configured generator; a **detector** then reads each response and decides whether it counts as a failure. A hit is that detector's verdict on one attempt — nothing more. Detectors are cheap on purpose, because they have to score thousands of attempts. Many are substring or regex matchers, some check for the *absence* of refusal language, some are small classifiers. All three fail in both directions: a response that quotes the request back, refuses in wording the detector does not recognise, or returns an API error string can all trip a hit without the model ever complying. So the workflow is: garak produces candidate evidence, you produce the finding. Before a hit goes in a report, open the per-attempt log, read the prompt and the untruncated response, and decide yourself whether the behaviour is what the probe claims. That transcript, not the counter, is what you show a stakeholder who pushes back.
go deeper
Says a hit is the detector's verdict on one attempt and that you must read the transcript before calling it a vulnerability.
Adds that probe and detector are separate components, names the cheap detector families, and gives concrete artefact cases such as prompt echo or error text.
Turns it into a triage procedure: sample per probe, verify the untruncated response, distinguish model output from guardrail output, and record what was actually read.
Frames it as evidentiary discipline — the report's credibility rests on verified transcripts and a stated sampling rate, not on tool counters.
**The mechanism, named at the level of the objects that do it.** A garak run is assembled from four plugin families, and knowing which family owns which decision *is* the answer to this question. A **probe** — selected with `garak --probes <family.Class>` — owns the attack shape: it holds a set of prompts and emits them as *attempts*. A **generator** — selected with `garak --model_type` plus `--model_name` — owns delivery only; it is the adapter to an OpenAI-style API, a local Hugging Face model, or a REST endpoint described by a JSON config, and it has no opinion about content. A **detector** owns judgment: it reads each returned string and emits a score per output, and a score past its threshold marks that output a **hit**. The harness then writes a per-attempt JSONL report plus an HTML summary (`--report_prefix` controls the filenames), and the per-probe number you eventually quote is *derived from those per-attempt verdicts* — nothing in that pipeline involved a human reading anything. So "a hit" decomposes to exactly this: *one detector, applied to one output string, said that string matched its rule for failure.* That is a machine opinion about text. A vulnerability is a claim that a system can be induced to do something it should not, reachable by a user who matters. The gap between those two sentences is the entire job of triage. **Why detectors are weak by construction.** A detector runs once per generated output across a whole sweep, so it has to be fast and dependency-light. The common families are (a) substring and regex matchers — "does this output contain the marker text"; (b) refusal-absence heuristics — "the output contains none of my known refusal phrases, therefore the model complied"; and (c) small classifiers fine-tuned to score toxicity, leakage or compliance. Every one of those is a cheap proxy for a semantic question. None of them can tell an echoed prompt from an authored answer, a cautionary paragraph from an instructional one, or a provider error body from a model utterance. **What it costs — and where that cost distorts the number.** The call volume of a run is explicit and multiplicative: `attempts ≈ (prompts in the selected probes) × garak --generations`, further multiplied if you sweep several probes. `--generations` does not default to 1 in every garak version, so read your version's default before budgeting rather than assuming it. A broad `garak --probes all` sweep against a metered chat endpoint is comfortably tens of thousands of calls: real money, hours of wall-clock once rate-limit backoff kicks in, and a report with more flagged rows than any person will read. `--parallel_attempts` buys wall-clock, not calls. The costly half is human. Opening a transcript, reading the untruncated output and deciding who authored the incriminating span is two to five minutes of an engineer's attention. Several hundred hits is therefore not a day of work — it is more than a day, which is precisely why the counter gets quoted instead of read. **Where the number misleads.** Three distortions, in the order they bite: 1. **The count scales with a config knob, not with the model.** Raise `--generations` and a single stubborn behaviour produces proportionally more hits. Two runs of the same model with different generation counts are not comparable as raw hit totals. 2. **Refusal-absence detectors count non-answers as compliance.** Empty completions, timeouts rendered as text, provider error bodies and refusals phrased outside the detector's phrase list all land as hits without a single harmful token in them. 3. **The denominator is attempts issued, not behaviours covered.** A pass rate answers "how many of the strings I sent were scored clean", never "how much of the attack surface did I reach". **What you actually check before a hit becomes a finding.** Open the per-attempt record in the JSONL and read the prompt and the *full* output, not the summary snippet. Then ask, in order: Is the matched span the model's own writing, or the prompt reflected back? Is the output a transport or provider error string? Was the output truncated before the part that would settle it? Is the "compliance" theatrical — a refusal wearing the requested costume, or fiction that conveys no capability? And did a guardrail in front of the endpoint author this text, in which case you have learned about the filter and nothing about the model? **And the symmetric warning.** The same cheapness means a *non*-hit is not evidence of safety: a detector that missed an unusual compliance phrasing produced a false negative that appears nowhere in the counters. Reading a small sample of non-hits from a probe you expected to land is the cheapest available check on that blind spot.
- Name two ways a garak detector can fire on a response where the model never complied.The response echoes the requested wording back inside a refusal, or the endpoint returned an error/placeholder string that happens to lack any refusal language a refusal-absence detector looks for.
- If you can only read a handful of transcripts, which hits do you open first?One or two per probe that reported hits — coverage across behaviours beats depth on one — plus any probe whose result contradicts what you expected from that endpoint.
- Does a hit tell you the severity of the issue?No. Severity comes from what the transcript shows and how the behaviour reaches a real user; the tool only reports that a detector's rule matched.
A garak detector is the metal detector at an airport gate: it beeps at a belt buckle exactly as loudly as at a knife. The beep tells you to open the bag; it never tells you what is inside.
saying these in an interview costs you the question
- Treating the hit count as the number of vulnerabilities found.
- Never opening a transcript before reporting.
- Assuming detectors are semantic judges of harm.
- Believing zero hits from a probe proves the model refuses that behaviour.