skip to content

Reading the Report

Turning per-probe pass rates and graded scores into a claim you will defend in a room, including the hits you had to open by hand. Interviewers ask because an unread report loses the room fastest.

on this pageshow

explore

questions

8

In a garak run, what does it actually mean when an attempt is marked a hit, and why is that not yet a confirmed vulnerability?

level: juniorimportance: must knowfreq 78%

answer

  1. probe elicits, detector judges
  2. hit = detector opinion, not evidence
  3. read the untruncated transcript
  4. echo, error text, truncation
  5. you write the finding, not the tool

basics

~20 s

It means garak's detector, not a human, judged that one model response matched its rule for failure. The detector may be a string match or a small classifier, so it can be wrong. Open the logged attempt, read the prompt and the full response, and confirm the model really did the thing.

solid answer

~50 s

A garak run has two independent halves. A **probe** generates prompts and sends them to the configured generator; a **detector** then reads each response and decides whether it counts as a failure. A hit is that detector's verdict on one attempt — nothing more. Detectors are cheap on purpose, because they have to score thousands of attempts. Many are substring or regex matchers, some check for the *absence* of refusal language, some are small classifiers. All three fail in both directions: a response that quotes the request back, refuses in wording the detector does not recognise, or returns an API error string can all trip a hit without the model ever complying. So the workflow is: garak produces candidate evidence, you produce the finding. Before a hit goes in a report, open the per-attempt log, read the prompt and the untruncated response, and decide yourself whether the behaviour is what the probe claims. That transcript, not the counter, is what you show a stakeholder who pushes back.

go deeper

for a junior

Says a hit is the detector's verdict on one attempt and that you must read the transcript before calling it a vulnerability.

for a middle

Adds that probe and detector are separate components, names the cheap detector families, and gives concrete artefact cases such as prompt echo or error text.

for a senior

Turns it into a triage procedure: sample per probe, verify the untruncated response, distinguish model output from guardrail output, and record what was actually read.

for a principal

Frames it as evidentiary discipline — the report's credibility rests on verified transcripts and a stated sampling rate, not on tool counters.

**The mechanism, named at the level of the objects that do it.** A garak run is assembled from four plugin families, and knowing which family owns which decision *is* the answer to this question. A **probe** — selected with `garak --probes <family.Class>` — owns the attack shape: it holds a set of prompts and emits them as *attempts*. A **generator** — selected with `garak --model_type` plus `--model_name` — owns delivery only; it is the adapter to an OpenAI-style API, a local Hugging Face model, or a REST endpoint described by a JSON config, and it has no opinion about content. A **detector** owns judgment: it reads each returned string and emits a score per output, and a score past its threshold marks that output a **hit**. The harness then writes a per-attempt JSONL report plus an HTML summary (`--report_prefix` controls the filenames), and the per-probe number you eventually quote is *derived from those per-attempt verdicts* — nothing in that pipeline involved a human reading anything. So "a hit" decomposes to exactly this: *one detector, applied to one output string, said that string matched its rule for failure.* That is a machine opinion about text. A vulnerability is a claim that a system can be induced to do something it should not, reachable by a user who matters. The gap between those two sentences is the entire job of triage. **Why detectors are weak by construction.** A detector runs once per generated output across a whole sweep, so it has to be fast and dependency-light. The common families are (a) substring and regex matchers — "does this output contain the marker text"; (b) refusal-absence heuristics — "the output contains none of my known refusal phrases, therefore the model complied"; and (c) small classifiers fine-tuned to score toxicity, leakage or compliance. Every one of those is a cheap proxy for a semantic question. None of them can tell an echoed prompt from an authored answer, a cautionary paragraph from an instructional one, or a provider error body from a model utterance. **What it costs — and where that cost distorts the number.** The call volume of a run is explicit and multiplicative: `attempts ≈ (prompts in the selected probes) × garak --generations`, further multiplied if you sweep several probes. `--generations` does not default to 1 in every garak version, so read your version's default before budgeting rather than assuming it. A broad `garak --probes all` sweep against a metered chat endpoint is comfortably tens of thousands of calls: real money, hours of wall-clock once rate-limit backoff kicks in, and a report with more flagged rows than any person will read. `--parallel_attempts` buys wall-clock, not calls. The costly half is human. Opening a transcript, reading the untruncated output and deciding who authored the incriminating span is two to five minutes of an engineer's attention. Several hundred hits is therefore not a day of work — it is more than a day, which is precisely why the counter gets quoted instead of read. **Where the number misleads.** Three distortions, in the order they bite: 1. **The count scales with a config knob, not with the model.** Raise `--generations` and a single stubborn behaviour produces proportionally more hits. Two runs of the same model with different generation counts are not comparable as raw hit totals. 2. **Refusal-absence detectors count non-answers as compliance.** Empty completions, timeouts rendered as text, provider error bodies and refusals phrased outside the detector's phrase list all land as hits without a single harmful token in them. 3. **The denominator is attempts issued, not behaviours covered.** A pass rate answers "how many of the strings I sent were scored clean", never "how much of the attack surface did I reach". **What you actually check before a hit becomes a finding.** Open the per-attempt record in the JSONL and read the prompt and the *full* output, not the summary snippet. Then ask, in order: Is the matched span the model's own writing, or the prompt reflected back? Is the output a transport or provider error string? Was the output truncated before the part that would settle it? Is the "compliance" theatrical — a refusal wearing the requested costume, or fiction that conveys no capability? And did a guardrail in front of the endpoint author this text, in which case you have learned about the filter and nothing about the model? **And the symmetric warning.** The same cheapness means a *non*-hit is not evidence of safety: a detector that missed an unusual compliance phrasing produced a false negative that appears nowhere in the counters. Reading a small sample of non-hits from a probe you expected to land is the cheapest available check on that blind spot.

  • Name two ways a garak detector can fire on a response where the model never complied.
    The response echoes the requested wording back inside a refusal, or the endpoint returned an error/placeholder string that happens to lack any refusal language a refusal-absence detector looks for.
  • If you can only read a handful of transcripts, which hits do you open first?
    One or two per probe that reported hits — coverage across behaviours beats depth on one — plus any probe whose result contradicts what you expected from that endpoint.
  • Does a hit tell you the severity of the issue?
    No. Severity comes from what the transcript shows and how the behaviour reaches a real user; the tool only reports that a detector's rule matched.

A garak detector is the metal detector at an airport gate: it beeps at a belt buckle exactly as loudly as at a knife. The beep tells you to open the bag; it never tells you what is inside.

saying these in an interview costs you the question

  • Treating the hit count as the number of vulnerabilities found.
  • Never opening a transcript before reporting.
  • Assuming detectors are semantic judges of harm.
  • Believing zero hits from a probe proves the model refuses that behaviour.

context

open as a page

A finished garak run prints, for each probe it ran, a pass/failure rate alongside an absolute score. What is that rate a fraction of, and does a higher failure rate mean the tested model did better or worse?

level: juniorimportance: must knowfreq 55%

basics

~20 s

It is a fraction of that one probe's own attempts, not of the whole run. An attempt is one prompt sent to the tested model; the detector attached to that probe labels it pass or fail. A higher failure rate means more attempts got through, so the model did worse on that probe.

open as a page

Which failure modes of garak's detectors produce false hits, and how do you recognise each one from the logged attempt?

level: middleimportance: must knowfreq 66%

basics

~20 s

Detectors that match a substring, or that infer compliance from the absence of a refusal, fire on responses that never complied: the model quoted the request back, refused in unusual wording, returned an error string, or was cut off mid-sentence. Classifier detectors add their own misfires. Read the response text to tell them apart.

open as a page

A garak sweep against a chat endpoint returns zero hits on every probe you ran. What can you conclude, and how do you check the run was not silently broken?

level: seniorimportance: must knowfreq 60%

basics

~20 s

Only that no probe you ran, judged by its own detector, found anything. First prove the run was real: check that attempts reached the endpoint, that responses are non-empty and not error text, and that the probes you needed were in the set you ran. A broken generator looks exactly like a clean model.

open as a page

In a garak report one probe lands in a poor band on its absolute score while its calibration score sits mid-pack against the reference models, and another probe scores well absolutely but far below the reference models. You are writing the run up. Which number do you act on in each case, and why?

level: seniorimportance: must knowfreq 45%

basics

~20 s

They answer different questions. The absolute score says how bad the behaviour is in itself; the calibration score says how unusual it is versus reference models on that same probe. Poor absolute, mid-pack relative is a hazard the whole field shares. Good absolute, far below the pack is this model's own outlier defect.

open as a page

A garak report shows each per-probe rate with an uncertainty interval beside it, but for several probes that interval is blank. What does a blank interval tell you about those rows, and what would you change about the run before quoting one of those rates?

level: middleimportance: should knowfreq 35%

basics

~20 s

Blank means too few attempts. Below a sample-size floor the tool declines to print an interval rather than show a meaningless one, so that rate is an estimate you cannot bound. Re-run the probe with more generations per prompt, or a longer prompt set, until the interval appears, then quote it.

open as a page

A single garak sweep leaves you several hundred hits and a day to triage them. How do you turn hits into report findings without over- or under-counting?

level: principalimportance: should knowfreq 47%

basics

~20 s

Group hits by the behaviour they demonstrate rather than writing one finding per hit. Read several per group, confirm the detector was right, and keep one finding with a verified example and the count behind it. Drop groups where every sampled hit was an artefact, and state what you sampled.

open as a page

Your team publishes garak's per-probe calibration scores in a quarterly report. Between two quarters the tested model was not changed, yet several of those relative scores moved. What could explain the movement, and how would you make a quarter-over-quarter comparison of these numbers trustworthy?

level: principalimportance: should knowfreq 25%

basics

~20 s

A relative score is a position against reference data that ships with the tool, so it moves when that reference data or the probe's prompts change, even with your model untouched. Sampling noise and a silently updated hosted endpoint also move it. Pin the tool and reference bundle per series, and trend absolute scores.

open as a page