skip to content

A garak scan against your deployed chat endpoint reports a high failure rate for one probe family. Before that number goes into a report, how do you establish that it reflects real model behaviour rather than detector error?

level: seniorimportance: must knowfreq 58%

answer

  1. read the raw prompt-and-reply records
  2. refusal that restates the request
  3. empty, truncated, error-body replies
  4. random fixed-size sample, criterion written first
  5. publish verified rate + detector + n

basics

~20 s

Open the logged prompt-and-reply records behind those failures and hand-read a sample. Check what the detector actually matched: refusal text, a stock disclaimer, an echo of the prompt, an empty or error reply, or a safety lecture can all trip a matcher. Then report the verified count, the detector used and the sample size.

solid answer

~1 min

Triage before you publish, in this order: 1. **Identify the adjudicator.** Find which detector produced those failures and what class of matcher it is. That tells you which artefacts to expect. 2. **Read the raw attempts.** Pull a random sample of the failing prompt-and-reply records from the run's log - not the summary - and score them yourself against a written definition of what would count as a real failure. 3. **Look for the usual artefacts.** Detector matches on refusals that quote the request, boilerplate warnings, the prompt echoed back, truncated or empty replies, and transport errors returned as text are the common inflators when the target is a deployed app rather than a raw model. 4. **Separate app behaviour from model behaviour.** If a guard or system prompt in front of the model rewrites replies, the detector is scoring the app's output, which is what you want, but say so. 5. **Report what you verified.** Give the hand-checked rate with its sample size next to the tool's headline number, and name the detector. A raw tool count is a starting queue of candidates, not a list of findings. If hand-scoring says the rate is roughly right, you now have evidence rather than a claim.

go deeper

for a junior

Knows not to quote the headline number straight, and that the logged prompts and replies can be opened and read.

for a middle

Runs a sample, names the specific artefacts to look for - refusals restating the request, boilerplate, echoes, empty or error replies - and reports the detector alongside the rate.

for a senior

Runs it as a method: fixed random sample, written success criterion, scores both sides, records the disagreement rate as the caveat, and reproduces the attempts they intend to quote.

for a principal

Makes triage a gate rather than a courtesy - no untriaged tool count reaches a stakeholder, and the disagreement rate is tracked over time as a property of the detector set.

### What the printed rate actually is Before triaging, be exact about the arithmetic, because half of the bad readings come from the denominator rather than the detector. garak sends each probe prompt `--generations` times, and each generated output is scored separately. The rate you see is therefore **outputs the detector scored above the cut, divided by outputs it saw** - not "prompts that broke the model". With `--generations 10`, 120 hits out of 1,200 attempts could be 120 distinct prompts each failing once, or 12 prompts failing every single time. Those are very different findings, and the summary number cannot tell them apart. Group the hit log by prompt before you interpret anything. ### The two things that inflate a rate, and how to tell them apart **Detector artefacts.** These are cheap to find and you find them by reading, not by re-running. Against a deployed chat product rather than a raw model, the recurring ones are: - a refusal that restates the request verbatim, so a literal matcher finds its string inside the refusal; - standing product boilerplate - a safety preamble, a disclaimer, a legal footer - that happens to share wording with what the detector watches; - the model echoing the prompt back before answering; - empty strings, truncations at a token cap, or an HTTP 429/500 body captured as the reply text - which an inverted, refusal-absence detector scores as a bypass; - a classifier detector meeting an output style unlike its training data and drifting. **Real behaviour.** If a hand-scored sample agrees with the detector, the rate stands - and now you can defend it in the room, because you can display three or four verified attempts with their prompts and replies. ### The method that survives a hostile read 1. **Fix the sample size before you look.** Pre-commit to n, then draw at random from the hit log. Choosing the interesting-looking failures produces whatever answer you set out to find, and a reviewer can always ask how the examples were picked. 2. **Write the success criterion down first**, in one sentence, so a second person scoring the same attempts would land in the same place. 3. **Score the passes too**, not only the hits. Hits reveal false alarms; only passes reveal misses. Without both sides you learn the rate is too high but never that it is also too low. 4. **Record the disagreement rate.** That single figure is the honest caveat for the whole probe family, and it is the evidence you produce when someone asks why you overrode a tool. 5. **Reproduce anything you intend to assert.** One clean confirmation run of the specific attempts you will quote turns a lucky sample into a demonstrated behaviour. ### What triage costs This is engineer time, and it is the part people fail to budget. Hand-scoring an attempt means reading a prompt and a reply and applying a written criterion: one to two minutes each once you are warmed up. Fifty attempts is therefore an hour or two per probe family, and you need it for hits *and* passes. Precision is bought expensively: with n=20 you cannot distinguish a 60% from an 85% true rate, so a small sample supports "the tool is roughly right" or "the tool is badly wrong", not a corrected decimal. Size the sample to the claim you intend to make, and if the budget only allows a small one, publish the tool's rate with a qualitative note rather than a spuriously precise corrected figure. ### Where the corrected number can still mislead Two traps remain after honest triage. First, correcting only the false alarms produces a *lower* number that is still a lower bound, because the misses you never sampled are untouched - a triaged rate is not a true rate. Second, if a guard or system prompt in front of the model rewrote the replies, the detector scored the **application**, not the model; that is usually what you want when testing a product, but the sentence in the report must say which was under test, or the number will be quoted against the base model later. ### What you publish The tool's raw rate; the hand-verified rate with its sample size; the detector name and matcher class; the probe selection and garak version; the disagreement rate; and reproduced examples for every claim you make as a finding. A raw tool count is a queue of candidates, not a list of findings, and calling it one is what gets a report walked back.

  • Why sample the passes as well as the failures?
    Failures only reveal false alarms. Reading passes is the only way to find attempts the detector missed, which is what tells you the rate is a floor rather than a measure.
  • Your hand-scored sample disagrees with the detector on roughly a third of the failures. What goes in the report?
    Both numbers, the sample size, the detector, and the disagreement rate as an explicit caveat on that probe family - plus reproduced examples for anything you assert as a finding.
  • Why fix the sample size before looking at the data?
    Otherwise you stop when the answer suits you. A pre-committed random sample keeps the verified rate defensible when someone asks how the examples were chosen.

saying these in an interview costs you the question

  • Publishes the scanner's rate with no hand-verification at all.
  • Cherry-picks which failures to inspect rather than sampling.
  • Inspects only the failures, never the passes, so misses are never discovered.
  • Blames the model for behaviour that turns out to be an error body or an empty reply captured as text.
  • Quotes the rate without stating which detector produced it.

context