skip to content

A garak sweep against a chat endpoint returns zero hits on every probe you ran. What can you conclude, and how do you check the run was not silently broken?

level: seniorimportance: must knowfreq 60%

answer

  1. empty report has three causes
  2. read raw responses, not the summary
  3. identical responses = error path
  4. sample non-hits from a probe you trust
  5. positive control makes clean falsifiable

basics

~20 s

Only that no probe you ran, judged by its own detector, found anything. First prove the run was real: check that attempts reached the endpoint, that responses are non-empty and not error text, and that the probes you needed were in the set you ran. A broken generator looks exactly like a clean model.

solid answer

~50 s

Zero hits is a claim about the instrument, not about the system. Three causes produce the same empty report. **The run never happened properly.** Auth failed, the endpoint rate-limited every call, the generator pointed at the wrong deployment, or responses came back empty. Check by reading raw attempt records: real model text, plausible lengths, variety across attempts. **The run happened but was not scored.** A detector mismatched to the response format — a wrapped or JSON-carried answer, say — zeroes out real signal. Check by reading a sample of *non*-hits from a probe you were confident would land. **The run happened but was narrow.** A subset of probes, one language, single-turn, against a stack with a filter in front of the model. Then the honest statement names what ran against what. Build in a positive control so an all-clean report is falsifiable.

go deeper

for a junior

Says zero hits does not mean safe and that you should check the tool actually talked to the endpoint.

for a middle

Separates 'run failed', 'scoring failed' and 'run was narrow', and names concrete log checks for each.

for a senior

Runs the whole validation: traffic, content variety, sampled non-hits, response format versus detector expectations, filter in front of the model, and a coverage statement with a denominator.

for a principal

Institutionalises it — positive controls in every scheduled run, a job that fails when the control is silent, and a reporting standard that forbids stating a clean result without scope.

**Why this is the question that separates operators from tool users.** A noisy garak report gets scrutinised line by line; a clean one gets believed and forwarded. That asymmetry means the highest-cost failure in the whole workflow is a run that quietly did nothing and was circulated as evidence of safety. Zero hits is a statement about the instrument first and about the system only afterwards, and the burden is on you to prove the instrument was alive. **What "zero hits" literally asserts.** For every attempt that a selected probe emitted, the generator returned something, the attached detector scored that something below its threshold, and the harness recorded a pass. Notice how many independent things must have worked for that sentence to mean anything: the probes had to be the right ones, the generator had to reach the right deployment with valid credentials, the responses had to be real model text, and the detector had to be able to *see* the model's answer in the string it was handed. Break any link and the report is indistinguishable from a perfectly safe system. **Three causes, one empty report.** | Cause | What actually happened | Cheapest check | |---|---|---| | The run never happened | Auth failed, the endpoint rate-limited or 4xx'd every call, `--model_name` pointed at the wrong deployment, responses came back empty | Read raw per-attempt records: non-empty text, plausible and *varying* lengths | | The run was not scored | The detector was mismatched to the response format — for instance the endpoint wraps the answer in a JSON envelope or a tool-call structure and the detector matches over the raw returned string | Read a sample of **non-hits** from a probe you were confident would land | | The run was narrow | A subset of probes, single-turn only, one language, and a filter sitting in front of the model | State scope explicitly: which probes, how many attempts, what surface | **Step 1 — prove traffic and content.** Inspect the per-attempt records in the JSONL report, never the HTML summary. You want non-empty outputs, plausible lengths, and genuine variety across attempts. Identical or near-identical short strings repeated across every attempt is the fingerprint of an error path, a cached reply or an auth failure — not of a well-behaved model. Then count attempts against what the selected probes multiplied by `garak --generations` should have produced; a shortfall means calls were dropped, retried away or cut short. **Step 2 — prove scoring.** Pick a probe whose attack you understand and read a sample of its *passes*. If you find outputs that plainly complied yet scored clean, your detector is mismatched, and the usual culprit is format: the substance sits inside a JSON field, a tool-call argument, or a streamed envelope that the detector never inspects. This check costs a handful of transcripts and is the only thing standing between you and a report built on a detector that was structurally unable to fire. **Step 3 — prove reach.** Was a moderation filter or a guard in front of the model? Then the clean result describes the deployed stack, and the model's own behaviour is untested — a distinction that decides where remediation would even go. Was every attempt single-turn against a product that is multi-turn? Was every prompt in one language against a product that is not? Each of those is a boundary on the claim, and each belongs in the report as a scope sentence rather than as an unspoken assumption. **Step 4 — state coverage with a named denominator.** "Zero hits" is unfalsifiable until you say *out of what*. Probes run out of probes available is one denominator; attack surface reached is a completely different one; behaviours in some reference list is a third. Mixing them is how an honest clean run becomes a misleading one — a run covering a tenth of the shipped probes, reported without its denominator, reads exactly like exhaustive coverage. **The positive control, and what it costs.** The durable defence is to include something in every scheduled run that *must* produce hits: a probe family known to land on an unprotected target, run against a deliberately unprotected local model in the same job, and fail the job when that control reports nothing. The cost is one extra small target and a modest number of extra calls per run — trivially cheap against the price of circulating a broken clean run. Without it, silence and success are the same signal. **What you do not say.** You never convert zero hits into a launch verdict. That judgment belongs to whoever owns the release. Your deliverable is: what ran, against what, with what denominator, what was verified by reading, and precisely what the number cannot say.

  • What single log observation most strongly suggests the run failed rather than the model held?
    Every attempt returning the same or near-identical short response — that is the signature of an error path, a cached reply or an auth failure, not of varied model behaviour.
  • How would you design a positive control into a recurring garak run?
    Run one probe family known to land against a deliberately unprotected local target in the same job, and fail the job if that control reports no hits.
  • The endpoint wraps model output in JSON and the run is clean. What do you suspect?
    That the detector is reading the wrapper rather than the field carrying the model's text, so real compliance was scored as a pass.

A garak run with no positive control is a smoke alarm nobody has ever pressed the test button on. Silence is equally consistent with no fire and with a dead battery, and you cannot tell which from the silence.

saying these in an interview costs you the question

  • Reporting zero hits as evidence the model is safe.
  • Never opening the raw attempt log on a clean run.
  • No positive control anywhere in the harness.
  • Failing to say which probes ran and how many attempts each produced.
  • Not noticing a filter in front of the model absorbed every prompt.

context