skip to content

A single garak sweep leaves you several hundred hits and a day to triage them. How do you turn hits into report findings without over- or under-counting?

level: principalimportance: should knowfreq 47%

answer

  1. tool counts attempts, report counts behaviours
  2. bucket, rank by consequence, sample n
  3. verified / artefact / ambiguous
  4. quote the read rate with the count
  5. unread is untested, not clean

basics

~20 s

Group hits by the behaviour they demonstrate rather than writing one finding per hit. Read several per group, confirm the detector was right, and keep one finding with a verified example and the count behind it. Drop groups where every sampled hit was an artefact, and state what you sampled.

solid answer

~60 s

The tool's unit is an attempt; the report's unit is a behaviour. Converting between them is the whole job. **Bucket first.** Group hits by probe and, within a probe, by the behaviour actually demonstrated. Hundreds of hits usually collapse into a handful of distinct behaviours plus a large tail of detector artefacts. **Sample, do not sweep.** Read a fixed number per bucket, classify each as verified, artefact, or ambiguous, and use that ratio to characterise the bucket. A bucket whose sample is all artefacts is not a finding; a bucket whose sample is all verified is one finding with a strong example and a count. **Budget by consequence, not by volume.** Spend the day on buckets whose behaviour would matter if real, not on the biggest bucket. Volume in a garak sweep reflects how many prompts a probe emits, not how bad the behaviour is. **Write down the sampling.** Every quoted count needs the read rate behind it. "Forty-one hits in this bucket; twelve read, eleven verified" survives challenge in a way that a bare number does not. Ambiguous buckets go in as untested, never as clean.

go deeper

for a junior

Knows hits are not findings one-for-one and that similar hits should be grouped.

for a middle

Buckets by probe and behaviour, samples per bucket, and keeps a verified transcript with each finding.

for a senior

Prioritises by consequence rather than volume, deduplicates across probes by remediation, and reports read rates and untested residue.

for a principal

Turns it into a standing triage policy — fixed sampling rules, which buckets are always read, upstream narrowing of probe selection — and keeps the launch decision with its owner.

**The counting problem, stated precisely.** garak's unit is the *attempt*: one prompt, one generated output, one detector verdict. A report's unit is the *behaviour a remediation would address*. Neither direction of naive mapping survives review. One finding per hit inflates, because a single probe emits many prompt variants and `garak --generations` multiplies each of them, so one stubborn behaviour can occupy dozens of rows — and the moment a stakeholder notices, the credibility of every other number in the document goes with it. One finding per probe deflates, because a probe can surface two genuinely different behaviours that need two different fixes, and collapsing them means one of the fixes never gets written. **A workable day, in order.** 1. **Bucket.** Group hits by probe, then split any bucket where sampled transcripts show materially different model behaviour. Several hundred hits typically collapse into a handful of distinct behaviours plus a long tail of detector artefacts. 2. **Rank by consequence if real** — deliberately ignoring bucket size. Bucket size is a function of how many prompt variants the probe emits and how eagerly its detector fires. Neither correlates with how much the behaviour would matter in production. 3. **Sample a fixed n per bucket**, top-ranked first, reading a spread rather than the head. Attempts are usually ordered by prompt variant, so the first rows of a bucket are one variant repeated and give you a falsely uniform picture. 4. **Classify each read hit** as verified, artefact, or ambiguous. 5. **Promote** buckets containing verified hits into findings: one behaviour, one verified untruncated transcript quoted as evidence, the bucket size and the read rate attached. 6. **Record the residue** honestly. Buckets you never opened are *untested*, not clean. **What it costs, and why the budget is the real constraint.** An untruncated transcript read plus a classification is two to five minutes of engineer attention. A day is therefore on the order of a hundred to a hundred and fifty reads at the outside, against several hundred hits — the arithmetic does not close, and pretending otherwise is how people end up quoting the counter. Sampling is not a shortcut you apologise for; it is the only honest way to spend a fixed budget, and the read rate is the thing that makes it auditable. **Deduplicating across probes.** The same underlying weakness frequently lands under several different probes. Merge those into one finding that lists the multiple routes with per-route counts, because a single remediation closes all of them. Merging hides nothing as long as the routes and their counts stay visible; not merging double-counts one issue and inflates the headline. **Where the number misleads — the paragraph that matters most.** Every count you publish carries three silent distortions. *First*, hit volume is an artefact of configuration: change the probe selection or `--generations` and the same system yields a different headline, so counts are never comparable across runs unless the configuration is quoted beside them. *Second*, the deduplicated finding count and the raw hit count measure different things, and a reader will happily conflate them — always label which is which. *Third*, and most damaging, a count published without its read rate implies verification that did not happen. "Forty-one hits in this bucket" is a claim about a detector. "Forty-one hits; twelve read, eleven verified" is a claim about the world, and it is the only one that survives someone opening the log behind you. The symmetric trap is the residue: an unexamined bucket omitted from the report reads to every reader as an absence of findings, which is precisely the inference you have no evidence for. **What you check before signing.** For each promoted finding: the quoted transcript is untruncated, the incriminating span was authored by the model rather than echoed from the prompt or written by an interposed guard, and the behaviour described is the behaviour a fix would target. For the document as a whole: every count carries its read rate, the untested list is present with sizes, and no sentence in it declares whether the residual risk is acceptable for launch — that decision belongs to the release owner, and a triage table is evidence for it, never a substitute. **The organisational tell.** If triage capacity is the binding constraint every single cycle, the fix is upstream of triage: narrow the probe selection per run, replace the detectors on the noisiest buckets, and set a standing rule for which buckets are always read in full and which are sampled thinly. That is a policy decision with a cost attached, not a heroics problem to be re-solved by whoever draws the report this quarter.

  • Why is bucket size a bad severity signal in a garak sweep?
    Bucket size mostly reflects how many prompt variants that probe emits and how eagerly its detector fires, neither of which correlates with how much the behaviour would matter in production.
  • Two probes surface what looks like the same weakness. One finding or two?
    One finding, listing both routes with their per-route counts, because a single remediation addresses both — splitting it would double-count the same issue.
  • How do you present buckets you had no time to open?
    As an explicit untested list with their sizes, so the reader can see the boundary of the work rather than inferring absence of findings from absence of mention.

saying these in an interview costs you the question

  • One finding per hit, or one finding per probe, with no thought about the unit.
  • Prioritising the biggest bucket because it is the biggest.
  • Quoting counts with no statement of how many were read.
  • Marking unexamined buckets as clean.
  • Reading only the first rows of a bucket, which are usually one prompt variant.

context