skip to content

When a garak scan repeats each prompt several times, what does the per-probe pass/failure rate in its report count over, and how do a probe that fails on every attempt and one that fails occasionally look different in that report?

level: middleimportance: should knowfreq 48%

answer

  1. denominator = prompts x repeats
  2. N=1 quantises the rate
  3. always-fails vs sometimes-fails collapse at N=1
  4. reconcile attempt totals, dropped != passed
  5. rate + attempts + N travel together

basics

~20 s

Attempts, not prompts. With the repeat count at N, each prompt yields N scored attempts and the probe's rate is over that larger total. A probe failing every time sits at the extreme of the range; one failing occasionally shows a small non-zero rate that a single-repeat run would have shown as zero.

solid answer

~50 s

The report's unit is the attempt — prompt times repeat — so the denominator grows with the repeat count while the number of distinct prompts stays fixed. That is what gives the rate any resolution at all. At a repeat count of 1 the per-probe rate is quantised to multiples of 1/(number of prompts), and every failing prompt looks equally bad because each contributes exactly one failed attempt. Raise the count and the two cases separate: the always-failing prompt contributes N failures, the occasional one contributes a fraction of N, and the probe's rate now carries information about *how often*, not just *whether*. The trap is comparison. A rate from a run at one repeat count is not comparable to the same rate at another — same number, different sample size, different noise. Always carry the repeat count and the attempt total alongside the rate, and check them against prompts times repeats to catch attempts that errored out and were dropped.

go deeper

for a junior

Knows the report gives a per-probe pass/failure rate and that repeats produce more scored attempts.

for a middle

States the denominator is prompts times repeats, and explains why a single-repeat run cannot separate always-fails from sometimes-fails.

for a senior

Reconciles attempt totals to catch dropped attempts, refuses to compare rates across different repeat counts, and drops to the per-prompt breakdown to classify a gap.

for a principal

Standardises the reported triple — rate, attempt total, repeat count — so numbers from different teams and runs mean the same thing.

**The denominator.** garak scores per attempt, and an attempt is one prompt sent once to the generator and judged once by a detector. With P prompts in a probe and a repeat count of N, a clean run produces P x N attempts, and the probe's pass/failure rate in the report is the share of those attempts the detector flagged. The distinct-prompt count is not the denominator, and no step in the pipeline collapses a prompt's N generations into one per-prompt verdict first. This stops being a pedantic distinction the moment N is not 1, because from then on the number of rows behind the rate and the number of prompts behind the run are different quantities that people habitually conflate. **What the extra denominator buys: resolution.** Take a probe with 20 prompts, one of which reliably elicits the behaviour and one of which does so about a fifth of the time. ``` N = 1 -> ~1 or 2 failed attempts out of 20 (rate 5-10%, coarse and noisy) N = 20 -> ~20 + ~4 failed out of 400 (rate ~6%, and the per-prompt pattern is now visible) ``` The aggregate rate barely moved. What changed is that at N = 20 you can look underneath it and see one prompt at 100% and another at roughly 20% — two completely different engineering problems. The always-failing prompt is a deterministic gap you can reproduce on demand, hand to an owner, and regression-test. The intermittent one is a rate question whose fix is judged statistically. At N = 1 both are flattened into "one failed attempt", and the report gives you no way to tell them apart. Note also what N = 1 does to the granularity of the rate itself: with 20 prompts it can only take the values 0, 5%, 10% and so on, so small real differences between targets vanish into the quantisation. **Where the number misleads.** Three specific misreadings, in rough order of how often they cause damage. - **Cross-run comparison at different counts.** Two scans reporting 6% mean different things if one drew 20 attempts and the other 400. The point estimate is the same; the interval around it is not remotely the same. Comparing them as equals — a baseline against a retest, one team's number against another's — is the most common reading error in this whole area, and nothing in the report warns you about it. - **Dropped attempts shrinking the denominator silently.** Timeouts, transport errors, throttled calls and abandoned retries can leave the scored total below P x N. A missing attempt is not a passing attempt, but arithmetic treats it as if it never existed, which quietly flatters the rate. If the totals do not reconcile, the denominator is wrong and so is every rate computed from it. - **Hit volume read as severity.** Repeats multiply false positives exactly as they multiply true ones. A detector with a modest false-positive rate produces N times its spurious hits, so a long hit list at a high count is a statement about N and the detector, not about how bad the target is. **What it costs to have the resolution.** The denominator is bought with queries: taking that 20-prompt probe from N = 1 to N = 20 is 20 calls becoming 400. That is negligible for one probe and decisive across a sweep — a 2,000-prompt selection at N = 20 is 40,000 calls, tens of millions of tokens, and hours of wall clock against a rate-limited endpoint. This is why the resolution is usually bought per probe rather than globally. **What to do.** Report the triple — rate, attempt total, repeat count — and never the rate alone; a bare percentage from this tool is uninterpretable. Reconcile the scored total against P x N before quoting anything, and account for the shortfall explicitly as errors rather than folding it in. And when the question is "is this a hard gap or a flaky one", do not read the probe-level rate at all: drop to the per-prompt attempt breakdown, where a prompt failing N times out of N and a prompt failing 3 times out of 20 are visibly different findings.

  • The report shows fewer attempts than prompts times the repeat count. What do you suspect?
    Attempts that errored, timed out or were dropped by the transport. They are missing data, not passes, and they shrink the denominator silently.
  • How do you tell a hard, reproducible gap from a flaky one in the same probe?
    Look at the per-prompt attempt breakdown, not the probe-level rate: one prompt failing on all N attempts is deterministic, one failing on a few is rate-limited behaviour.

Rating a probe over prompts rather than attempts is like grading a class by how many students ever failed a quiz instead of how many quizzes were failed: the first number cannot tell you whether one student failed every single time or twenty students each slipped once.

saying these in an interview costs you the question

  • Assuming the rate is computed over distinct prompts regardless of the repeat count.
  • Comparing rates from runs with different repeat counts as though they were the same measurement.
  • Treating a lower attempt total than prompts times repeats as harmless.
  • Reading a high hit count as high severity when it is just the repeat count multiplying one flaky prompt.

context