A garak scan configured with one generation per prompt reports a probe as passing; the identical scan rerun the next day fails it. What about the tool's sampling explains the flip, and how do you pick a repeat count that makes it unlikely?
answer
- each attempt is a Bernoulli draw
- miss probability (1-p)^N
- N=1 detects only near-deterministic failures
- solve N backwards from smallest p
- clean scan is not absence
basics
~20 sThe target samples, so a prompt that fails only sometimes is a coin flip. With one repeat you draw once: a behaviour firing on one attempt in ten is missed nine times out of ten and filed as a pass. Pick the repeat count from the rarest failure you still want to catch.
solid answer
~50 sEach attempt is an independent draw from a stochastic system — the model's decoding, and often a guardrail whose verdict also varies. If a prompt elicits the failing behaviour with probability *p* per attempt, a run with N repeats misses it entirely with probability `(1 - p)^N`. At N = 1 and p = 0.1 that is a 90% chance of a clean report on a real defect. At N = 10 it is about 35%. At N = 30 it is around 4%. So the count is chosen backwards from the smallest firing rate you care to detect and the miss probability you will tolerate — not from a round number that felt cheap. The honest caveat is that this is a per-prompt calculation, and the cost is per-scan. Multiplying every prompt in a full sweep by 30 is usually not affordable, which is why teams tier the count rather than picking one value for everything.
go deeper
Says the model is random so the same prompt can pass once and fail once, and that running it more times helps.
Gives the per-attempt probability model and the (1-p)^N miss probability, and picks N from the rate they want to detect.
Adds that N is chosen against a budget, that rare behaviours cannot be ruled out at any affordable N, and rules out configuration drift before blaming sampling.
Turns it into a stated detection floor for the programme: which firing rates the standing scan can see, and what it explicitly cannot.
**The mechanism.** A garak attempt is one prompt sent once and scored once by that probe's detectors. Each attempt is recorded on its own; nothing in the pipeline collapses the repeats of a prompt into a single per-prompt verdict, and no repeat votes on another. So the per-attempt result is a Bernoulli variable and the scan as a whole is a sampler. Why the variable is not a constant: decoding samples (temperature and nucleus settings on the generator), serving stacks introduce their own nondeterminism through batching and routing, guardrail classifiers sit at a threshold that a small token-level difference can cross, and hosted endpoints roll backend versions underneath a stable model name. Any one of these is enough to make a prompt fail on Tuesday and pass on Wednesday with no code change anywhere. **The arithmetic you should be able to do at a whiteboard.** With per-attempt firing probability *p* and N repeats of that prompt, the probability the run sees nothing is `(1 - p)^N`: ``` p = 0.10: N=1 -> 0.90 miss N=10 -> ~0.35 N=30 -> ~0.04 p = 0.02: N=1 -> 0.98 miss N=10 -> ~0.82 N=30 -> ~0.55 ``` Inverted, the count you need for a tolerated miss chance m is `N = ln(m) / ln(1 - p)`. For a 5% miss chance: about 29 repeats at p = 0.10, and about 149 at p = 0.02. Two things fall out. First, `-g 1` is not a cheap version of the scan; it is a different instrument, one that only reliably detects near-deterministic failures. Second, rare behaviours are brutally expensive to rule out — at p = 0.02 even thirty repeats leaves better-than-even odds of a clean sheet, so "the scan was clean" must never be restated as "the behaviour does not occur". **What it costs to buy that confidence.** The counts above are per prompt, and the bill is per selection. Re-running one 20-prompt probe at `-g 30` is 600 calls — minutes and cents, and you should simply do it whenever a flip matters. Running a 2,000-prompt selection at `-g 30` is 60,000 calls: at a rough 600 tokens per call that is 36M tokens and, at three requests per second sustained, most of a working day of wall clock. That gap between "cheap on one probe" and "unaffordable across the sweep" is the entire reason teams tier the count instead of picking one value for everything. **Where the number misleads.** The headline trap is the clean report. A `-g 1` sweep that comes back with no hits reads, to everyone downstream, as evidence of absence; it is nothing of the sort, and it is the same document six months later when someone cites it in a risk register. The symmetric trap is the single hit: at `-g 1` you cannot distinguish a behaviour that fires on every attempt from one that fires two percent of the time, because both appear as exactly one failed attempt. Those are different engineering problems with different owners and different fix dates, and distinguishing them is precisely what raising the count buys. A third misreading is quieter: comparing yesterday's rate with today's when the counts differed. Two rates from different sample sizes are two different measurements, and the difference between them mixes a real change with a measurement change in a proportion nobody can recover afterwards. **What to check before you call the flip sampling noise.** Configuration drift produces flips that look exactly like randomness and are not. 1. **Same probe selection and same repeat count** in both runs — read both from the runs' own recorded configurations, not from the command you think you typed. 2. **Same generator decoding settings.** A temperature change moves *p* itself. 3. **Same endpoint and model version.** A hosted model name can point at a new backend without notice, and a pinned snapshot is worth having for exactly this reason. 4. **Same detector behaviour.** A garak upgrade between the two runs can change a detector, and the flip is then in the judge, not the target. Only when those four match is "the target samples" the actual explanation. And if it is, the cheap next move is not to re-run the sweep — it is to re-run the single suspect prompt at a high count and measure its firing rate directly, then report the rate with its count attached.
- How many repeats would you need to have roughly a 95% chance of seeing a behaviour that fires on 10% of attempts?About 30. (1 - 0.1)^30 is around 0.04, so the miss chance is about 4%.
- Your budget only allows one repeat per prompt. What do you write in the report?That the run measured one draw per prompt, so it detects near-deterministic failures only, and a clean result rules out nothing intermittent. Give the count next to the rate.
- Before concluding a day-to-day flip is sampling noise, what would you rule out first?Configuration drift: different probe selection, a different repeat count, changed decoding settings, or a moved model endpoint behind the same name.
saying these in an interview costs you the question
- Calling the flip a bug in garak rather than sampling in the target.
- Treating a clean single-repeat scan as evidence the behaviour does not occur.
- Picking the repeat count as a round number with no reference to the firing rate they want to detect.
- Never checking that the two runs actually used the same probe set, repeat count and decoding settings before blaming randomness.