You rerun a garak probe that builds its prompts during the run against the same endpoint after shipping a mitigation, and the failure count drops. How do you establish whether the mitigation worked rather than the probe simply having sent different prompts?
answer
- generated prompts = fresh sample each run
- replay saved failing prompts = the regression test
- n reruns per build, compare spreads
- pin selection, attempts, package, endpoint
- report range, not a point
basics
~20 sA probe that mints prompts each run does not send the same prompts twice, so one lower count is a different sample, not a result. Replay the specific saved prompts that failed before against the patched system, and rerun the generating probe several times on both builds before claiming an improvement.
solid answer
~50 sSplit the question in two, because one run answers neither half. **Did the specific defect close?** That is a replay, not a scan. Take the concrete prompts that produced hits in the earlier run out of its logs, send exactly those at the patched build, and check the same ruling. This is deterministic and is what you attach to the ticket. **Did the family get harder?** That needs repetition. Run the generating probe several times against the old build and several against the new one, and compare the distributions rather than two single numbers. If the spread across reruns of one build overlaps the gap between builds, you have measured noise. **Also rule out the boring explanations** before claiming credit: the probe selection or attempt count differed, the endpoint or its configuration moved underneath you, or the generation step itself depends on something that changed. Pin what you can pin, and say in the write-up which half of the claim each piece of evidence supports.
go deeper
Recognises that the probe did not send the same prompts twice, so the two counts are not directly comparable.
Proposes replaying the saved failing prompts and repeating the generating probe, and knows that one before/after pair proves nothing.
Runs both tracks, controls selection, attempt count, package version and endpoint configuration, and reports spreads with the specific closed defect called out separately.
Sets the standard that a non-deterministic probe never gates a release on its own, and that every finding graduates into a pinned deterministic regression case.
**Why the naive comparison fails.** A garak probe that mints prompts during the run produces a fresh sample each time it executes. Comparing run A on the old build with run B on the new one therefore changes two things at once: the build *and* the prompt sample. The report offers no field that separates them, so a count that halved is entirely consistent with an unchanged system that happened to be asked easier questions. This is not a subtle statistical point — it is the ordinary case, and it is why "the number went down after we shipped the fix" is the single most common unsound claim in a scanner-driven remediation review. **The two-track method.** *Track 1 — regression on the concrete defect.* Pull the exact prompts that produced hits out of the earlier run's attempt records and send precisely those at the patched build, with the same detector ruling. Fixed inputs make this deterministic and cheap — a handful of calls, seconds of wall-clock, and no generative machinery — and it is the only evidence that answers "is the bug I filed gone?". Keep those prompts as a permanent pinned case rather than hoping the generator rediscovers them; rediscovery is probabilistic, and a generator that does not sample that region again will read as a fix. *Track 2 — did the family get harder?* This needs repetition, not a rerun. Execute the generating probe n times against the old build and n times against the new one, holding the probe selection, the attempt or turn budget, `garak --generations` and the package version fixed, and compare the two *distributions*. Report a range, not a point. If the spread within one build overlaps the gap between builds, you have measured noise. The cost is the honest constraint here: each replicate spends the full generative budget — target calls plus attack-model calls at every turn or search node — so n = 5 on a family that costs a few thousand calls per run is a five-figure call count and a nontrivial bill. If you cannot afford enough replicates to see the spread, say so and call the trend unproven rather than quoting one pair as a trend. **Confounders to eliminate before claiming credit.** - The probe selection or the attempt/generations budget differed between the runs — the two numbers then are not even the same measurement. - The endpoint moved: a different deployed configuration, a guardrail toggled, a model version rotated behind the same URL, or simply non-deterministic decoding at the target. - The generating machinery's own inputs changed — a different attack model, judge, or temperature is a change in the *instrument*, and its output moving tells you nothing about the system. - The garak package was upgraded between the runs. This one invalidates the comparison outright rather than merely weakening it, because both the probe's prompts and the detector's ruling can change with the version; the same reply can be ruled differently and the same probe name can mean a different prompt set. - Errors were silently counted as passes: throttling, timeouts or auth failures produce attempts with no usable output. Check the error/failed-attempt count in the report before believing a drop, because a target that answered fewer times looks safer. **How to report it.** "The three prompts that produced hits in the earlier run no longer produce them at the patched build, same detector, same package version. Across five reruns of the generating probe per build, failure counts fell from a range of X-Y to a range of Z-W, with selection, budget and endpoint held fixed." That separates the closed defect from the claimed trend, and a reviewer can weigh each independently. A single before-and-after pair from a non-deterministic probe cannot be weighed at all — and presenting it as a fix is how a defect that is still live gets closed on paper.
- Why keep the specific failing prompts as a pinned case instead of relying on the generating probe to find the issue again?Because rediscovery is probabilistic. A pinned prompt is a deterministic regression test that fails loudly if the mitigation regresses; the generator may simply not sample that region again.
- How many reruns are enough to claim the family got harder?Enough that the spread within a build is visibly smaller than the gap between builds. There is no fixed number — if you cannot afford that many runs, report the replay result and call the trend unproven.
Comparing one pre-fix run with one post-fix run of a generating probe is like weighing yourself before and after a diet on two different scales. Something moved; you cannot tell which of the two things it was.
saying these in an interview costs you the question
- Declares the mitigation effective from a single before/after pair.
- Has no saved prompts from the original failing run to replay.
- Changes the probe selection or attempt count between the two runs and still compares them.
- Ignores that a package upgrade between runs can change how replies are ruled.
- Quotes a point estimate from a probe with visible run-to-run variance.