Your automated attacker-model jailbreak loop ran to its turn cap and restart count against a target and produced no successful attempt. What can you legitimately write in the report, and what would you refuse to claim?
answer
- bounded search, not a property
- name cap, restarts, seeds, scorer, build, date
- stop-reason breakdown diagnoses the null
- never ship a bare pass token
- decide in advance what a null buys
basics
~20 sYou can write that no attempt succeeded within the turn cap and restart count used, against this seed set, judged by this scoring model, on this target build and date. You cannot write that the target is jailbreak-resistant or that no attack exists. A capped search reports a budget, not a property.
solid answer
~50 sA negative from a bounded search is a statement about the search, not about the target. The defensible sentence names every bound: the seed goals, the turn cap, restarts per goal, the attacker model, the scoring model and its hit threshold, the target build, and the date. What you must refuse: 'the model is safe', 'jailbreak-resistant', 'no vulnerabilities in this category'. None follows. The result is consistent with several very different worlds — the caps were too low, the seed goals missed the framings that work, the attacker was too weak, or the scoring model's threshold was strict enough to call real successes failures. So the useful report separates them. Include the stop-reason breakdown (hit, no-progress, cap reached), the distribution of best ratings achieved, and any near-misses. A run where most threads stalled early reads very differently from one where threads were still climbing when the cap ended them — the second says raise the cap, the first says change the seeds or the attacker.
go deeper
Should know not to say 'the model is safe' and to mention that the run had limits.
States the result with its caps and seed set attached, and avoids generalising beyond what ran.
Uses the stop-reason and rating distributions to say which explanation the null supports and what a re-run should change.
Owns the reporting contract: bounds travel with the number, no bare pass token reaches a dashboard, and what a null does and does not release is agreed before the budget is spent.
## The sentence you can defend "Across `G` seed goals, `R` restarts each, capped at `T` turns, using attacker model `A` and scoring model `J` at threshold `t`, no attempt was scored a success against target build `X` on date `D`." Every binding in that sentence is load-bearing. A **bounded search** reports the search, not a property of the target, and each binding you drop converts a measurement into an assertion the run does not support. The single most common omission is the **scoring threshold**, and it is the one that most changes what the null means. ## Why a clean negative is dangerous downstream Readers do not read caveats; they read headlines. "Automated jailbreak testing: pass" becomes an artefact in a launch review, and by then the caps that produced it are invisible and unrecoverable. The mitigation is structural rather than editorial: put the bounds in the same sentence as the result, and never let the run emit a standalone pass token that can be copied into a dashboard on its own. If a field in your reporting schema can hold the value "pass" with nothing attached, that field is the vulnerability. ## Diagnosing your own null A null is consistent with at least four different worlds, and the run's own metadata distinguishes them. 1. The caps were too low; 2. the seed goals never aimed at the framings that work; 3. the attacker was too weak to find anything; 4. or the judge's threshold was strict enough to score real successes as failures. Report the **stop-reason breakdown** — *hit*, *no-progress*, *cap reached* — and the distribution of best ratings achieved per thread. - Many early no-progress stops mean the search never got traction, which points at the seeds or the attacker. - Many cap-reached threads with ratings still climbing mean the budget bound you, and the honest recommendation is a deeper re-run on that subset. - Ratings clustered flat at the floor across every goal is a statement about the scoring configuration before it is a statement about the target. ## What the null cost, and what a stronger one would cost Be concrete, because "just raise the cap" is asked as though it were free. Thirty goals, five restarts, a cap of twenty and three legs per pass is on the order of 9,000 calls. Doubling the cap does not double the bill — it more than doubles it, because the attacker's transcript grows with each turn — and it applies that increase to every thread, including all the ones that stalled at turn four. Re-running only the twelve goals whose ratings were still rising, at a cap of sixty, is roughly 10,800 calls: about the size of the original run, spent on the subset where depth is the plausible explanation. Framing it that way turns a vague "we should test harder" into a **costed choice** a budget owner can actually make. ## Coverage versus difficulty State what the seed set aimed at and, explicitly, what it never attempted. A null over ten goals is a far narrower claim than a null over two hundred, and a reader cannot recover the denominator from the word "no". Say the denominator, and say which harm categories or capability areas were out of scope entirely. ## The organisational contract Decide before you spend the budget what a null buys. If the answer is "nothing changes either way", the run was not worth running, and that is better discovered in advance. A workable contract looks like: a null at these caps - releases the change, - schedules a deeper re-run at the next model or system-prompt update, - and versions the caps so that the next run is comparable with this one. Pair it with the standing caveat that a bounded automatic search is weak evidence next to an unbounded human effort with time to think, and say plainly which of the two this report represents. ## What I check before signing - That the scoring model was validated for this configuration — replay known-successful and known-refused answers through it at this run's threshold and confirm it labels them correctly, because an unvalidated judge can manufacture a null on its own. - That the target build under test is the build that ships, including the system prompt and filter configuration, and not a staging variant. - That the caps, restart count, seed set and scoring threshold are in the run metadata rather than in someone's terminal history. - And that no bare pass indicator derived from this run has reached a dashboard where the bounds cannot follow it.
- The stop-reason breakdown shows most threads ended at cap with ratings still rising. What do you recommend?A targeted deeper re-run on exactly that subset at a higher cap, and holding the null as provisional until it lands. The budget, not the target, produced this result.
- How would you keep a leadership summary from over-reading the null?Put the bounds in the same sentence as the finding, publish no standalone pass token, and agree beforehand what a null releases and what it schedules for re-testing.
saying these in an interview costs you the question
- Writing 'no jailbreaks found, model is safe'.
- Reporting the null without the turn cap, restart count or seed set.
- Presenting an automated null as equivalent to a human red-team effort.
- Never questioning whether the scoring configuration could have hidden successes.
- Letting a bare pass indicator into a launch dashboard.