You must choose the maximum number of turns for an attacker-model jailbreak loop (an attacker proposes a prompt, the target answers, a scoring model rates the answer). How do you pick that cap from evidence rather than guessing, and why do extra turns buy less and less?
answer
- turn-to-first-hit distribution
- cumulative hits, find the knee
- late turns drift to cosmetic rewrites
- report success rate at cap N
- deep control arm
basics
~20 sRun a pilot with a generous cap and record, for every thread that succeeded, the turn on which it first succeeded. Most hits land early and the tail is thin. Put the cap where new hits stop accruing, then spend the freed calls on more restarts and more seed goals instead of deeper threads.
solid answer
~50 sMeasure, do not guess. In a pilot with a deliberately high cap, log the turn index of each thread's first success. Plotting cumulative hits against turn index gives a curve that rises steeply and then flattens; the cap belongs near the knee, plus a small margin. The flattening has a cause. A thread that has not landed by turn ten is usually not a slow winner — it is a thread whose seed goal or opening framing the target refuses cleanly, and the attacker keeps rewriting around a wall. Later turns also drift: with a long transcript of failures in context, the attacker often produces cosmetic variants rather than genuinely new framings. The cap is per target, per seed set and per attacker model, so re-measure when any of those change. And record the cap alongside every result, because a hit at turn twenty-eight is not comparable with a run capped at ten.
go deeper
Should at least know a cap exists to bound cost and that you can look at when successes actually happened.
Expected to describe the turn-to-first-hit pilot, the flattening curve, and reinvesting the freed calls in restarts and seeds.
Explains why returns fall — selection plus context-driven degradation — keeps a deep control arm, and refuses cross-run comparisons at different caps.
Treats the cap as a versioned parameter of the measurement, re-derived when the target, attacker or seed set changes, and insists it travels with every reported number.
## What you are actually measuring The quantity that decides a turn cap is the *turn-to-first-hit* distribution: for every thread in a pilot, the turn index at which the judge — the scoring model that rates whether the target's answer met the goal — first called that answer a success. Threads that never succeed contribute no index at all. They are **censored**, and forgetting that is the first mistake: the mean first-hit turn over successful threads tells you where winners land, not how likely a win is. The two questions are separate, and only the first one sets a cap. ## Running the pilot, and what it costs 1. Take a subset of seed goals, 2. set a cap far above what you intend to ship — deep enough that the tail is visible, typically three to four times your guess — 3. and run it once. Fifteen goals, two restarts each, a cap of fifty, three legs per pass is around 4,500 calls if every thread runs full depth, less in practice because hits terminate threads early. That is a real spend, but it is spent once per target rather than once per run, and it replaces a round number chosen in a meeting. ## Reading the curve Plot the cumulative fraction of eventual hits found by turn `n`. It rises steeply over the first handful of turns, bends, and then crawls. Put the production cap just past the **bend**, with a margin for pilot noise. The margin matters more than the bend does: if the whole curve rests on eighteen successful threads it is a step function, not a curve, and reading a precise knee off eighteen points is the most common way this exercise goes wrong. If you cannot see a bend, you did not have enough hits, not a flat world. ## Why the curve bends Two effects, and they compound. - **Selection:** threads that were ever going to work usually work early, because the attacker's first rewrites make the largest changes in framing and the later ones make the smallest. - **Degradation:** as the transcript of its own failures fills the attacker's context, it starts producing cosmetic variants of a losing framing rather than abandoning it. So marginal hit probability per turn falls while cost per turn rises with the growing transcript — the two curves move against each other, which is why the cap has a defensible location at all. ## How the resulting number misleads - The success rate you publish is censored by the cap you chose, so it is a rate *at cap N* and nothing else. Comparing your run at cap ten against a run at cap forty says nothing about the two targets; comparing it against a published figure is worse still, because that figure's judge, threshold, seed set and attacker all differ too and none of those is usually stated. - The subtler bias is in the pilot itself: a knee measured on the goals you already know work sits too far left, so the shipped cap is too shallow for exactly the hard goals anyone cares about — and the loss is invisible in the aggregate, because it reads as "those goals are simply resistant" rather than "those goals were cut off". - Finally, the knee is a property of the whole quadruple of attacker, target, seed set and judge configuration. Swap any one and the number you derived no longer applies. ## Complements to the cap A hard cap bounds the worst case but does nothing about a thread that stopped learning at turn four and rides out fifteen more. Pair the cap with a **no-progress stop** so converged threads return their remaining turns to the pool, and spend the freed calls on more restarts and more seed goals rather than on depth — past the knee, a fresh draw returns more hits per call than another turn. ## What I would check - That the pilot's seed goals are representative rather than the ones already known to land. - How many successful threads the knee actually rests on. - That the cap is written into the run metadata and travels with every quoted rate, so no one downstream can compare it against a number taken at a different depth. - That a small **deep control arm** — a slice of goals run at several times the shipped cap — is refreshed whenever the target build changes, because a target update moves the knee and a shipped cap that is quietly now too shallow produces a falling success rate that looks like a security improvement.
- You swap in a noticeably stronger attacker model. What do you expect to happen to the right cap?Hits generally arrive earlier, so the knee moves left and the cap can come down. The calls that frees are better spent on more restarts and broader seed coverage than on depth.
- Your pilot shows a thin but real tail of hits at turn thirty-plus. Do you raise the cap?Usually not for the main run — a thin tail costs full depth on every thread. Keep the low cap and run a small deep arm on the goals most likely to need it, then report both.
saying these in an interview costs you the question
- Picking a round number with no pilot behind it.
- Assuming a deeper cap is always better because more search can only help.
- Comparing an attack success rate against a published or historical number taken at a different cap.
- Never re-measuring after swapping the attacker model.