You have a fixed query budget against one target and must split it between running each attacker-model jailbreak thread for more turns and starting more sequential restarts from fresh seed framings. How do you decide the split, and what does each side buy?
answer
- depth exploits, restarts explore
- past the knee, restart wins
- restart clears attacker context
- stateful target favours depth
- report (cap, restarts) with every rate
basics
~20 sDepth buys refinement of one framing; restarts buy independent draws at the same goal. Since most hits land in early turns, past the knee of the turn-to-first-hit curve a restart returns more hits per call than another turn. Set the cap at that knee and put the remaining budget into restarts.
solid answer
~50 sThink of it as one budget bought in two currencies. **Depth** lets the attacker exploit a framing that is partially working — the rating is climbing, the target is softening. It is worth paying for exactly while the rating still moves. **Restarts** re-roll the opening framing and clear the attacker's context. They buy independent draws, which raises the chance one framing is the right one for this goal and turns a per-goal result into a distribution rather than a single lucky draw. The rule follows the pilot data: cap turns at the knee of the turn-to-first-hit curve, then spend the rest on restarts, because past the knee a fresh draw beats another turn. Two things push back toward depth — a stateful target where value genuinely accumulates across a conversation, and a seed set with no other framings left to try. Report the split: three restarts of twelve turns and one thread of thirty-six cost the same and are not the same experiment.
go deeper
Should see that both consume the same budget and that running one thread forever is not obviously best.
Explains depth as exploitation and restarts as exploration, and ties the split to where hits actually land in a pilot.
Adds context anchoring as the reason restarts beat a bigger cap, names the stateful-target exception, and reports the split alongside the rate.
Argues the split from what the result must support — a quotable per-goal interval and a reproducible number — not only from hits per call.
## The budget identity Fix a total call budget `B`. At three billed legs per pass — attacker, target, judge — you have `B/3` passes to tile across `goals x restarts x turns`. Every split of the same `B` is affordable; they are not the same experiment, and they do not return the same number of hits. ## What each side buys - ***Depth*** is exploitation. It lets the attacker refine a framing that is partially working: the judge's rating is climbing, the target's refusals are getting softer or shorter, and each rewrite has something real to react to. It is worth paying for exactly as long as the rating still moves. - ***Restarts*** are exploration. A restart re-rolls the opening framing for the same goal and clears the attacker's context, buying an independent draw rather than one more edit of the draw you already have. ## Why restarts usually win past the knee The per-turn probability of a first hit falls sharply after the early turns — winners land early, and the attacker's rewrites get more incremental as its context fills with its own failures — while cost per turn rises, because the transcript fed back to the attacker keeps growing. Marginal value of turn `n+1` therefore decays on both axes at once. A restart's first turns sit back at the fat part of that same curve, so past the **knee** of the turn-to-first-hit distribution a fresh thread returns more hits per call than another turn on an old one. ## What a restart actually fixes Two distinct failure modes, which is why it beats simply raising the cap. 1. The first is a **bad opening framing**: the seed phrasing put the target straight into a hard refusal, and no amount of rewriting inside that thread recovers. 2. The second is **context anchoring**: an attacker shown a long transcript of its own failures produces steadily more incremental variants of the losing line. Raising the cap addresses neither; restarting addresses both. ## When depth genuinely wins - Where the target carries **conversational state** and the attack builds on it, a fresh thread cannot reconstruct the accumulated position in its first turn, so restarting throws away real progress. - Where the **attacker model is weak**, its early candidates are poor and it needs turns to reach anything, which pushes the knee to the right. - And where the **seed set is exhausted** — no further distinct opening framings remain for the goal — a "restart" is a replay, and a replay buys nothing that depth would not. ## What the split does to the number This is where the reporting goes wrong. With one thread per goal, a per-goal result is a single yes-or-no draw and carries no information about variance; several restarts per goal give a per-goal hit fraction you can put an interval around, and make the run reproducible in the only sense available — repeat it and the number should land in the same neighbourhood. But restarts also create a **denominator** that is easy to swap by accident. A run of thirty goals at five restarts has 150 threads, and "40% success" can mean - sixty threads landed or - twelve goals landed at least once. Those are different quantities and the second is always the larger. A number quoted without saying whether the denominator is threads or goals cannot be compared with anything, and the goal-level reading rises automatically with restart count even against an unchanged target — which is exactly how a run gets reported as finding more while proving nothing new. ## Pseudo-independence A restart that replays the same opening framing with the same sampling seed is a repeat, not a draw, and averaging repeats fakes an interval that is far too narrow. Real independence needs: 1. a cleared transcript, 2. a re-rolled opening framing, 3. and a different sampling seed on the attacker. It also needs a target with no **cross-session memory**: a system that retains a user profile, a per-account safety score, or a cached moderation decision links your supposedly separate restarts, and the later ones are measuring a target the earlier ones have already changed. ## What I would check - That restarts are independent in all three senses above. - That the target has no memory or rate-based adaptation across sessions. - That the pair `(turn cap, restarts per goal)` is published with every success rate, along with which denominator the rate uses. - And that no comparison across runs silently mixes two different splits of the same total budget.
- When does depth genuinely beat restarts?When the target carries conversational state that the attack builds on, so a fresh thread cannot reconstruct the accumulated position, and when the attacker model is weak enough that its early candidates are poor.
- What makes a restart actually independent?A cleared transcript, a re-rolled opening framing rather than a replay, a different sampling seed on the attacker, and a target with no memory carried across sessions.
Depth is redrafting the same letter over and over; a restart is writing a fresh one from a different angle. Once the redrafting has stopped changing the reply you get back, the next hour is worth more spent on a new angle than on another draft.
saying these in an interview costs you the question
- Running exactly one long thread per goal and quoting a success rate from it.
- Claiming more turns strictly dominate more restarts.
- Reporting an attack success rate without the turn cap and restart count.
- Calling restarts independent while replaying the same opening framing.
- Ignoring that a stateful target may carry memory across supposedly separate restarts.