skip to content

In a propose-score-rewrite jailbreak loop, one thread has run fifteen turns with the scoring model's rating flat and the target repeating essentially the same refusal. Which signals tell you the thread is stalled, and what do you do with its remaining turns?

level: seniorimportance: should knowfreq 50%

answer

  1. flat rating plus duplicate candidates
  2. target refusal unchanged
  3. stop reason: no-progress vs cap vs hit
  4. restart resets attacker context
  5. check judge variance before trusting a plateau

basics

~20 s

Flat ratings, near-duplicate candidates from the attacker, and an unchanging refusal from the target together say the thread has converged on nothing. Stop it early, keep the transcript, and return the remaining turns to the pool for a fresh restart from a different seed framing rather than more rewrites of a losing line.

solid answer

~50 s

Three signals, and you want them together rather than any one alone: - **Rating plateau** — the scoring model's number has not improved over a window of turns. Alone this is weak: loops often sit flat and then jump. - **Candidate similarity** — the attacker's recent proposals are near-duplicates by any cheap text-similarity measure. This is the strongest signal: it is editing wording, not exploring. - **Response stability** — the target keeps returning the same refusal shape, so the loop gets no new information to react to. When all three hold, kill the thread and restart from a fresh seed framing rather than continuing. Log the stop reason as *no-progress*, distinct from *cap reached* and *hit*, because those outcomes mean different things when you later tune the cap. The risk is cutting a thread that was one framing away, so keep the window generous — never shorter than the attacker needs to try a genuinely different approach.

go deeper

for a junior

Should notice that the loop is repeating itself and that continuing costs calls for nothing.

for a middle

Names the signals — flat rating, near-duplicate candidates, unchanging response — and stops the thread early rather than riding out the cap.

for a senior

Calibrates the no-progress window from pilot data, records a distinct stop reason, restarts to reset attacker context, and knows a saturated judge fakes a plateau.

for a principal

Makes stop reasons part of the run's reported metadata so that success rates can be read against how many threads ended each way.

**Why stalls, not caps, are where the money goes.** A turn cap bounds the worst case. It does nothing about the ordinary case, which is a thread that stopped learning anything at turn four and then rides out its remaining allowance in silence. The arithmetic is blunt: a thread that converged at turn four but runs to a cap of twenty wastes sixteen passes, and at three billed legs per pass — attacker, target, judge — that is 48 calls. Across sixty threads in a run, most of which stall, that is comfortably over half the run's spend buying near-duplicate paraphrases. A no-progress stop returns those turns to the pool, where a restart on a different seed framing can use them. **The three signals, and why you want them together.** Per turn, compute three things. - **Rating plateau** — the judge's score has not improved over a window of turns. On its own this is weak evidence: loops genuinely sit flat and then jump when the attacker finds a new framing. - **Candidate similarity** — the attacker's last few proposals are near-duplicates of one another by any cheap lexical similarity measure. This is the strongest of the three, because it says directly that the loop is editing wording rather than exploring. - **Response stability** — the target keeps returning the same refusal shape, so there is no new information for the attacker to react to. Stall when the rating has been flat across the window *and* both similarities are high. The similarity measure does not need to be clever; it is a control signal for a scheduler, not evidence in a report, so a token-overlap score is fine and an embedding model is an unnecessary fourth billed leg. **Calibrating the window.** Too short and you cut threads mid-pivot, which shows up later as a lower measured success rate with no visible cause. Calibrate from pilot data rather than intuition: for threads that eventually succeeded, find the longest run of flat, near-duplicate turns that preceded the hit, and set the window comfortably above the bulk of that distribution. Then treat the window as a versioned parameter — it belongs in the run metadata next to the cap, because a rate measured with an aggressive stop is not comparable with one measured without it. **Restart rather than continue.** When a thread stalls, the useful move is a restart with a genuinely different opening framing for the same goal, not a nudge to the attacker. A restart clears the attacker's context, and that context is frequently the actual problem: a long transcript of failures anchors the attacker to the losing line it has been polishing. Restarts also give independent draws at the same goal, which is what turns a per-goal result into something more than one lucky or unlucky sample. **Where these signals lie to you.** A saturated judge is the big one. If the scoring model returns the same value for everything — every attempt at the floor, or every attempt at some indifferent middle — then "the rating has not improved" is trivially true on turn two for every thread in the run, and the loop will stall everything at once. Check the rating's variance before you trust a plateau; a stall rate that jumped after a scoring change is a fact about the scorer, not the target. Greedy or near-deterministic decoding on the attacker makes near-duplicate candidates the normal case rather than a symptom, so either raise the attacker's sampling temperature or drop the similarity signal. And a target that has started rate-limiting returns errors that look like beautifully stable refusals, so exclude non-answers — HTTP errors, truncations, empty completions — from the similarity computation, or a quota problem will present as universal target robustness. **How the numbers move when you add the stop.** Expect the measured success rate to fall slightly — some cut threads would have landed — and the mean turns per thread to fall a lot. Reporting only that "cost per hit improved" folds those two changes together and hides the first. Report the stop-reason breakdown instead: how many threads ended as *hit*, how many as *no-progress*, how many as *cap reached*. Those three counts are what let a later reader tell a budget-bound run from a traction-bound one. **What I would keep.** Every stalled transcript, tagged with its stop reason and the window that fired. Stalls cluster by goal and by opening framing, and that clustering is often the most useful thing a run produces on a day when nothing lands.

  • Every thread stalls at turn three. What do you suspect first?
    The scoring model, not the attacker. If the rating has almost no variance — everything scored the same — the plateau condition is always true. Check the rating distribution before touching the loop.
  • Why restart rather than simply continue with a nudge to the attacker?
    The accumulated transcript of failures is usually what is anchoring the attacker to a losing framing. A restart clears that context and yields an independent sample of the same goal.

saying these in an interview costs you the question

  • Stalling on a flat rating alone, with no similarity check.
  • Letting every thread run to the cap because 'it might still land'.
  • Recording a no-progress stop and a cap-reached stop as the same outcome.
  • Not noticing that the scoring model returns the same value for everything.
  • Treating rate-limit errors from the target as stable refusals.

context