skip to content

An automated jailbreak search has spent two-thirds of its query allowance against a metered endpoint and has produced no new distinct successful template for the last fifth of those queries, while judged successes keep arriving. How do you decide between stopping, reseeding, and letting it run?

level: seniorimportance: should knowfreq 52%

answer

  1. plateau, judge lock-in, or search collapse
  2. no-new-cluster window set before the run
  3. trip once diagnose, trip twice redirect
  4. reseed is not re-choosing the method
  5. record curve, threshold, last new template

basics

~20 s

Stop paying for repeats. First check the flat stretch is a real plateau, not a stalled judge or a collapsed search. If it is real, spend the remaining allowance on a different seed set or objective rather than more of the same, and record the queries spent and the last new template so the stop is evidence.

solid answer

~50 s

Three questions, in order. **Is the plateau real?** A flat distinct-template curve with successes still arriving is the exploitation signature, but it is also what a scoring step stuck on one surface cue looks like, and what a search whose population has collapsed to a single lineage looks like. Sample recent successes and check both before believing the curve. **Is it long enough to act on?** New templates arrive in bursts, so a few hundred flat queries prove nothing. The no-new-cluster window should have been fixed before the run and scaled to how sparsely discoveries arrived earlier in that same run. **What is the remaining allowance worth elsewhere?** The choice is rarely stop versus continue; it is continue here versus reseed with a different objective, corpus or conversational surface. An exhausted search has near-zero marginal value. Whichever you pick, record the queries spent, the curve and the last new template, so the stop is evidence rather than a hunch.

go deeper

for a junior

Says the run should stop or change when it stops finding new things, and that continued successes alone are not a reason to keep spending.

for a middle

Names the exploitation signature, proposes a no-new-cluster window fixed in advance, and prefers reseeding over letting the allowance drain.

for a senior

Separates genuine exploitation from judge lock-in and search collapse, diagnoses each from evidence, and treats the remaining allowance as capital with an opportunity cost.

for a principal

Makes the stopping rule a programme standard so campaigns are comparable, and ensures the record distinguishes an exhausted search from an exhausted budget.

## The decision is economic, not emotional The failure this question hunts for is treating the stop as a feeling. Operators watch a dashboard, see judged successes rolling in, and let a metered run drain because stopping while it is "still working" feels premature. The discipline is to write the stopping rule into the run plan before any query is spent, then apply it as written. **A workable rule.** Pick a *no-new-cluster window* expressed in queries — proportional to the run rather than an absolute number — and decide in advance what happens when it trips. Trip once: sample and diagnose. Trip twice consecutively: redirect. Publishing the window up front removes the mid-run temptation to extend it because the curve might turn. With no prior campaign to calibrate against, scale the window off the discovery spacing observed early in the same run, for example a multiple of the largest gap between new clusters so far, and write the multiple down so it cannot be quietly grown later. ## What the remaining third is worth Size it in the units you are billed in before deciding. A third of the allowance is a third of the queries, and in a judged attacker-model loop each attempt is roughly three billed model calls — attacker, target, judge — multiplied again by turn depth in a multi-turn loop. So the question is never "shall I let it finish"; it is "what is the best use of that specific pile of money and rate-limited wall-clock". A confirmed exhausted search has near-zero marginal value, which means almost any redirection dominates continuing. That is a strong claim and it is why the diagnosis below matters: it must be a *confirmed* exhausted search. There is a human cost on the other side of the ledger too. Reseeding is not free — new seed corpora, a restated objective, a re-calibrated scoring step — and if the remaining allowance is small, the setup hours can exceed what the queries are worth. Stopping outright is then the honest answer. ## Three causes that look identical on the curve | cause | lineages of recent successes | accepted responses | correct action | |---|---|---|---| | Genuine exploitation | descend from the known hit, a small set of parents | vary in wording, and a hand read confirms the behaviour is really there | redirect the remaining allowance | | Judge lock-in | span several unrelated lineages | suspiciously uniform in shape, and a hand read finds no actual behaviour | re-score a sample by hand or with a second method before any stop decision | | Search collapse | all descend from a single parent; population diversity is gone | ordinary, nothing anomalous | widen mutation or restart the population, then continue | Judge lock-in means the scoring step began accepting on a proxy cue that correlates with the behaviour under test without being it. It is the dangerous case because it corrupts the numerator of every metric above it, and a saturation reading is only ever as good as the judge beneath it. Search collapse means the population lost diversity: mutation too conservative, one parent dominating selection, or a generating step sampling too narrowly. It is the one case where continuing after a change is right, because the flat curve describes the search's own degeneracy rather than the target. ## Reseeding is not re-choosing the method Redirecting the remaining allowance means new seeds, a different target behaviour, or a different conversational surface, run with the *same* search. Switching to a different class of search — a gradient-based optimiser instead of an attacker-model loop, say — is a decision about the access and compute you actually hold, made when the campaign is planned. It is not a reaction to a flat curve, and treating it as one usually means discovering mid-run that you do not have the weights or the GPU hours the new class assumes. ## Where the stop gets reported wrongly The single worst sentence is "the target is resistant". A stop says *this search stopped discovering, with these seeds, this objective, this scoring step and this clustering rule, after N queries*. Nothing about the target's remaining weaknesses follows from it. The second worst is a record that omits the queries spent, so a later reader cannot tell an exhausted search from an exhausted budget — two facts with opposite implications for whether to fund another campaign. ## What to record Queries spent and the billed-call multiplier behind them, the distinct-template curve, the clustering rule and threshold with their version, the last new template and the query index at which it arrived, the diagnosis that was made, and the reason for the stop. That record is what turns a judgement call into evidence, and it is the only thing that makes two campaigns comparable later.

  • Recent successes span several lineages but the accepted responses all look near-identical. What do you check first?
    The judge. Diverse prompts converging on one accepted response shape is judge lock-in, not discovery. Re-score a sample by hand or with a second method before trusting the curve at all.
  • How do you set the no-new-cluster window without a prior run to calibrate against?
    Scale it off the discovery spacing observed early in the same run — a multiple of the largest gap between new clusters so far — and write the multiple into the run plan so it cannot be quietly extended later.

saying these in an interview costs you the question

  • Letting a metered run continue because judged successes are still arriving.
  • Extending the stopping window mid-run because the curve might turn.
  • Reading a flat curve as saturation without checking whether the judge or the search population is the cause.
  • Reporting the stop as evidence the target has no further weaknesses.
  • Having no record of queries spent or the clustering rule, so the stop cannot be defended or reproduced.

context