skip to content

A branching attacker-model jailbreak search runs against a metered chat endpoint with a fixed query allowance. It exhausts the allowance while only the first third of your seed behaviours have been searched at all; the rest were never attempted. How do you diagnose that, and what do you change?

level: seniorimportance: should knowfreq 33%

answer

  1. global allowance + sequential seeds = prefix coverage
  2. per-seed cap plus small reserve
  3. instrument queries per seed per level
  4. interleave levels for uniform coverage
  5. early stop on hit and on stall

basics

~20 s

The allowance is global and the search walks seeds one at a time, so early seeds spend everything. Confirm it by counting queries per seed and per level. Fix it by dividing the allowance into a per-seed cap, capping frontier width, and stopping a seed early once it hits or clearly stalls.

solid answer

~60 s

Diagnose first: emit a counter per seed behaviour and per level — queries sent, candidates expanded, level reached. Three shapes are common. **Sequential drain**, where seeds are processed in order under one global allowance and the early ones consume it. **Width blow-up**, where the frontier cap is not actually enforced so one seed's tree grows exponentially. **Constant inflation**, where retries on transport errors, malformed attacker output or repeated conversation carry-over multiply the per-node cost far above the estimate you sized with. The fix is to make the allowance a resource that is divided rather than consumed first-come. Give each seed its own cap, derived from allowance divided by seed count, with a small shared reserve for seeds that look promising. Enforce the width cap. Add early termination: stop a seed on a confirmed hit, and stop it when its best score has not improved for a set number of levels. Then run seeds in an interleaved order so a truncated run still covers the whole seed set shallowly instead of a third of it deeply.

go deeper

for a junior

Notices the run stopped early and that later seed behaviours were never tried.

for a middle

Identifies a global allowance consumed sequentially and proposes a per-seed cap.

for a senior

Instruments queries per seed and per level, distinguishes sequential drain from an unenforced width cap from an inflated per-node constant, and adds early stop, stall detection and interleaving.

for a principal

Sets the policy that a partial run reports its own coverage denominator, and decides how the allowance is divided across seeds, teams and engagements before the run starts.

### What actually went wrong Nothing here is a budget failure; it is a scheduling failure wearing a budget failure's clothes. The loop is **depth-committed**: it takes seed behaviour one, spends whatever that tree costs, then moves to seed two, and so on, decrementing one global counter. Under a fixed allowance the guaranteed outcome is *prefix coverage* — the seeds at the front of the list get everything, the tail gets nothing, and the report describes a third of the intended scope while looking like a completed run, because the allowance was consumed exactly to zero. ### Diagnose it with counters, not reasoning Instrument, per seed behaviour: queries sent, candidates expanded, deepest level reached, and a terminal reason (confirmed hit, stalled, per-seed cap, allowance exhausted, error). Then read the shape. | What the counters show | Diagnosis | |---|---| | Per-seed counts roughly equal, then simply stop at seed N | Sequential drain of a global allowance | | One or two seeds orders of magnitude above the rest | The width cap is not binding — check it is applied per level, not per parent | | Total queries far above candidates x expected calls per node | The per-node constant is wrong: retries, re-scoring, or malformed attacker output being regenerated | | Token spend far above call spend | Deep lines re-sending accumulated conversation; a level-6 call costs several times a level-1 call | | The run's counter below the provider's usage figures | Retries after timeouts and 429s that your counter never incremented for | ### What the run cost, and what to change The fix is to stop treating the allowance as something consumed first-come and start treating it as something *divided*, in roughly this order of payoff. 1. **Per-seed cap = allowance / seed count**, enforced as a hard stop, plus a small shared reserve (10-15%) a seed may draw from only after showing partial movement. 2. **Assert the width cap in the loop.** Pruning each parent's children to w instead of pruning the whole level to w restores exponential growth in one seed and drains everything. 3. **Early stop on a confirmed hit.** Once a behaviour is demonstrated, further queries on that seed buy a duplicate, not a finding. 4. **Stall detection.** No improvement in best score across k consecutive levels ends the seed. 5. **Interleave.** Run every seed to level one, then every seed to level two, and so on. A run truncated at any point then has uniform coverage at a known depth instead of an arbitrary prefix. Interleaving is free in queries and costs a little engineer time in state management, which is the trade worth making: it converts an unreportable run into a reportable one. ### Where the number misleads The dangerous artefact is the report, not the run. **Untested reads as clean.** A hundred and thirty-three behaviours were never attempted, and unless the report says so explicitly, a reader sees "no findings for those" and concludes the model resisted them. The sentence "the model resisted our jailbreak suite" is not supportable by this run at all. Two subtler traps. **A 100%-of-allowance figure looks like completion** — the counter hitting its limit is exactly what a well-sized full run also looks like, so allowance utilisation is not evidence of coverage. And **hits per seed is ambiguous**: divided by seeds *attempted* it flatters the run, divided by seeds *planned* it is honest, and the two differ by a factor of three here. Say which denominator you used. If the seed list happens to be ordered by category or severity, prefix coverage is also *biased* coverage — you may have searched every prompt-injection seed and no data-exfiltration seed, which is worse than a random third. ### What to check before believing the next run Do the sizing sum in advance — seeds x per-seed cap x calls per candidate — and compare it to the allowance; if projected spend exceeds it, the run is already known to truncate and you fix the plan, not the run. Reconcile the tool's counter against the provider's usage for the window. Assert frontier size after each prune. And run the whole thing at a tiny allowance first: a dry run that truncates should show *every* seed at level one, not the first few at full depth. That single smoke test catches the scheduling bug before it costs an engagement. Provider-side rate limiting and concurrency tuning are a different problem; this is purely about how the search divides the queries it is permitted to make.

  • Why interleave levels across seeds rather than finishing each seed in turn?
    Because a truncated interleaved run still has uniform coverage of every seed at a known depth, which is reportable; a truncated sequential run covers an arbitrary prefix and nothing about the rest.
  • One seed shows a query count a hundred times the others. What do you check first?
    Whether the width cap is applied per level or per parent. Pruning each parent's children to w instead of the whole level to w restores exponential growth.
  • How should the finished report describe coverage after this run?
    As seeds attempted over seeds planned, plus the depth reached, and explicitly that untested seeds are untested rather than clean.

saying these in an interview costs you the question

  • Raises the allowance without changing how it is divided — the same prefix just gets longer.
  • Reports the run as evidence the model resisted, without stating how many seeds were attempted.
  • Blames the endpoint for slowness when the counters show one seed's tree consumed the queries.
  • Adds parallelism as the fix, which spends the same allowance faster rather than spreading it.
  • Keeps searching a seed after it has already produced a confirmed hit.

context