You are running a token-level suffix search against a local open-weights model and must set the step budget up front, knowing each run bills GPU hours per model and per target behaviour. How do you set it, and what tells you to stop early or restart rather than spend the remaining steps?
answer
- GPU hours first, steps second
- measure seconds per step locally
- outer check on an interval, exit on hit
- plateau means restart, not more steps
- negative result is bounded by the budget
basics
~20 sSet it from the GPU hours you can spend divided across the behaviours you must cover, not from a number in a paper. Watch the loss curve and run the real success check periodically: stop the moment a verified hit lands, and restart from a fresh initial suffix when the loss has plateaued rather than buying more steps.
solid answer
~50 sThe budget is a scheduling decision, not a hyperparameter you inherit. Start from the GPU hours available for the engagement, divide across the behaviours and models you must cover, and convert to steps using a per-step cost measured on your own hardware — each step is a backward pass plus a batch of candidate forward passes, so candidate batch size and suffix length move the arithmetic as much as the step count. Then instrument the run so it can end early. Two signals matter: - **A verified hit.** Run the full-reply check on an interval, not only at the end, and terminate on a confirmed success. - **A plateau.** If total and per-token losses have been flat with no hit, the greedy search is stuck; a restart from a different initial suffix or length usually beats buying more steps. Record the budget with the outcome: "no suffix found in this many steps" is a bounded negative, never a robustness claim.
go deeper
Should know the run has a step limit and that more steps cost more GPU time.
Explains per-step cost as a backward pass plus a candidate batch, and knows to stop on a verified hit rather than always running to the limit.
Plans the budget across behaviours from GPU hours, restarts on plateaus, caps wall clock per behaviour, and reports the budget with every negative result.
Frames it as recurring capacity per checkpoint and decides the broad-shallow versus narrow-deep allocation according to what the engagement is meant to answer.
**What one step is, concretely.** A single iteration of a token-level suffix search is: one backward pass through the target model to get the ranking signal over suffix positions, then a *batch* of forward passes — one per candidate substitution being evaluated — to pick the winner exactly. The candidate batch usually dominates. So per-step cost scales with candidate batch size, with suffix length (more positions to consider), and with total sequence length (prompt plus suffix plus target, which sets the cost of every pass). "Steps" is therefore not a portable unit. Five hundred steps at a batch of 512 on one GPU and five hundred steps at a batch of 64 on another are different experiments with the same label. **Budget in GPU-hours, then convert.** The only honest planning unit is GPU-hours, because that is what you are actually buying and what you can actually compare across behaviours and hardware. The order of operations is: (1) take the GPU-hours the engagement has; (2) measure seconds-per-step on *your* hardware, at *your* batch size, with a short pilot — never inherit a step count from a write-up; (3) decide the allocation across behaviours; (4) divide into steps. A step budget copied from someone else's setup is a number with the units stripped off. **Allocation is the decision that matters.** You almost never have one behaviour. The real question is broad-and-shallow versus narrow-and-deep. Many behaviours at a modest budget each finds the easy ones and maps where refusal training is thin — useful for prioritising defensive work. A few behaviours at a large budget answers whether one specific behaviour is reachable at all — useful when someone has asked exactly that. Pick which question the engagement is answering before you fix per-run numbers, because the two allocations produce non-comparable outputs and you cannot retrofit one into the other. **Early exit and restart discipline.** - Evaluate the outer success check on a schedule, not only at the end, and terminate on a *verified* hit. Running to the limit after a confirmed success is pure waste, and it happens constantly because the loop was written to run a fixed count. - Log per-target-token loss alongside the total. A single stubborn token that never falls usually means the target string is wrong for this model, not that the budget is short — and no number of extra steps fixes a wrong target. - On a plateau, restart. The candidate step is greedy and gets stuck in a basin; a fresh initial suffix, a different suffix length or an adjusted target explores a different one. Several short restarts frequently beat one long run for identical GPU-hours. - Cap wall-clock per behaviour so one hard case cannot silently consume the engagement's whole allocation. **The recurring cost nobody budgets for.** The artefact is bound to the exact checkpoint it was optimised against. A fine-tune, a model merge, a quantisation change or a new release invalidates it, and the search re-runs from scratch. So the correct answer to "what does it cost to keep testing this?" is a per-checkpoint recurring line item, not a one-off. Teams that budget it once and then discover the suffix stopped working after a routine weight update have mispriced the whole programme. **Where the number misleads.** Two failures, in opposite directions. A *budget-exhausted run with no hit* is a bounded negative — no suffix was found for that behaviour, that target string, that checkpoint, that budget, that initialisation. It is not evidence of robustness, and the slide that turns "search failed" into "model resistant" is the most common dishonest move in this area, usually made in good faith by someone who did not write the run. Conversely, a *hit at step 40* reported as "cheap to break" ignores the restarts and target-string iterations that preceded the run that worked; the honest cost is the whole search programme for that behaviour, not the lucky run. **What to check before you believe a result.** Publish and demand: step count, candidate batch size, suffix length, target completion string, initialisation, hardware, wall-clock and GPU-hours, and the number of restarts. Without those, a negative result is unfalsifiable and a positive one is unreproducible — and either way, nobody downstream can price the next run.
- Why is a plateau usually a reason to restart rather than to extend the budget?The candidate step is greedy and gets stuck in a basin. A fresh initial suffix or a different suffix length explores a new basin, and several short restarts often beat one long run for the same GPU hours.
- A run finishes its budget with no hit. What may you say about the model?Only that no suffix was found for that behaviour, that target, that checkpoint and that budget. It is a bounded negative and not evidence of robustness.
- Why is the budget a recurring cost rather than a one-off?The suffix is tied to the checkpoint it was optimised against; a fine-tune, merge or quantisation change means re-running the search, so the spend repeats every time the model changes.
A plateaued search is a car with its wheels spinning in one rut: pressing the accelerator harder is buying more of the same. Backing out and starting from a different spot is what actually moves you, and it usually costs less fuel than the spinning did.
saying these in an interview costs you the question
- Copies a step count from a write-up without measuring per-step cost locally.
- Reports a budget-exhausted run as evidence the model is robust.
- Runs to the step limit even after a verified success.
- No plan for re-running when the checkpoint changes.