skip to content

Before a gradient-guided prompt search can take its first optimisation step against an open-weights chat model, what must the team actually have on hand, and why does that same list have to be assembled again for every additional model in scope?

level: middleimportance: must knowfreq 58%

answer

  1. checkpoint, tokenizer, chat template
  2. licence and weight custody
  3. VRAM: params plus backward activations
  4. hours to days per model
  5. gradients belong to one weight tensor

basics

~20 s

You need the specific checkpoint's weights on a machine you control, its tokenizer and chat template, permission to hold those weights, and a GPU with room for the model plus backward-pass activations. The gradients belong to that one weight tensor, so nothing carries over: each new model means fresh weights, fresh setup and fresh GPU hours.

solid answer

~50 s

The shopping list splits into three parts. **Artefacts.** The exact checkpoint you intend to attack, not a sibling of it; its tokenizer; and the chat template that turns a conversation into the token sequence the model actually sees. Get the template wrong and you optimise a string the deployment never constructs. **Rights.** A licence or engagement clause letting you hold and run those weights, plus storage under the client's handling terms. **Compute.** GPU memory for the parameters *and* the backward pass. An input-token search carries no optimiser state for the weights, but it holds activations for the backward pass and scores batches of candidate substitutions, so the footprint sits well above plain inference — and the run is hours to days, not seconds. The cost repeats per target because a gradient is taken of one particular parameter set. A second model, or the same model after a point-release, is a different function; the search restarts from scratch.

go deeper

for a junior

Should name weights on local hardware and a GPU, and understand the run is not instant.

for a middle

Expected to give the full list — exact checkpoint, tokenizer, chat template, rights to hold the weights, memory beyond inference — and explain that gradients are tied to one parameter set so cost recurs per target.

for a senior

Adds operational discipline: recorded weight version, template matched to production, a trial run to measure headroom, an agreed stopping rule, and an explicit statement of which targets did not get this treatment.

for a principal

Turns the per-target recurrence into a budgeting rule for the engagement and a criterion for what capability the team keeps in house.

**Why this is an interview question rather than trivia.** Scoping an engagement means converting the sentence "we will attempt automated jailbreak search" into hardware, hours and custody obligations. Candidates who have only read about the method describe the algorithm; candidates who have run it describe the prerequisites, because the prerequisites are what size the plan or kill it. **Artefacts, ordered by how often they are wrong.** | Artefact | What it is | What goes wrong without it | |---|---|---| | The exact checkpoint | The specific parameter file that serves traffic | A base model, a sibling release or a public version of a client fine-tune is a *different function*; optimising against it is an unplanned transfer experiment | | The tokenizer | The vocabulary and merge rules that turn text into token ids | Substitutions are over tokens, not characters; a mismatched tokenizer produces a string that re-tokenises differently at serving time | | The chat template | The role markers and formatting that wrap a conversation into one sequence | You optimise a bare user string the deployment never constructs | | The system prompt | The fixed preamble production prepends | Without it you search against a materially easier model, since the preamble itself shifts refusal behaviour | | The numerical build | Precision and quantisation of the served weights | Full-precision local weights are not the reduced-precision build users hit | **Rights.** A licence or contract clause permitting you to hold and execute those weights, plus storage under the client's handling terms and a deletion commitment you can evidence. Open-weights licences carry use restrictions; a client's own fine-tune carries commercial terms. This is a real gate, not paperwork: some clients will simply refuse to release weights, and that answer arrives after the plan is written unless you ask first. **Compute, honestly.** GPU memory must cover three things at once, not one: the parameters, the activations retained for the backward pass, and the batch of candidate substitutions re-scored each step. As a rough anchor, a 7-billion-parameter model in 16-bit occupies about 14 GB for weights alone, so a 24 GB card is marginal once activations and candidate batches land on it, and a 40–80 GB accelerator is the comfortable floor. A 70-billion-parameter model is roughly 140 GB in the same precision and needs sharding across several accelerators before the first step runs. Wall clock is dominated by forward passes: hundreds to low thousands of iterations, each re-scoring a batch of candidates, which lands at hours per behaviour on a high-end accelerator — and an engagement measures a *suite* of behaviours, with restarts for the ones that do not converge. **Why it does not amortise.** The gradient is taken of one particular parameter tensor. A second model is a different function; the same model after a point release is a different function. The artefact you produced describes the thing you differentiated, and nothing else. That is the practical reason a five-target engagement cannot casually promise this treatment to all five: you either buy five times the hours or you state in the plan which targets got the search and which got query-driven testing only. **Where the number misleads.** The classic error is quoting a published per-target GPU-hour figure as if it were the engagement's budget. Those figures are usually *per behaviour*, on a small open-weights model, on the paper authors' hardware, counting only runs that converged. Rescale it and the real budget is that figure times a model-size factor, times the number of behaviours in scope, times the restart rate for non-converging runs — routinely one to two orders of magnitude above the quoted line. The mirrored error is on the output side: reporting a success rate whose denominator is the runs that converged, silently dropping the ones that hit the iteration cap. Both errors flatter the method, and both are caught by insisting that every rate carries its denominator in the sentence. **What to check before you spend hours.** Record the weight file's hash and version so the report says which function was tested. Confirm the template and system prompt against what production actually sends. Measure VRAM headroom on a short trial rather than inferring it from parameter count. Agree a stopping rule and an hour cap up front, so one non-converging behaviour cannot quietly consume the whole allocation. And fix the denominator before the run, not after the results are in.

  • You have the weights but not the production system prompt. What do you do?
    Ask for it, and say so in the plan if you cannot get it. Optimising without it searches a different input than production builds; at minimum, record the assumption and re-verify any hit against the real deployment.
  • The model does not fit on your GPU. What are the options, ranked?
    Run on bigger hardware, shard across GPUs, or pick a smaller checkpoint and accept that you are no longer attacking the target directly. The last is a scope change and belongs in the plan, not in a footnote.

A gradient search is a key cut for one lock: the hours buy a result against that exact checkpoint, and the next model — or the same model after a point release — starts again from a blank.

saying these in an interview costs you the question

  • Listing only 'a GPU' and treating the tokenizer, template and exact checkpoint as details.
  • Assuming inference-sized VRAM is enough for a backward pass.
  • Believing one successful run yields an artefact reusable across models or checkpoint updates.
  • No mention of the right to hold client weights, or of where they will be stored and when deleted.

context