skip to content

A client hands you the weight file for the model behind their product, and your gradient-guided prompt search converges on your workstation. Production serves a quantised build of that checkpoint, with a fixed system prompt prepended and a separate input classifier in front. Which of your preconditions were actually satisfied, and what would you change locally before spending more GPU hours?

level: seniorimportance: should knowfreq 42%

answer

  1. access held, fidelity did not
  2. quantised build is a different function
  3. template and system prompt as fixed context
  4. classifier outside the loop, inside the eval
  5. replay survivors end-to-end

basics

~20 s

Only one precondition held: you had weights you could differentiate. You optimised a different numerical build, on an input that omits the production preamble, for a path that has no classifier in front. Before more GPU hours, load the quantised build, include the real system prompt as fixed context, and search only the attacker-controlled region.

solid answer

~50 s

The access precondition was met — you could run a backward pass. The *fidelity* preconditions were not. **Numerics.** A quantised build is a different function from the full-precision one. A brittle token sequence found against full precision can dissolve under quantisation. **Input construction.** Production prepends a system prompt and wraps everything in a chat template. If your local run optimised a bare user string, you searched over a sequence the deployment never builds. **Defence path.** The input classifier is not differentiable through your loop, and probably should not be: it is a separate component that either passes or blocks the string. A hit that never reaches the model is not a hit. The fix is to make the local rig look like the serving path: quantised weights, real template and preamble held fixed, optimisation restricted to the span an attacker controls, and every surviving candidate replayed against the live deployment through the classifier before anything is written up as a finding.

go deeper

for a junior

Should notice at least that the local model and the production model are not identical.

for a middle

Names the three mismatches — precision, formatted input, filter in front — and says the fix is to rebuild the local rig to match the serving path.

for a senior

Expected to sequence it: match the build, freeze real scaffolding, calibrate on a short run, filter candidates through the classifier, replay survivors end-to-end, and report only confirmed findings.

for a principal

Frames the mismatch as a contracting problem: what artefacts the engagement must require up front so the team never spends its allocation on the wrong function.

## Split the scenario into two questions - *Could the search run?* Yes — you held weights on hardware you control, so a backward pass was available and the access precondition was genuinely met. - *Was the function you differentiated the function that serves users?* No, in three separate ways. Each one costs you something different, and only the first is about the model at all. ## Mismatch one: numerics Production serves a **quantised build** — the same trained parameters stored at reduced precision (8-bit or 4-bit is typical) to cut memory and latency. Quantisation is a deployment economics choice, not a safety control, but it does change the computed function slightly at every layer. Optimised token sequences of this kind are brittle by construction: the search drives the input onto a sharp, narrow feature of the loss surface, which is exactly the structure that small numerical perturbation disturbs. A sequence that reliably lands against full precision can dissolve against the 4-bit build and vice versa. The fix is cheap when possible — obtain and load the serving build — and when the client cannot release it, the mismatch becomes a stated limitation in the report rather than a surprise during the demo. ## Mismatch two: the constructed input The model never receives a bare user string. It receives a formatted sequence: role markers from the chat template, a fixed system preamble, often retrieved context and tool descriptions. If your local loop optimised only the raw user text, the gradients were taken through a sequence the deployment never builds. Two consequences follow. - The **preamble** materially shifts refusal behaviour, so a search without it was run against an easier model and its convergence rate is inflated. - And if the preamble is confidential, holding it is one more **custody obligation** to negotiate. ## Mismatch three: the filter in front A separate **input classifier** sits ahead of the model and decides pass or block before a token is generated. - Keep it *outside* the differentiable objective — it is a distinct component, usually not differentiable in your loop, and forcing it in muddles the objective and misrepresents the deployment. - But do not keep it outside the *evaluation*: score every candidate the search emits against the classifier as a pass/fail gate. High-entropy optimised token strings are cheap for a filter to notice precisely because they look nothing like natural traffic, so this is where the local yield most often collapses. ## Where the number misleads — the highest-stakes part The local run produces a **convergence count**: how many target behaviours the optimiser drove to the chosen continuation on your rig. It is tempting to report that as an attack success rate against the product. It is not one, because the denominator and the system both changed. A representative shape: 40 of 50 behaviours converge locally; replay through the quantised build with the real preamble and the classifier leaves 6 that reproduce end to end. Reporting 40 makes a claim about a system that was never exercised, and it is the specific misreading this scenario exists to catch. - The findings section carries 6. - The local 40 belongs in method notes, as evidence that the search itself worked, alongside the **attrition** at each stage — that attrition is genuinely useful to the client, because it shows which of their controls did the work. ## The rig, in order, and what it costs - (1) Reproduce the serving stack locally: quantised weights, exact template, real system prompt. - (2) Freeze all of it and let the optimiser vary only the span an attacker actually controls. - (3) Run a short **calibration search** on two or three behaviours before committing the allocation, and measure how many survivors clear the classifier; if that number is near zero, the answer is a different method, not more GPU hours. Calibration costs a few hours and routinely saves the rest of the budget. - (4) Replay every survivor against the live deployment, with the client's knowledge and inside the authorised window, and report only the confirmed set. ## What you tell the client That the local result demonstrates the search converges against *this* build under *these* conditions, and that a finding becomes a finding only when it reproduces through the path that serves their users.

  • The client cannot give you the quantised build. What now?
    Run against what you have, calibrate expectations down, and state the mismatch plainly as a limitation. Then lean harder on end-to-end replay: only candidates confirmed against the live deployment get reported.
  • The system prompt is confidential and they will not share it. How do you proceed?
    Optimise with a placeholder preamble of similar length and role, treat the result as weaker evidence, and ask whether they will run the search themselves in-environment with the real preamble.

saying these in an interview costs you the question

  • Reporting local convergence as a product finding without replaying through the production path.
  • Assuming a full-precision result carries to a quantised serving build unchanged.
  • Optimising a bare user string when the deployment always prepends a system prompt.
  • Trying to differentiate through the input classifier instead of using it as a pass/fail filter on candidates.
  • Spending the whole GPU allocation before any candidate has been checked against the real path.

context