skip to content

White-Box Search

A gradient costs weights on your own hardware and GPU hours per target, and returns a brittle, filterable string. Interviewers ask because candidates quote white-box results at query-only endpoints.

on this pageshow

explore

questions

14

A teammate proposes running a jailbreak method that optimises the prompt by taking gradients of a refusal-related loss with respect to the input tokens, and aiming it at a vendor's hosted chat endpoint you can only send text to and read text back from. Why can that class of method not run there at all?

level: juniorimportance: must knowfreq 72%

answer

  1. backward pass needs weights in memory
  2. text out is not a derivative
  3. log-probs are outputs, not gradients
  4. hosted target implies query-driven only
  5. check the access grant before the plan

basics

~20 s

Computing that gradient needs a backward pass through the model's weights. A hosted endpoint returns generated text, not derivatives with respect to your input tokens. With no weights loaded on hardware you control, there is nothing to differentiate, so the search never starts. Against that endpoint you are limited to query-driven attacks.

solid answer

~50 s

The method is a numerical optimisation over token positions: at each step it needs the derivative of a loss (roughly, how strongly the model is heading toward a refusal versus a compliant continuation) with respect to the embedding of each candidate token. That derivative only exists if you can run a forward *and* a backward pass through the actual parameters. A text-in/text-out endpoint gives you neither. Even an endpoint that returns top-k token log-probabilities gives you outputs, not gradients — log-probs support estimating a direction by sampling, but that is a different, query-hungry class of method, and every probe costs money and shows up in the vendor's abuse telemetry. So the practical rule: gradient-guided prompt search is available against a model whose weights you can load locally, and unavailable against a model you can only call. Proposing it for a hosted target is a sign someone has not checked what access the engagement actually grants.

go deeper

for a junior

Should say gradients require the model's weights locally, and that an API returns text, so the method cannot run against a hosted endpoint.

for a middle

Adds what each iteration actually needs (forward plus backward pass at the input embeddings) and why returned log-probabilities are not a substitute.

for a senior

Frames it as a scoping check: confirm the access class before choosing methods, and record in the report which families were unavailable so a clean result is not over-read.

for a principal

Treats it as a portfolio question — which target types the team faces, and therefore whether local-weights capability is worth owning at all.

## The mechanism, one iteration at a time A **gradient-guided prompt search** — the family that includes published methods such as GCG — treats the prompt as the variable being optimised and holds everything else fixed. One iteration does three things. - (1) *Forward pass.* The candidate prompt is tokenised, mapped to embedding vectors, and run through the network to produce a scalar loss. The loss is usually the negative log-likelihood the model assigns to a target continuation the attacker picked in advance — a compliant-looking opening phrase — so "lower loss" means "the model is more inclined to begin the way I want instead of refusing". - (2) *Backward pass.* Automatic differentiation propagates that loss back through every layer to the input embedding matrix, yielding a gradient vector at each attacker-controlled token position: a local, linear estimate of which direction in embedding space lowers the loss. - (3) *Candidate ranking and re-scoring.* Because the search space is a discrete vocabulary rather than a continuous vector, the gradient cannot be applied directly; it is used only to shortlist plausible token substitutions, and each shortlisted candidate is then re-scored with a further forward pass. Step 2 is the step that requires the **parameters**. It is not optional dressing on the method — it is the method. ## What a hosted endpoint actually returns A text-in/text-out chat API gives you - generated tokens, - sampling controls such as temperature and top-p, - sometimes a top-k log-probability field per generated token, - sometimes moderation labels, - and sometimes an embeddings endpoint that serves a *different* model entirely. Not one of those is a derivative of the serving model's loss with respect to your input embeddings. There is no request shape that turns a completion into a backward pass, because the vendor never computes one for your request and would have to expose its parameter graph to do so. ## Why "just estimate the gradient from queries" is not a workaround, in numbers **Derivative-free estimation** over a discrete vocabulary means comparing alternatives by measuring them. - A modern tokenizer has roughly 32,000 to 256,000 entries; - a control span is typically ten to thirty positions; - a single well-informed step compares thousands to tens of thousands of alternatives; - and this class of search converges, when it converges, over hundreds to low thousands of steps. Multiply and you are in the millions-to-tens-of-millions of endpoint calls **for one behaviour against one model**. At even a fraction of a cent per call that is a five- to six-figure bill, and at a typical commercial rate limit it is months of wall clock — before the vendor's abuse telemetry notices tens of thousands of near-identical adversarial prompts from one key and suspends it. That cost gap is exactly why the **query-driven families** exist: an iterative-rewrite loop, where a second model rephrases a failed prompt and retries, converges or gives up in tens of queries, not millions. ## Where the number misleads Two readings go wrong, and both matter more than the mechanism. - First, the *silent null*: a report that says "automated jailbreak search produced no successful attack" against a query-only target is describing the team's access, not the model's robustness. Nothing ran. A reader — often an executive or an auditor — converts that sentence into "the model withstood optimisation-based attack", which is a claim no one made and no evidence supports. - Second, the *quiet substitution*: a team runs the search against an open-weights model of similar architecture and reports the hits it lands on the real target as though they were results of a search against it. That is a **transfer measurement**, with a much lower and much more fragile success rate that decays with every checkpoint update on the target side, and it must be labelled as such. ## What you check before committing - Read the access grant, not the statement of work's adjectives: does the engagement deliver a checkpoint file, or only an API key? - If a checkpoint, is it the build that serves traffic, and are you licensed and contractually permitted to hold it? - If only a key, write the method plan as query-driven from the start and cap the query budget against the vendor's rate limits and the client's spend. - Confirm in writing that automated adversarial probing of that endpoint is authorised, since a third party's terms of service usually forbid it by default. - Then, in the report, name the families you could not run and why, so the clean result is read as the bounded thing it is.

  • The endpoint returns top-5 token log-probabilities. Does that change your answer?
    Not for gradient search. Log-probs are model outputs; they let you score candidates and estimate a search direction by sampling, which is a different, far more query-hungry family. You still have no derivative with respect to your input tokens.
  • The client will not share weights but will run your code inside their environment. Is the method back on the table?
    Potentially yes — the requirement is a backward pass through the parameters, not physical possession. If they will host your job next to the loaded model and return only the resulting artefacts, the search can run under their controls.
  • What do you write in the report about the methods you could not run?
    State them explicitly, with the reason (query-only access). Otherwise a reader assumes the model was tested against everything the team knows and treats the clean result as stronger than it is.

Estimating a gradient from queries alone is like mapping a hillside in thick fog by walking to a point, feeling the slope, and walking back — possible in principle, and the number of walks is the whole reason nobody does it against a metered endpoint.

saying these in an interview costs you the question

  • Claiming you can get gradients from an API by asking for log-probabilities.
  • Saying the method 'just needs enough queries' against a metered endpoint, with no sense of the query cost or the rate limits.
  • Confusing an embeddings endpoint with access to the serving model's parameters.
  • Planning the engagement without ever asking whether weights are part of the deliverable.

context

open as a page

A white-box token search appends a tuned suffix to a request and marks a trial successful when the model's reply begins with a preselected affirmative phrase. Why is that success criterion only a proxy, and what do you check before recording the trial as a real result?

level: juniorimportance: must knowfreq 60%

basics

~20 s

Because the check only reads the reply's first few tokens. A model can open with the agreed phrase and then refuse, stall, or produce useless text. The search optimises exactly that prefix, so it overfits it. Before recording a hit, read or grade the whole continuation.

open as a page

You ran a gradient-based search against an open-weights model on your own GPUs to produce an adversarial suffix, then sent that frozen string unchanged to a hosted chat endpoint whose weights you cannot inspect. What does it mean to say the attack "transferred", and why is a failure at the hosted endpoint less informative than a failure on the local model?

level: juniorimportance: must knowfreq 55%

basics

~20 s

Transfer means the string you optimised on the local model still produces the disallowed behaviour on the hosted endpoint, with no re-optimisation. Failure there is uninformative because you get no gradients and no internals back, only the reply text. On the local model a failure still shows you the loss and lets you keep searching.

open as a page

Before a gradient-guided prompt search can take its first optimisation step against an open-weights chat model, what must the team actually have on hand, and why does that same list have to be assembled again for every additional model in scope?

level: middleimportance: must knowfreq 58%

basics

~20 s

You need the specific checkpoint's weights on a machine you control, its tokenizer and chat template, permission to hold those weights, and a GPU with room for the model plus backward-pass activations. The gradients belong to that one weight tensor, so nothing carries over: each new model means fresh weights, fresh setup and fresh GPU hours.

open as a page

In a gradient-guided adversarial suffix search against a model whose weights you hold locally, why must the optimisation objective be a fixed target completion string rather than a compliance verdict from a judge model, and how does the choice of that string change the run?

level: middleimportance: must knowfreq 55%

basics

~20 s

Gradients need a differentiable number. Cross-entropy on a fixed target completion gives one; a judge's yes or no is a discrete label from another model with no usable gradient. The string you pick sets the difficulty: a short generic opener is easy but weak evidence, a long specific one is slower and rarer.

open as a page

Before spending GPU hours optimising a jailbreak string against a locally run open-weights model in the hope it fires at a hosted endpoint you cannot inspect, how do you choose which local model to optimise against, and why do practitioners often optimise against several local models at once?

level: middleimportance: must knowfreq 48%

basics

~20 s

Pick a local model as close as you can guess to the hidden one: similar tokenizer, similar chat formatting, similar safety tuning. The string is optimised for those exact tokens and that exact refusal behaviour. Optimising against several local models at once forces the string onto behaviour they share instead of one model's quirks, which usually survives the jump better.

open as a page

A suffix produced by a token-level search is a run of unrelated characters and word fragments no person would type. A deployment adds an input check that rejects prompts whose per-token perplexity under a small language model is far above normal traffic. Why does that check defeat this class of result cheaply, and what does it cost the defender?

level: seniorimportance: must knowfreq 55%

basics

~20 s

The search optimises token probabilities, not readability, so the winning string is statistically bizarre — exactly what a perplexity score measures. Scoring it needs one pass through a tiny model, far cheaper than the search that produced it. The cost is false positives on legitimately odd input and a threshold that must be tuned per traffic mix.

open as a page

A quarter ago your report recorded that a gradient-optimised suffix produced the disallowed behaviour on roughly two of five attempts against a hosted chat endpoint. Re-running the same frozen string today, it almost never works. What are the plausible causes, and what should the original entry have recorded so this is diagnosable at all?

level: seniorimportance: must knowfreq 42%

basics

~20 s

The endpoint changed under you: a new model build, an altered system prompt, or a filter added in front or behind it. Your string did not decay. The entry should have stamped the date, the endpoint and any build identifier returned, decoding settings, the number of attempts, and how a hit was judged. Without those, nothing is attributable.

open as a page

An adversarial suffix optimised against a locally run open-weights model fires reliably at one vendor's hosted chat endpoint but does nothing at another vendor's hosted endpoint of comparable capability. What mechanisms explain the difference, and how would you work out which one is responsible without any access to either system's internals?

level: middleimportance: should knowfreq 38%

basics

~20 s

Comparable capability does not mean comparable internals. Different tokenizers re-split your string, different safety tuning gives a different refusal to suppress, hidden system prompts differ, and one endpoint may be a pipeline with classifiers around the model. You separate them from the outside by comparing the shape of the failures: wording, variation, latency, and whether any output streamed at all.

open as a page

A client hands you the weight file for the model behind their product, and your gradient-guided prompt search converges on your workstation. Production serves a quantised build of that checkpoint, with a fixed system prompt prepended and a separate input classifier in front. Which of your preconditions were actually satisfied, and what would you change locally before spending more GPU hours?

level: seniorimportance: should knowfreq 42%

basics

~20 s

Only one precondition held: you had weights you could differentiate. You optimised a different numerical build, on an input that omits the production preamble, for a path that has no classifier in front. Before more GPU hours, load the quantised build, include the real system prompt as fixed context, and search only the attacker-controlled region.

open as a page

You are running a token-level suffix search against a local open-weights model and must set the step budget up front, knowing each run bills GPU hours per model and per target behaviour. How do you set it, and what tells you to stop early or restart rather than spend the remaining steps?

level: seniorimportance: should knowfreq 40%

basics

~20 s

Set it from the GPU hours you can spend divided across the behaviours you must cover, not from a number in a paper. Watch the loss curve and run the real success check periodically: stop the moment a verified hit lands, and restart from a fresh initial suffix when the loss has plateaued rather than buying more steps.

open as a page

You lead a red team whose targets are mostly third-party hosted chat endpoints, with one or two open-weights models deployed in house. A senior engineer proposes buying GPU capacity and standing up a weight-handling process so the team can run gradient-guided prompt searches. How do you decide, and what recurring costs sit behind the hardware line item?

level: principalimportance: should knowfreq 30%

basics

~20 s

Decide from the target mix: the capability only applies where you hold weights, which here is one or two systems. Beyond hardware you take on weight custody and deletion duties, licence review, per-target GPU hours that never amortise, re-runs after every checkpoint update, and staff time to keep the rig matching each serving stack.

open as a page

As the lead of an engagement you receive one artefact: a nonsensical suffix, found by a token search on weights your organisation hosts, that reliably drives that model to produce disallowed content. Which defensive decisions does that result legitimately support, and which does it not?

level: principalimportance: should knowfreq 30%

basics

~20 s

It supports layered defence: input anomaly checking, output-side review, and a regression case kept for future checkpoints. It shows refusal training does not hold off-distribution. It does not support severity claims about a live surface, statements about systems you hold no weights for, or any headline that the model is broadly unsafe.

open as a page

You lead a red-team program with in-house GPUs for optimising attack strings against local models, and a set of hosted endpoints you can only send prompts to. Transferred strings tend to stop working after vendor changes. How do you decide how much of a program's effort goes into surrogate-optimised transfer attacks versus attacks discovered by querying the hosted endpoints directly, and what evidence would make you cut transfer work entirely?

level: principalimportance: should knowfreq 28%

basics

~20 s

Decide from your own history, not from the literature: track what fraction of locally optimised strings ever fired at a hosted endpoint and how long they kept working. If that half-life is shorter than your reporting cycle, transfer buys exhibits that are stale on delivery. Keep it where it produces cheap candidates; cut it when the endpoints are guard-dominated.

open as a page