skip to content

The Cost of Gradients

Gradients need weights on hardware you control, so this class is routine against an open-weights model and out of reach against a hosted one. Interviewers ask candidates who quietly conflate the two.

on this pageshow

explore

questions

4

A teammate proposes running a jailbreak method that optimises the prompt by taking gradients of a refusal-related loss with respect to the input tokens, and aiming it at a vendor's hosted chat endpoint you can only send text to and read text back from. Why can that class of method not run there at all?

level: juniorimportance: must knowfreq 72%

answer

  1. backward pass needs weights in memory
  2. text out is not a derivative
  3. log-probs are outputs, not gradients
  4. hosted target implies query-driven only
  5. check the access grant before the plan

basics

~20 s

Computing that gradient needs a backward pass through the model's weights. A hosted endpoint returns generated text, not derivatives with respect to your input tokens. With no weights loaded on hardware you control, there is nothing to differentiate, so the search never starts. Against that endpoint you are limited to query-driven attacks.

solid answer

~50 s

The method is a numerical optimisation over token positions: at each step it needs the derivative of a loss (roughly, how strongly the model is heading toward a refusal versus a compliant continuation) with respect to the embedding of each candidate token. That derivative only exists if you can run a forward *and* a backward pass through the actual parameters. A text-in/text-out endpoint gives you neither. Even an endpoint that returns top-k token log-probabilities gives you outputs, not gradients — log-probs support estimating a direction by sampling, but that is a different, query-hungry class of method, and every probe costs money and shows up in the vendor's abuse telemetry. So the practical rule: gradient-guided prompt search is available against a model whose weights you can load locally, and unavailable against a model you can only call. Proposing it for a hosted target is a sign someone has not checked what access the engagement actually grants.

go deeper

for a junior

Should say gradients require the model's weights locally, and that an API returns text, so the method cannot run against a hosted endpoint.

for a middle

Adds what each iteration actually needs (forward plus backward pass at the input embeddings) and why returned log-probabilities are not a substitute.

for a senior

Frames it as a scoping check: confirm the access class before choosing methods, and record in the report which families were unavailable so a clean result is not over-read.

for a principal

Treats it as a portfolio question — which target types the team faces, and therefore whether local-weights capability is worth owning at all.

## The mechanism, one iteration at a time A **gradient-guided prompt search** — the family that includes published methods such as GCG — treats the prompt as the variable being optimised and holds everything else fixed. One iteration does three things. - (1) *Forward pass.* The candidate prompt is tokenised, mapped to embedding vectors, and run through the network to produce a scalar loss. The loss is usually the negative log-likelihood the model assigns to a target continuation the attacker picked in advance — a compliant-looking opening phrase — so "lower loss" means "the model is more inclined to begin the way I want instead of refusing". - (2) *Backward pass.* Automatic differentiation propagates that loss back through every layer to the input embedding matrix, yielding a gradient vector at each attacker-controlled token position: a local, linear estimate of which direction in embedding space lowers the loss. - (3) *Candidate ranking and re-scoring.* Because the search space is a discrete vocabulary rather than a continuous vector, the gradient cannot be applied directly; it is used only to shortlist plausible token substitutions, and each shortlisted candidate is then re-scored with a further forward pass. Step 2 is the step that requires the **parameters**. It is not optional dressing on the method — it is the method. ## What a hosted endpoint actually returns A text-in/text-out chat API gives you - generated tokens, - sampling controls such as temperature and top-p, - sometimes a top-k log-probability field per generated token, - sometimes moderation labels, - and sometimes an embeddings endpoint that serves a *different* model entirely. Not one of those is a derivative of the serving model's loss with respect to your input embeddings. There is no request shape that turns a completion into a backward pass, because the vendor never computes one for your request and would have to expose its parameter graph to do so. ## Why "just estimate the gradient from queries" is not a workaround, in numbers **Derivative-free estimation** over a discrete vocabulary means comparing alternatives by measuring them. - A modern tokenizer has roughly 32,000 to 256,000 entries; - a control span is typically ten to thirty positions; - a single well-informed step compares thousands to tens of thousands of alternatives; - and this class of search converges, when it converges, over hundreds to low thousands of steps. Multiply and you are in the millions-to-tens-of-millions of endpoint calls **for one behaviour against one model**. At even a fraction of a cent per call that is a five- to six-figure bill, and at a typical commercial rate limit it is months of wall clock — before the vendor's abuse telemetry notices tens of thousands of near-identical adversarial prompts from one key and suspends it. That cost gap is exactly why the **query-driven families** exist: an iterative-rewrite loop, where a second model rephrases a failed prompt and retries, converges or gives up in tens of queries, not millions. ## Where the number misleads Two readings go wrong, and both matter more than the mechanism. - First, the *silent null*: a report that says "automated jailbreak search produced no successful attack" against a query-only target is describing the team's access, not the model's robustness. Nothing ran. A reader — often an executive or an auditor — converts that sentence into "the model withstood optimisation-based attack", which is a claim no one made and no evidence supports. - Second, the *quiet substitution*: a team runs the search against an open-weights model of similar architecture and reports the hits it lands on the real target as though they were results of a search against it. That is a **transfer measurement**, with a much lower and much more fragile success rate that decays with every checkpoint update on the target side, and it must be labelled as such. ## What you check before committing - Read the access grant, not the statement of work's adjectives: does the engagement deliver a checkpoint file, or only an API key? - If a checkpoint, is it the build that serves traffic, and are you licensed and contractually permitted to hold it? - If only a key, write the method plan as query-driven from the start and cap the query budget against the vendor's rate limits and the client's spend. - Confirm in writing that automated adversarial probing of that endpoint is authorised, since a third party's terms of service usually forbid it by default. - Then, in the report, name the families you could not run and why, so the clean result is read as the bounded thing it is.

  • The endpoint returns top-5 token log-probabilities. Does that change your answer?
    Not for gradient search. Log-probs are model outputs; they let you score candidates and estimate a search direction by sampling, which is a different, far more query-hungry family. You still have no derivative with respect to your input tokens.
  • The client will not share weights but will run your code inside their environment. Is the method back on the table?
    Potentially yes — the requirement is a backward pass through the parameters, not physical possession. If they will host your job next to the loaded model and return only the resulting artefacts, the search can run under their controls.
  • What do you write in the report about the methods you could not run?
    State them explicitly, with the reason (query-only access). Otherwise a reader assumes the model was tested against everything the team knows and treats the clean result as stronger than it is.

Estimating a gradient from queries alone is like mapping a hillside in thick fog by walking to a point, feeling the slope, and walking back — possible in principle, and the number of walks is the whole reason nobody does it against a metered endpoint.

saying these in an interview costs you the question

  • Claiming you can get gradients from an API by asking for log-probabilities.
  • Saying the method 'just needs enough queries' against a metered endpoint, with no sense of the query cost or the rate limits.
  • Confusing an embeddings endpoint with access to the serving model's parameters.
  • Planning the engagement without ever asking whether weights are part of the deliverable.

context

open as a page

Before a gradient-guided prompt search can take its first optimisation step against an open-weights chat model, what must the team actually have on hand, and why does that same list have to be assembled again for every additional model in scope?

level: middleimportance: must knowfreq 58%

basics

~20 s

You need the specific checkpoint's weights on a machine you control, its tokenizer and chat template, permission to hold those weights, and a GPU with room for the model plus backward-pass activations. The gradients belong to that one weight tensor, so nothing carries over: each new model means fresh weights, fresh setup and fresh GPU hours.

open as a page

A client hands you the weight file for the model behind their product, and your gradient-guided prompt search converges on your workstation. Production serves a quantised build of that checkpoint, with a fixed system prompt prepended and a separate input classifier in front. Which of your preconditions were actually satisfied, and what would you change locally before spending more GPU hours?

level: seniorimportance: should knowfreq 42%

basics

~20 s

Only one precondition held: you had weights you could differentiate. You optimised a different numerical build, on an input that omits the production preamble, for a path that has no classifier in front. Before more GPU hours, load the quantised build, include the real system prompt as fixed context, and search only the attacker-controlled region.

open as a page

You lead a red team whose targets are mostly third-party hosted chat endpoints, with one or two open-weights models deployed in house. A senior engineer proposes buying GPU capacity and standing up a weight-handling process so the team can run gradient-guided prompt searches. How do you decide, and what recurring costs sit behind the hardware line item?

level: principalimportance: should knowfreq 30%

basics

~20 s

Decide from the target mix: the capability only applies where you hold weights, which here is one or two systems. Beyond hardware you take on weight custody and deletion duties, licence review, per-target GPU hours that never amortise, re-runs after every checkpoint update, and staff time to keep the rig matching each serving stack.

open as a page