skip to content

Estimating the Gradient

Withholding the weights does not withhold the gradient, it prices it: a large query count reads as a barrier and is an invoice. Interviewers ask because teams treat a closed model as a boundary.

on this pageshow

explore

questions

4

With no weights, only a fraud API's returned risk score, how does an attacker get a search direction?

level: juniorimportance: must knowfreq 62%

answer

  1. the response is still observable from outside
  2. the endpoint answers what you pay for
  3. difference two returned scores
  4. a free operation became a metered one

basics

~20 s

The attacker buys the direction instead of computing it. Sending slightly varied transactions and differencing the returned risk scores shows how the score moves - the same information a gradient gives, at one paid call per probe.

solid answer

~50 s

An attacker holding the weights gets the direction for free: one pass gives the gradient of the score with respect to the input, which says how to change the transaction to lower its risk score. Without weights that computation is unavailable, but the information is not. The endpoint answers every question it is paid to answer, so the attacker sends a base transaction and a set of slightly varied ones and differences the returned scores; the pattern of changes estimates the same direction. The estimate is noisy and approximate rather than exact, and each sample is a metered call on a commercial contract. So the honest statement is not `no weights, no attack` but `a free operation became a metered one` - the barrier moved from possibility to price. This family is usually called score-based, or zeroth-order, in the literature.

go deeper

for a junior

Be ready to say that a returned score is itself information about how the model responds, and that differencing scores across slightly varied inputs substitutes for a gradient the attacker cannot compute.

for a middle

An interviewer expects you to explain that the wanted gradient is with respect to the input rather than the weights, that the bought version is approximate, and that each sample is one paid call.

for a senior

Show that you reframe the question from can-they to what-does-it-cost: state the price per call, the samples per direction and the steps, and say when that total stops being worth one evaded transaction.

for a principal

Own the framing that withholding weights changes an attacker's cost structure rather than their capability, and that any assurance argument resting on we never publish the model is an economic argument that has to be priced.

## What direction is being talked about In ordinary training, a gradient is taken with respect to the **weights**: it says how to change the model to fit the data better. An evasion attacker wants a different one, taken with respect to the **input**: holding the model fixed, which way should this transaction be nudged so the returned risk score drops? It is the same arithmetic pointed the other way, and it is the reason a single small, structured change can move a confident score while random noise of the same size does essentially nothing. Noise spreads its magnitude in no particular direction; the attack direction is chosen because it is the one the score is most sensitive to. ## Why holding the weights makes it free An attacker with the model file computes that input direction directly. It costs one evaluation, uses no network, leaves no log line anywhere, and can be repeated as often as the attacker's own hardware allows. That is the white-box vantage, and it is an **access assumption** - weights, architecture and gradients granted - not an event that happened. ## Why not holding them does not end the attack The common wrong answer at this point is that a third party with no weights cannot compute a gradient and therefore cannot attack. The first half is true and the second does not follow. A gradient is a statement about how an output responds to small changes in an input, and a hosted scoring service publishes exactly that response, one paid answer at a time. If the same transaction is sent again with one field moved slightly and the returned score comes back different, the difference between the two scores is a measurement of the model's sensitivity in that direction. Repeat that over the fields, or over a set of mixed directions, and what comes back is an approximation of the direction the white-box attacker computed for nothing. Nothing about the model has to be known for this to work. The architecture, the framework, the feature preprocessing and the calibration step are all inside the box. The attacker never sees them and never needs to, because the estimate is taken through the endpoint's behaviour rather than through its internals. ## What the attacker actually gets, and what it costs Three properties define the family: - **The estimate is approximate.** It is a measurement, not a derivation, and it carries sampling error. That is usually tolerable - the search takes many small steps, and a direction only has to be roughly right for progress to continue. - **Every sample is a query.** On a metered commercial API that is a line on an invoice: a price per call, multiplied by the samples spent on each direction estimate, multiplied by the number of steps the search takes. - **The whole attack becomes a counting problem.** With weights in hand, the interesting question is how good the optimiser is. Without them, the interesting question is how many calls the direction costs and whether that number fits a budget. This is why score-based work is quoted in thousands of calls rather than in milliseconds. ## How to say this in an interview The sharp framing is that access does not change what is computable, it changes what is billable. A defender who reasons `we never publish the weights, so nobody can take a gradient against our model` has described a change in the attacker's cost structure and mistaken it for a change in the attacker's capability. The right follow-on question is never `can they` but `what does it cost them, and does the value of one evaded transaction exceed it`. A related trap runs the other way. Because the estimate is bought rather than derived, some candidates describe the probing perturbations as noise. They are not noise: noise is what you add when you have no information about direction, and the entire point of probing is to acquire that information. The word matters because the contrast between a structured direction and same-sized random noise is the single most load-bearing fact in this whole area. ## Where it stops The limits of the family are all limits on the meter, not on the mathematics. A narrow input, a cheap call and a generous budget make the estimate practically free; a very wide input, an expensive call or a tight spend cap can price it out entirely. And a returned number that carries very little precision can starve the differencing of signal, because two probes that come back as the same printed value have measured nothing. None of those make the attack impossible - they make it more or less worth doing.

  • Does the estimated direction have to be accurate for the attack to work?
    No. It has to be right often enough to make progress. The search takes many small steps, and errors in individual estimates partly cancel across steps, so a coarse direction that agrees with the true one most of the time still walks the input toward a lower score. Accuracy buys fewer steps, not the difference between working and not working.
  • Why is the probing perturbation not just random noise?
    Noise is what you send when you have no directional information; probing is how directional information is acquired. Random changes of the same magnitude essentially never flip a trained model, because they spread their budget across directions the model is insensitive to. The probes are used to find the direction the score actually responds to, and only that direction is then followed.
  • Does the attacker need to know the model's architecture first?
    No. Everything from feature preprocessing to the final calibration sits inside the endpoint, and the estimate is taken through the endpoint's observed behaviour. That is precisely why this family is priced in queries rather than in reverse-engineering effort: the attacker substitutes paid measurement for knowledge of the internals.

You cannot read the recipe, but you can order the dish a hundred times with one ingredient changed each time and learn which ingredient the taste depends on. The kitchen stays closed; the bill grows.

saying these in an interview costs you the question

  • Says no weights means no gradient, therefore no attack
  • Claims the attacker must first steal or rebuild the weights
  • Calls the probing perturbations random noise
  • Assumes the estimate must match the true gradient exactly
  • Treats query count as free because inference is cheap

context

open as a page

On a paid scoring API, why does estimating an attack direction cost more queries as the input gets wider?

level: middleimportance: should knowfreq 45%

basics

~20 s

Each estimate is built from probes, and covering the input coordinate by coordinate needs at least one probe apiece. Widen the feature vector and the calls per direction rise in proportion, then multiply by every step of the search.

open as a page

Your black-box evasion test on a metered API ran out of budget with no evasion - what does that establish?

level: seniorimportance: should knowfreq 35%

basics

~10 s

A budget-exhausted run bounds the spend, not the model: at this price per call, this sample count per direction and this many steps, no evasion was reached. It says nothing about a better-funded adversary.

open as a page

A fraud API returns risk scores rounded to two decimals - how does that affect a probe-and-difference attack?

level: seniorimportance: nice to knowfreq 28%

basics

~20 s

Probes whose effect falls below the rounding step return the same printed number, so the difference is zero and that paid sample bought nothing. The attacker must probe harder or buy more samples, which raises the bill.

open as a page