skip to content

A fraud model randomly quantises features before scoring, so a resubmitted transaction scores differently - why does that stall an attacker?

level: middleimportance: nice to knowfreq 32%

answer

  1. the model is no longer a function
  2. one call is one draw
  3. the difference you measure is the draw
  4. average it away, at a price
  5. square root, so four times for half

basics

~20 s

Because each reading is one draw from a distribution, not the quantity being optimised. The difference an attacker measures between two nearby candidates is swamped by the defence's randomness, so the search follows noise instead of a direction - while the misclassified transactions stay reachable.

solid answer

~40 s

Randomising the model turns the attacker's objective into an expectation, and a search that evaluates each candidate once is optimising a single sample of it. Two nearby candidates may differ by less than the spread the quantisation introduces, so the measured difference is mostly the draw, and steps taken from it go in arbitrary directions. That is why the attack appears to converge nowhere and the reported success rate collapses. It is a *precision* problem, not a robustness one: an adversary who averages the score over repeated evaluations recovers a usable signal, at a cost multiplier set by how much precision they need. Because the standard error of an average falls with the square root of the number of draws, halving the uncertainty costs roughly four times the evaluations - expensive, and paid once.

go deeper

for a junior

Recall that a randomised scoring step makes the model return different numbers for the same input, so any single reading an attacker takes is one sample rather than the value itself.

for a middle

Be ready to explain why small measured differences drown in that spread, and why averaging over draws restores the signal at a cost that grows with the square of the precision wanted.

for a senior

Show you would treat a finding that reproduces once in five tries as evidence of randomisation and re-run with an averaged objective before reporting anything.

for a principal

Own the framing: this is a query-cost control, and its value depends entirely on whether the adversary pays per call or holds the weights locally. Say which case you are buying.

## What randomising actually changes A deterministic scorer answers the same question the same way every time, so an attacker comparing two candidate transactions attributes any difference in score to the difference between the candidates. Insert a randomised step - say a quantisation of the normalised feature vector whose rounding boundaries are drawn per request - and the model is no longer a function of the input alone. The same transaction resubmitted returns a different number. The attacker's objective has quietly changed. What they care about is the model's *typical* behaviour on a candidate: an expectation over the defence's randomness. What they can observe in one call is a single draw from that distribution. ## Why one draw per candidate breaks the search An iterative attack works by making a small change and reading how the objective responds. The response to a small change is, by construction, small. If the defence's randomness contributes a spread comparable to or larger than that response, then the measured difference between two nearby candidates is dominated by which draw happened to come back. The direction the attacker infers is then largely arbitrary, the steps it produces do not accumulate, and the attack wanders. This holds whether the attacker reads the input gradient through the randomised step or infers a direction from returned scores. The failure is not about which vantage they have; it is that their measurement is being taken at a precision the defence has made unavailable in a single call. ## Why the misclassified inputs are untouched Nothing in this story moved a decision boundary. The set of transactions the model gets wrong on a typical draw is the same set it got wrong before, and for most candidates the randomisation changes the score by far less than the margin. The defence added variance to the *measurement channel*, and the measurement channel is the attacker's instrument, not the model's weakness. Hence *masked, not fixed*. ## What an adversary who has read the defence does They stop treating a single call as the objective and treat it as a sample of one. Averaging repeated evaluations of the same candidate reduces the spread of the estimate; the attack then optimises the averaged quantity, which is close to the expectation they actually cared about. The examples come back, and the corrected failure rate is typically near what it was before the defence existed. The cost is real and it is worth quoting precisely, because it is the entire value of the defence: - The standard error of an average falls with the **square root** of the number of draws. To halve the uncertainty you take roughly four times as many evaluations. - That multiplier applies at every step of an iterative search, so a fifty-step attack that now needs thirty draws per evaluation costs thirty times what it did. - Against a paid endpoint that is money and it is rate limits; against a downloaded weight file it is only wall-clock, and wall-clock is cheap. So the defence's honest description is a **query-cost multiplier**, and the multiplier is very different depending on whether the adversary must pay per call or holds the model locally. ## The tells this produces from the outside A randomised defence leaves a distinctive trace in an evaluation, and it is worth recognising because it is easy to mistake for flaky infrastructure: - The same attack against the same build returns noticeably different success rates on different seeds, so a finding *reproduces once in five tries*. - Increasing optimisation steps buys almost nothing, because more steps of a noisy direction do not accumulate. - A search that ignores gradients entirely may match or beat one that reads them, since neither has a reliable signal but the former never expected one. Triage is to treat non-reproducibility as data rather than as a broken harness: check whether the served pipeline is randomised at all, then re-run with the objective averaged across draws before reporting any number. ## What to write down A defensible line reads: *against an adversary who averages over the defence's randomness, the failure rate is X at radius R; the defence multiplies the required evaluations by roughly N.* A line reading *success dropped to 3%* with no mention of the randomisation is a measurement of the harness, and anyone planning against it is planning against a number that no adaptive adversary will ever reproduce.

  • Does it matter whether the attacker holds the weights or only sees returned scores?
    It changes the price, not the outcome. Both are taking a measurement the randomness has made imprecise, and both fix it by averaging over draws. The weights-holder pays in local compute, which is cheap; someone paying per call to a metered endpoint pays in money and rate limits. The defence is therefore worth far more against a metered adversary than against one holding a downloaded weight file.
  • How would you word what this defence bought, for a reviewer?
    As a cost multiplier with the corrected failure rate beside it: the adversary needs roughly N times the evaluations to reach the same success, and once they spend it the failure rate is X at radius R. Wording it as a drop in attack success, with no mention of the randomisation or the adaptive re-run, states a robustness claim the evaluation never supported.

Weighing two nearly identical parcels on a scale that jitters by more than the difference between them. The scale is fine after enough repeat weighings; it never changed which parcel was heavier.

saying these in an interview costs you the question

  • Calls the randomness robustness rather than imprecision
  • Assumes randomising moved the decision boundary
  • Treats non-reproducible findings as a broken harness
  • Ignores that averaging removes the defence's effect
  • Quotes the same cost for metered and local adversaries

context