skip to content

A fraud API returns risk scores rounded to two decimals - how does that affect a probe-and-difference attack?

level: seniorimportance: nice to knowfreq 28%

answer

  1. the returned number is the whole channel
  2. two probes, one printed value
  3. a zero difference is not a zero response
  4. bigger probes measure a wider region

basics

~20 s

Probes whose effect falls below the rounding step return the same printed number, so the difference is zero and that paid sample bought nothing. The attacker must probe harder or buy more samples, which raises the bill.

solid answer

~50 s

The estimate is built from differences between returned scores, so the returned precision sets the smallest change that is visible at all. If a probe moves the underlying score by less than the rounding step, both calls print the same value and the differencing measures zero - not because the model is insensitive there, but because the signal was quantised away. The attacker's options are to push probes further, which measures the response over a wider region and so gives a coarser, more biased direction, or to spend more calls per coordinate to pull a weak signal out. Both raise the invoice. The diagnostic tell is a large share of probes returning identical values. Reporting this correctly matters: the finding is that the endpoint's returned precision priced the estimate up, not that the model resisted anything.

go deeper

for a junior

Remember that the attacker only ever sees the printed score, so a change too small to show up in that printed number is a change they cannot detect at all.

for a middle

Be able to explain the two responses - larger probes or more samples - and why each degrades something: bias in the first case, cost in the second.

for a senior

Demonstrate that you would spot identical score pairs mid-run and re-price rather than burn the budget, and that your report distinguishes a starved measurement from a resistant model.

for a principal

Own the distinction between a control that raises an adversary's cost and one that changes what they can reach, and insist that any black-box result quoted to you states the returned precision it was obtained under.

## The measurement channel, not the model Everything a score-based attacker learns arrives through one narrow channel: the number the endpoint prints back. The direction estimate is a set of differences between such numbers, so whatever the channel cannot represent simply never reaches the attacker. Rounding a calibrated risk score to two decimals means the smallest observable change is one hundredth. Any probe whose true effect on the score is smaller than that returns a value identical to the reference, and the difference computed from that pair is exactly zero. The important reading is that the zero is an artefact of the channel. The model's response in that direction may be perfectly ordinary; the attacker has just been handed a number too blunt to show it. Confusing those two is how a test report ends up claiming robustness that was never measured. ## What the attacker does about it There are two moves, and both cost. **Probe harder.** If small probes are invisible, use larger ones. This works, but it changes what is being measured. A difference taken over a wide separation describes the average behaviour across that whole span, not the local response at the point of interest, so the resulting direction is biased toward whatever dominates over the wider region. In an area where the score bends, that direction can point somewhere useless. It also runs into a hard ceiling on any realistic engagement: probes have to stay inside changes the attacker is actually allowed to make to a transaction, and a probe big enough to clear the rounding step may already be outside what a plausible record can contain. **Spend more calls.** Aggregate over many probes so that the handful which do cross the rounding boundary carry the estimate. This preserves locality but multiplies the sample count per coordinate, and that multiplier lands on top of a bill that already scales with the width of the input and the number of search steps. In practice the attacker mixes both, and the outcome is the same in either case: a coarser direction for more money. Whether that ends the attack is an economic question about the value of one evaded transaction, not a structural one. ## Reading the run correctly The tell is easy to see from the attacker's side of the wire: a large fraction of probe pairs returning byte-identical scores, and an estimated direction that comes back mostly zeros. A tester who does not recognise it can spend a whole budget on samples that measured nothing, then write up `the attack did not converge`. The honest finding has a specific shape, and it separates three different sentences that are constantly collapsed into one: - **The channel was too coarse.** Probes below the printed precision returned no signal; the estimate needed larger probes or more samples, at a stated multiple of the original cost. - **The budget ran out.** The attack was priced above what the engagement funded. This bounds the spend, not the model. - **The model resisted.** Nothing in a quantised run supports this, because the search never got a usable direction to follow. Only the first is what a coarse score establishes. Writing the third when the evidence supports the first is the reporting failure this question exists to catch, and a reviewer who reads a black-box result should always ask what precision the endpoint returned before believing any conclusion about the model. ## The economic frame Seen from a distance, returned precision is one more term in the same cost expression as everything else in this family: price per call, samples per direction, coordinates covered, steps taken. Coarse scores multiply the samples term. They do not remove the attack, they do not change the model, and they do not change what a determined adversary can eventually reach - they change what the attempt costs, and therefore which adversaries find it worth attempting. A result that reads `no evasion found` without stating the precision it worked against has left out one of the numbers that made it come out that way.

  • How would you notice this during a run rather than after it?
    Watch the distribution of probe differences. A healthy run shows small non-zero differences scattered across coordinates; a starved one shows a large share of exactly-equal score pairs and an estimated direction that is mostly zeros. That pattern says the channel is too coarse for the probe size long before the budget is exhausted.
  • Why does simply using larger probes not solve it cleanly?
    A difference taken over a wide separation describes average behaviour across that span rather than the local response, so the direction is biased wherever the score bends. Larger probes also have to stay within changes a real transaction could plausibly contain, and that ceiling can arrive before the rounding step is cleared.
  • What should the write-up say if the run stalled on coarse scores?
    That the endpoint's returned precision starved the estimate and raised the cost by a stated factor, with the probe sizes and sample counts that were tried. It must not say the model resisted the attack: the search never obtained a usable direction, so the run carries no evidence either way about the model's behaviour.

saying these in an interview costs you the question

  • Reads a zero difference as the model being insensitive there
  • Reports a starved estimate as evidence of robustness
  • Ignores that larger probes measure a wider, coarser region
  • Assumes coarse scores end the attack rather than repricing it
  • Omits returned precision from a black-box result write-up

context