skip to content

How does an attacker optimise through a cleaning stage that has no usable gradient, and what does that cost them?

level: middleimportance: should knowfreq 44%

answer

  1. the forward pass stays honest
  2. only the direction needs an approximation
  3. a transform that keeps content is near-identity
  4. randomised transforms get averaged over draws
  5. the price is steps, paid once

basics

~20 s

They replace the non-differentiable stage with a smooth stand-in when computing their search direction, and the natural stand-in is the identity because a content-preserving transform barely changes its input. The cost is extra optimisation steps, offline and cheap.

solid answer

~50 s

Stages like bit-depth reduction, rounding or a lossy re-encode have derivatives that are zero almost everywhere or meaningless, so a search that needs a direction through the pipeline stalls. The adversary's answer is to keep the real transform in the forward direction - so the change is scored against what the pipeline actually does - and substitute a differentiable approximation of it when working out which way to move. The identity is usually good enough, precisely because a cleaning transform is designed to leave content intact and is therefore close to the identity by construction. In the literature this backward-pass approximation is known as BPDA. If the transform is randomised per request, they average the direction over draws so the update works on average rather than for one draw. Either way the defence converted a wall into a bill, payable in optimisation steps on the attacker's own hardware.

go deeper

for a junior

Recall that an attack needs a direction through everything between the submission and the decision, and that a stage with no usable derivative interrupts the search rather than removing the vulnerable inputs.

for a middle

Explain the split between scoring a candidate with the real transform and choosing a direction with a smooth stand-in, and say why a content-preserving transform is close to the identity and therefore easy to stand in for.

for a senior

Be ready to turn this into an evaluation plan: same norm, same radius, same success criterion, transform inside the attacker's loop, and a reported cost multiple instead of a robustness number.

for a principal

Decide what a measured cost multiple is worth next quarter, knowing it is a one-off bill for the attacker and a standing per-prediction cost for you, and whether that trade justifies keeping the claim in a customer-facing document at all.

## The problem the attacker faces An attack search needs to know which way to move an input to make the model's decision worse. That direction comes from how the loss responds to changes in the *input*, propagated back through everything between the submission and the decision. A cleaning stage sits in that path. If the stage is a smooth function, nothing changes conceptually - the attacker's direction now accounts for the stage and the search continues. The interesting case is the stage that is deliberately or incidentally non-smooth: quantising to fewer bits, rounding to integer pixel values, median filtering, a lossy re-encode with block-wise quantisation. Their derivative is zero almost everywhere, with jumps in between, so the direction that comes back is either zero or uninformative and the search appears to stop. Defences of this class are often reported on that basis: "the attack fails against our pipeline". What failed is the *search*, not the existence of inputs that break the model. ## The property that rescues the search A cleaning transform has a design requirement that is also its weakness: it must leave the decision-relevant content intact, or the model's clean accuracy falls apart and the defence is unshippable. "Leaves the content intact" is a strong statement. It means the transform's output is close to its input on exactly the part of the signal that matters. A function that is close to the identity is well approximated, for the purpose of choosing a search direction, *by* the identity. So the adversary splits the two roles the pipeline plays in their search: - **scoring** - is this candidate actually misread? - uses the real transform, so the answer is honest about what the deployed pipeline does; - **direction** - which way should I move next? - uses a smooth stand-in for the awkward stage. The published name for this backward-pass approximation is BPDA. The point to carry into an interview is not the acronym but the reason it works: the defender chose a transform that changes as little as possible, and "changes as little as possible" is the same statement as "is easy to approximate". Where a stage is not close to the identity - a heavy reconstruction, say - a smooth stand-in fitted to imitate it plays the same role. The stand-in never has to be faithful everywhere; it only has to be good enough near the inputs being searched over to keep pointing roughly the right way. ## Randomised transforms A common upgrade is to randomise: draw the resize factor, the crop, the noise level or the quantisation offset per request. Now the pipeline is a different function on every submission and a direction fitted to one draw is partly wasted on the next. The adversary's answer is to stop optimising for one draw: average the direction over several draws of the randomness so the candidate is pushed toward one that works in expectation, and check candidates against fresh draws. This costs more work per update - roughly in proportion to how many draws are averaged - and it lowers the reliability of any single submission. It does not change what is achievable; it changes the price, and it charges the defender too, since a randomised pipeline gives non-reproducible decisions on identical inputs, which is unattractive anywhere decisions are audited. ## What it costs, stated as a multiple The useful way to report all of this is not "the defence was broken" but "the defence costs an adaptive attacker *k* times the steps". Measure it the way you would measure anything else: fix the norm and radius, fix the success criterion, run the same attack against the bare model and against the defended pipeline with the stage inside the attacker's loop, and report the steps or queries needed to reach the same success rate on each. The number is usually a small multiple rather than an order of magnitude, and it is paid once per pipeline, offline, on hardware the defender does not control. Against that, the defender pays a per-prediction transform cost and a permanent clean-accuracy loss. That asymmetry - a one-off bill for the attacker, a standing rate for the defender - is the reason this defence class is described as masked rather than fixed. ## What this does not mean It does not mean preprocessing is pointless. Normalisation and format handling are ordinary engineering, and removing sensor artefacts may genuinely improve accuracy. It means the *robustness* claim attached to the stage has to be earned by an evaluation in which the attacker optimised through it, and stated as a cost, not a guarantee.

  • Why doesn't randomising the transform per request close the gap?
    It raises variance, and variance is answered by averaging: the attacker fits a change that works across draws rather than for one draw, paying more work per update. It also charges the defender, because identical inputs now get non-identical decisions, which is hard to defend anywhere the decision is audited or appealed.
  • How would you measure the cost multiple honestly?
    Fix the norm, the radius and the success criterion. Run the same attack against the bare model and against the defended pipeline with the cleaning stage inside the attacker's optimisation, and report steps or queries to reach equal success. Report that ratio, not a robust-accuracy figure produced by an attacker who ignored the stage.
  • If the stage is genuinely hard to approximate, has the defence worked?
    It has raised the bill further, and that is worth reporting - but hardness of approximation is not a property you can bound. It is measured only by how hard your evaluators tried. Treat a large measured multiple as a finding with an expiry date, not as a limit on what an adversary can do.

saying these in an interview costs you the question

  • Says a non-differentiable stage makes the attack impossible
  • Thinks the attacker must abandon the real transform when scoring candidates
  • Believes randomising the transform removes the attack rather than raising its cost
  • Confuses the gradient with respect to inputs with the one used in training
  • Reports a broken search as evidence of a robust model

context