skip to content

Scrubbing Before Inference

Compressing, quantizing or reconstructing an input before the model sees it, and why an attacker optimizing through the step pays almost nothing extra. Interviewers ask because it sounds free.

on this pageshow

explore

questions

4

A vision pipeline denoises and re-encodes every submitted image before the classifier sees it - which attacker does that stop, and which does it merely charge?

level: juniorimportance: must knowfreq 58%

answer

  1. a perturbation is fitted to a function
  2. change the function, refit the direction
  3. it was never built to survive this
  4. cleaning that keeps the signal keeps the attack
  5. a bill, not a wall

basics

~20 s

It stops an attacker whose perturbation was fitted against the bare classifier and never had to survive the cleaning step. An attacker who knows the step is there fits a perturbation through it and pays only extra optimisation steps.

solid answer

~50 s

A cleaning stage - denoise, blur, reduce bit depth, resize, lossy re-encode - is deployed on the theory that an adversarial perturbation is small and fragile, so a lossy transform wipes it out. That is true of a perturbation built against the classifier alone: a perturbation is a *direction* fitted to one specific function, and if you change the function the direction no longer points anywhere useful. It is not true of an attacker who knows the stage exists. Their target is the whole chain, cleaning stage included, so they fit a change that survives it. They can do that with nothing but the ordinary submission path plus a local copy of the transform, which they run offline. The premium is extra optimisation steps, not a wall. So the honest claim is: it defeats non-adaptive attackers and reused published perturbations, and it charges an adaptive one a modest multiple.

go deeper

for a junior

Be ready to say what a perturbation is fitted to. Recall that it is a direction computed against a specific function, so a cleaning step defeats one built for a different function and not one built for this pipeline.

for a middle

Explain the composition: the attacker's real target is the cleaning stage and the model together, and that function is available to them offline. Say why the defender's dial is bounded on both sides.

for a senior

Show that you would re-run the evaluation with the transform inside the attacker's optimisation before letting the number into a review, and that you would report a cost multiple rather than a robustness claim.

for a principal

Own the framing that this class of defence buys a cost increase against an adaptive adversary and an accuracy loss against your own users, and decide whether that trade is worth funding for this deployment at all.

## What a cleaning stage is An input-cleaning (preprocessing, or "purification") defence sits between the request and the model. Every input is transformed before the classifier sees it: smoothed or median-filtered, resized, quantised to fewer bits, re-encoded lossily, or reconstructed and re-emitted. The reasoning behind it is intuitive and half correct. An adversarial perturbation is a small, precisely structured change; lossy transforms are exactly the thing that destroys small, precise structure; therefore the model receives a clean input. ## Why the first half of that reasoning is right An adversarial perturbation is **not noise**. Random change of the same magnitude essentially never flips a trained classifier - that contrast is the single most useful fact in this area. What flips it is a *direction*, read off how the model's loss responds to changes in the input, and that direction is fitted to one particular function. If an attacker fitted their direction to the bare classifier `f`, and you actually deploy `T` then `f`, the change they crafted is being evaluated at a point it was never fitted for. Much of its precise structure does not survive `T`, and what arrives behaves closer to noise than to an attack. Measured against that attacker, the defence really does work, and the numbers people report are real numbers about a real (if lazy) adversary. ## Why the second half is wrong The adversary's target was never `f`. It is whatever function turns their submission into a decision, and that function is `f` composed with `T`. Nothing about `T` is out of reach: - it is code, usually inherited, often documented, sometimes visible in a client library or an integration guide; - where it is not documented it is guessable, because the space of ordinary cleaning chains is small - resize to a fixed input size, normalise, re-encode; - and it runs offline on the attacker's own machine at no cost to them. Once the attacker includes `T` in what they optimise against, the perturbation they produce is by construction one that survives `T`. Cleaning has become a *constraint on the search*, not a barrier to it. Constraints cost steps. ## Why the window is so narrow The defender's dial is bounded on both sides, and both bounds are tight. `T` must be strong enough to destroy a perturbation of the size the threat model allows, and weak enough to leave the decision-relevant content intact - otherwise clean accuracy collapses and the deployment is worse off with the defence than without it. But "leaves the decision-relevant content intact" means `T` is close to the identity on precisely the part of the input the model reads. That surviving part is the channel the attacker now works in. There is no setting of the dial where the content the model needs gets through and nothing an adversary can shape gets through with it. ## What it costs each side The asymmetry is what makes this a weak defence rather than a cheap one. The defender pays a *rate*: every prediction runs the transform, and every marginal input - the faint defect, the low-contrast case - is closer to being misread than it was. The attacker pays a *one-off*: a bounded amount of offline work to make their search account for the transform, after which every subsequent attempt against that pipeline costs roughly what an attempt against the undefended model cost. Reported adaptive evaluations of this defence class collapse it to near the undefended model's numbers. ## How to state the claim honestly Do not write "the pipeline destroys adversarial perturbations". Write something you can defend: - it removes attacks that were not built against this pipeline, including copies of published perturbations; - it raises an adaptive attacker's cost by a measured multiple, which you should actually measure rather than assume; - it does not change the worst case, and it charges clean accuracy for what it does buy. And evaluate it the way you would state it: the attack you run against the defended pipeline must be an attack that includes the cleaning stage. An evaluation in which the attacker ignored the defence is an evaluation of an attacker who did not read your code, and no deployed pipeline gets to assume that adversary.

  • Does your answer change if the cleaning chain is proprietary and the attacker cannot read it?
    It raises the bill, not the class. Ordinary cleaning chains are drawn from a small space, and the endpoint's own behaviour on probe inputs narrows it further; documentation, client code and contractors leak the rest. Treat secrecy of the transform as friction you cannot audit, and never report it as robustness - the number you quote should assume the adversary has read it.
  • What is the honest one-line claim for a cleaning stage in a security review?
    "It removes attacks not built against this pipeline and raises an adaptive attacker's cost by a measured multiple; it does not change the worst case, and it charges clean accuracy." If you cannot supply the multiple, say the defence is unevaluated rather than quoting a number produced by an attacker who ignored it.
  • Why does random noise of the same size not flip the classifier when the perturbation does?
    Noise spreads its magnitude in an arbitrary direction, and almost every direction leaves the decision unchanged. The perturbation spends the same magnitude along the one direction the model's loss is most sensitive to. That is why "we add noise" and "we remove noise" are both weak answers to a structured, direction-seeking adversary.

It is a lock that only defeats burglars who did not know the door was there. Anyone who has seen the door plans around it once, and walks in every night after that.

saying these in an interview costs you the question

  • Says lossy re-encoding destroys any adversarial perturbation
  • Treats an adversarial perturbation as small random noise
  • Assumes the attacker does not know the preprocessing chain exists
  • Quotes robustness measured with the defence outside the attacker's loop
  • Calls a cost increase a guarantee

context

open as a page

How does an attacker optimise through a cleaning stage that has no usable gradient, and what does that cost them?

level: middleimportance: should knowfreq 44%

basics

~20 s

They replace the non-differentiable stage with a smooth stand-in when computing their search direction, and the natural stand-in is the identity because a content-preserving transform barely changes its input. The cost is extra optimisation steps, offline and cheap.

open as a page

You turn an inspection line's input-cleaning stage up until an attacker's perturbations stop landing - what did that buy?

level: seniorimportance: should knowfreq 36%

basics

~20 s

A false-reject and re-image rate that lands on your hardest real parts, plus a cost increase for an adaptive attacker. The window between destroying a perturbation and destroying the defect signal is narrow, and the attacker works inside whatever survives.

open as a page

You inherit a pipeline documented as 91% accurate under attack with its cleaning stage enabled - what do you ask before trusting that?

level: seniorimportance: nice to knowfreq 27%

basics

~20 s

Ask whether the cleaning stage was inside the attacker's optimisation, and what norm, radius, steps and restarts the attack used. If every row was measured against an attacker who ignored the stage, the figure describes that attacker, not the pipeline.

open as a page