An attack run against your wrapped image classifier flips almost every test image, but the same adversarial files fail when uploaded to the live service, which resizes and re-encodes each upload before inference. What is wrong with the wrapper, and how do you fix it?
answer
- wrapper input != service input
- resize and re-encode kill the perturbation
- boundary at the attacker's control point
- non-differentiable steps stay in predict
- round-trip assertion on every success
basics
~20 sThe wrapper exposes the bare model, so the attack optimised a tensor the live service never receives. The resize and re-encode step sits outside the wrapper and destroys the perturbation. Fix it by moving the whole serving preprocessing chain inside the wrapper, so the attack's input is the uploaded file rather than the tensor.
solid answer
~50 sYou attacked a model nobody serves. The service's input is a file; your wrapper's input was a decoded, resized, normalised tensor several steps downstream. Everything between those points — decode, resize, crop, lossy re-encode, quantisation, normalisation — is a filter the perturbation must survive, and a high-frequency perturbation tuned on the tensor usually does not survive resampling or lossy compression. Redraw the wrapper boundary at the service's own input: the wrapper takes the file, runs the identical preprocessing the service runs, then the model. The declared input domain becomes the raw image, and the budget is expressed in a space the attacker controls. That raises a real problem: parts of the chain are not differentiable. Use a differentiable approximation and say so, or keep the faithful chain and drive attacks that need only predictions. Do not quietly drop the step to keep gradients flowing — that is how you got here.
go deeper
Should spot that the harness and the service feed the model different things, and that the missing preprocessing is the suspect.
Names the specific destroyers — resampling and lossy re-encode against a high-frequency perturbation — and proposes moving preprocessing into the wrapper.
Confirms the mechanism with a controlled round-trip test, then handles the differentiability split explicitly rather than dropping the step, and expects the honest success rate to fall.
Institutionalises it: the wrapper boundary is defined by where the attacker's control starts, and every reported success must survive the real ingestion path before it is counted.
### The diagnosis This is the most common way an adversarial robustness harness produces a number that means nothing: a **fidelity gap** between the wrapper's input and the service's input. The service's input is a file of bytes. Your estimator wrapper's input was a decoded, resized, normalised float tensor several steps downstream of that. Everything in between — decode, resize or crop, colour conversion, lossy re-encode, integer quantisation, normalisation — is a filter that a perturbation has to survive, and it was never in the optimisation. A perturbation tuned freely on the tensor concentrates energy wherever the loss falls fastest, which is typically high spatial frequency; resampling and lossy compression are precisely the operations that discard high spatial frequency. The harness's near-total success rate and the service's near-total failure are the same fact seen from two sides. Work it in this order rather than guessing: 1. Write down the service's real path, byte for byte, from what it accepts to what the model consumes. Include steps hiding in a CDN transform, a client SDK, or an upload-size limit. 2. Write down the wrapper's path. The set difference is your suspect. 3. **Confirm the mechanism instead of assuming it.** Take one adversarial example the harness scored as a success, push it through only the missing steps, and re-query through the wrapper. If the label reverts, those steps are the cause. This is what separates a preprocessing gap from the other candidates — a stale checkpoint, a different model version behind the endpoint, a quantised export. Rule out the checkpoint separately by comparing *clean* predictions between wrapper and service on identical inputs: matching clean answers with diverging adversarial ones points at preprocessing, not weights. ### The fix, and its differentiability split Redraw the wrapper boundary at **the point where the attacker's control begins**. For an upload endpoint that is the file. The declared input domain becomes the raw image, and the perturbation budget is expressed in the space the attacker actually manipulates. Then split the chain: - **Differentiable steps** (resize, crop, colour conversion, normalisation) go inside the wrapper *and* inside the gradient path, so the attack optimises through them. In ART this is what the estimator's own `preprocessing` argument and the `Preprocessor` API are for. - **Non-differentiable steps** (lossy encode, integer quantisation, format conversion) still go inside the prediction path, so success is scored honestly. For the gradient path you either supply an approximate backward pass — ART's `Preprocessor` exposes `estimate_gradient` for exactly this — and record in the report that you did, or you drop gradients for this evaluation and drive attacks that need only predictions. What you must not do is delete the awkward step to keep gradients flowing. That is the move that produced the original wrong number. ### What it costs Two distinct bills. Engineering: reimplementing the serving chain faithfully, often across a language boundary, is days rather than hours, and it is a maintenance liability because the chain drifts. Compute: every candidate now pays decode, resize and encode on top of the forward pass, and an iterative attack evaluates thousands of candidates per example, so wall-clock per example can rise by an order of magnitude. If the chain contains randomness — a random resized crop, for instance — a single evaluation is no longer a reliable signal, and the standard remedy of averaging over sampled transforms multiplies the per-step cost again by the number of samples. Budget for the wide wrapper to run on a smaller sample than the narrow one did. ### Where the new number misleads Expect the success rate to fall, sometimes from ninety-something percent to single digits, and the budget needed to achieve anything to grow. **That drop is not a regression; the earlier number was the fiction.** But do not over-trust the new one in the other direction either. Because part of the chain is now approximated or opaque to gradients, the attack driving it is weaker than a determined adversary's would be. The corrected figure is a **floor** on what is achievable, not a bound on what is possible. Reporting it as "the service is robust" repeats the original sin with the sign flipped. ### What you would check, permanently Make the round-trip an assertion rather than a habit: an example is only counted as a success after it has been serialised through the real ingestion path, re-read, and re-scored. Fail the harness loudly when the wrapper's clean predictions stop matching the live service's, because that is the signal that a CDN transform, a resize filter or an upload limit changed underneath you — and the alternative is a harness that keeps quietly reporting a comfortable number about a pipeline that no longer exists.
- Which perturbations tend to survive a resize and lossy re-encode, and which do not?High-frequency, per-pixel noise is largely removed by resampling and lossy compression. Lower-frequency, spatially larger or structured perturbations survive better, at the cost of being more visible.
- You use a differentiable approximation for the re-encode step. What must the report say?That gradients were taken through an approximation of a non-differentiable step, and that every reported success was re-scored through the exact serving chain on the prediction path.
- How would you rule out a stale checkpoint as the cause instead?Query the live service and your wrapper with the same clean inputs and compare outputs. Matching clean predictions but diverging adversarial ones point at preprocessing, not weights.
saying these in an interview costs you the question
- Removes the awkward preprocessing step so gradients keep flowing, and reports the old number.
- Blames the model or the endpoint without testing the preprocessing hypothesis in isolation.
- Treats the drop in success rate as a bug to be tuned away rather than as the corrected result.
- Claims the harness result still holds because the perturbation is 'small enough' to survive compression, without testing it.