When you wrap a trained image classifier so an adversarial robustness library can attack it, the wrapper makes you declare the valid input value range. What does the library do with that range, and what goes wrong if you declare it wrong?
answer
- declared range = feasible set
- clipped after every step
- budget units follow the range
- 0-1 vs 0-255 mismatch
- float example, uint8 file
basics
~20 sIt tells the library the legal bounds of an input, so every perturbed candidate is clipped back into them. Declare 0-255 when your model is actually fed 0-1 and the attack searches a space your service never accepts, so the perturbation size you report means nothing.
solid answer
~50 sThe declared range is the library's only knowledge of what a legal input looks like. Attacks project each candidate back into it after every step, so it both defines the feasible set and gives the perturbation budget its units. If you declare 0-1 for a model that is served 0-255 pixel values, a budget of 8/255 becomes either invisible or a full repaint, and the success rate you report is for a threat model nobody has. The range must match the tensor the deployed model actually receives, at the point where your wrapper hands data over. A second trap sits underneath it: real images are 8-bit integers, so a float example that survives clipping can still lose its perturbation when it is written back to a file. Clipping to the declared range is not the same as being representable in the format the service ingests.
go deeper
Should say the range bounds legal inputs, the attack clips into it, and it must match what the model is actually fed.
Adds that the perturbation budget is expressed in the same units, so a wrong range makes the reported epsilon incomparable.
Raises quantisation and the round-trip test: an example that only exists as a float tensor may not survive being written to the format the service ingests.
Frames it as deliverable integrity — every robustness number in a report must carry the input domain and units it was measured in, or it cannot be compared across engagements.
### Where this argument lives An adversarial robustness library never touches your model directly. It sees an **estimator wrapper**: in Adversarial Robustness Toolbox (ART) that is a class such as `PyTorchClassifier` or `TensorFlowV2Classifier`; in Foolbox it is `foolbox.PyTorchModel`. The wrapper is built from a small set of constructor arguments that together are the library's entire description of the system under test — the callable, the input shape, the number of classes, the loss object, and the legal value range. ART names that last argument `clip_values` and takes a `(minimum, maximum)` pair; Foolbox names it `bounds`. There is no second channel. If the range is wrong, the library is wrong about the model, and nothing in the run will tell you. ```python # ART: the range is the estimator's external input domain, before its own preprocessing clf = PyTorchClassifier(model=net, loss=criterion, input_shape=(3, 224, 224), nb_classes=1000, clip_values=(0.0, 1.0), preprocessing=(IMAGENET_MEAN, IMAGENET_STD)) ``` Note the ordering that trips people: ART applies the `preprocessing` tuple *inside* `predict` and `loss_gradient`, so `clip_values` describes the tensor you hand the estimator, not the normalised tensor the network finally consumes. Declaring the post-normalisation range (roughly -2.1 to 2.6 for ImageNet statistics) when you are feeding it 0-1 images is the same class of error as the 0-1 versus 0-255 mix-up. ### The mechanism Attacks in these toolkits are constrained optimisers. Each iteration computes a step, applies it, and then **projects the candidate back into the feasible set** — into the epsilon-ball around the original input, and into the declared value range. In ART that clip is unconditional; several attack classes refuse to construct at all without `clip_values`, and some derive their per-iteration step size from the range's width. So the declared range does two jobs at once: it is the feasible set, and it is the **unit system** in which the perturbation budget is read. An `eps` of 8/255 means one thing against a 0-1 domain and something 255 times different against a 0-255 domain. The library will not object; it has no independent view of what a pixel is. ### What it costs The argument itself is free — two floats. Getting it wrong costs the whole run. A PGD-style iterative attack at 40 iterations over 1,000 test images is on the order of 40,000 forward and 40,000 backward passes; that is minutes on a GPU and an hour or more on CPU, plus the engineer-hours spent writing up a number that has to be withdrawn. The re-run is the cheap part. The expensive part is a delivered robustness figure whose units nobody can reconstruct three months later. ### Where the number misleads **Range declared too wide** (you said 0-255, the model is fed 0-1). The clip never binds. The attack is free to push values to 40 or 200 on a network that has never seen anything above 1.0, activations saturate, and essentially every image flips. Attack success rate approaches 100 percent and reads as a devastating finding. It is a finding about inputs the serving path cannot even produce. **Range declared too narrow** (you said 0-1, the model is fed 0-255). Every candidate — and the clean input, since the attack clips it too — is crushed into a near-black image. Clean accuracy collapses first. If your success metric counts only rows the model classified correctly to begin with, the denominator collapses with it, and a huge success rate is computed over a handful of surviving examples with an enormous confidence interval. **Representability.** The declared range is continuous; the artefact usually is not. Pixels are uint8, audio samples are quantised, a tabular field may be an integer count or a one-hot column that has to stay one-hot. A perturbation that lives between two quantisation levels vanishes on the round-trip to the wire format. A float tensor you never serialised and re-read is not evidence that an attacker can deliver anything. ### What you check before believing the run Draw a real batch from the actual serving path, print `x.min()` and `x.max()`, and confirm the pair you declared brackets them and nothing wider. Then compare clean accuracy through the wrapper against the live service on the same inputs — if they disagree before you attack anything, the domain is wrong and every downstream number is void. Finally, take one adversarial example the harness scored as a success, serialise it exactly as an attacker would have to deliver it, load it back through your own ingestion path, and re-query. If the label reverts, the range or the quantisation is the reason, not the attack.
- Your wrapped model takes a normalised tensor, but the service takes a raw uploaded image. Where should the normalisation live?Inside the wrapper, as part of its preprocessing, so the declared input domain is the raw image and the attack optimises in the space the attacker actually controls.
- How would you catch a quantisation problem without reading the library's source?Round-trip a successful adversarial example through the real serialisation path, load it back and re-query the model. If the label flips back, the perturbation was not representable.
Quoting a perturbation budget without the declared input range is like quoting a price without a currency: 8 is cheap or ruinous depending on a unit the number does not carry.
saying these in an interview costs you the question
- Treats the range as cosmetic metadata the library does not really use.
- Cannot say what clipping does during the attack loop.
- Reports an epsilon without stating the units or the input domain it applies to.
- Assumes a successful float tensor automatically survives being saved as an image file.