skip to content

Domain Constraints

You can hand the library a permitted value range and a mask of which features may move; a categorical or checksum rule it cannot express you must filter for yourself. Interviewers ask which you used.

on this pageshow

explore

questions

5

When you wrap a model for an adversarial-example library such as the Adversarial Robustness Toolbox, you declare a permitted minimum and maximum for the input values. What does that declaration change about the examples the attack generates, and what does it not?

level: juniorimportance: must knowfreq 70%

answer

  1. box constraint, clipped each step
  2. one global min/max, not per-feature
  3. no integrality, no one-hot, no checksum
  4. scale mismatch silently changes strength
  5. image-shaped assumption

basics

~20 s

When you wrap a model for an adversarial-example library you declare the legal minimum and maximum for input values. The attack clips every generated example back inside that box, so pixels stay in range. It is one global range over all features, and it knows nothing about what any individual feature means.

solid answer

~50 s

The declared range is a **box constraint applied after each optimisation step**. Whatever the gradient step produces, the values are pushed back between the minimum and maximum before the next iteration and before the example is returned. That is why an image attack does not hand you pixels at 3.7 or -0.4. What it does not do: - It is **one range for the whole input**, not a per-feature range. A tabular row where age lives in 18-90 and amount in 0-50000 cannot be described by it. - It expresses **no relationships**: no integrality, no one-hot exclusivity, no "this field is derived from that one", no checksum. - It does not stop a feature from moving; it only bounds how far. So the range keeps generated inputs inside the tensor's legal box, and everything about whether the row is a *constructible* record is still your problem.

go deeper

for a junior

Says the range keeps generated values inside the legal input interval, e.g. valid pixel values, and that you must set it to match your preprocessing.

for a middle

Adds that clipping happens every iteration, that it is one global box rather than per-feature, and that perturbation budgets are read relative to it.

for a senior

Treats a range/scale mismatch as the first diagnosis for a dead attack, and knows the box says nothing about record validity outside the image case.

for a principal

Frames it as the library's image-shaped domain assumption, and decides which domains are worth driving through these wrappers at all versus writing a constrained search.

## The object that actually holds the range An adversarial-example library never attacks your model directly; it attacks a **wrapper** you build around it. In the Adversarial Robustness Toolbox (ART) that wrapper is an estimator such as `art.estimators.classification.PyTorchClassifier`, and one of its constructor arguments is `clip_values`, a two-element tuple `(min, max)`. Foolbox spells the same idea `bounds=(0, 1)` on `foolbox.PyTorchModel`. torchattacks does not ask at all: it documents that your model must accept inputs already in `[0, 1]` and assumes it. Three spellings, one statement — *one* pair of numbers describing *every* coordinate of the input tensor. ## The mechanism, step by step An iterative attack such as ART's `ProjectedGradientDescent` repeats three moves: compute the gradient of the loss with respect to the input, take a step along it, project back. The projection is two clamps, not one: 1. back inside the **perturbation ball**, so the example stays within `eps` of the original input; 2. back inside **`clip_values`**, so no coordinate leaves the declared input box. Both run on every iteration, not once at the end. That ordering matters. A final-only clamp would return a different and generally weaker example, because the search would have spent its steps optimising in a region the model is never fed. ## What it costs The clamp itself is free — an elementwise minimum and maximum over a tensor, invisible next to a forward and backward pass. What costs money is getting the value wrong, because nothing tells you. `clip_values` is not checked against the data you pass to `generate()`; you can declare `(0, 1)` and hand over 0-255 images and the run completes normally. Order-of-magnitude: a 1,000-example PGD sweep at 40 iterations against a mid-size vision model is tens of minutes to a few GPU-hours; the black-box attacks in the same library spend thousands to hundreds of thousands of billed queries per sweep against a metered endpoint. A unit mismatch throws all of that away, plus the engineer-day of triage that follows a result nobody can defend. ## Where the number misleads This is the paragraph worth memorising. In ART, `eps` and `eps_step` are expressed in **input units, not as a fraction of the declared range**. So the same call means different things depending on the scale your model is actually fed: | data really in | `clip_values` passed | `eps` passed | what you actually measured | |---|---|---|---| | 0-1 | (0, 1) | 0.03 | the intended budget, roughly 8/255 | | 0-255 | (0, 1) | 0.03 | a perturbation ~255x smaller than intended; success near zero | | 0-1 | (0, 255) | 8 | an effectively unbounded attack; success near 100% | Both failure rows produce a *plausible* number. The first reads as "the model is robust at a small budget" and gets written into a report. The second reads as "the model breaks under a trivial perturbation" and gets written into a different report. Neither is a property of the model; both are a property of a mismatch between the wrapper's declared box and the tensor the model receives. The closely related trap is *where normalisation lives*. If mean/standard-deviation normalisation sits inside the callable you wrapped, the attack optimises in raw-pixel space and the declared range should be the raw one. If it sits outside, the attack sees normalised values that legitimately run from roughly -2.5 to +2.5, and a declared `(0, 1)` box crushes every example on the first projection. ## What one global interval cannot say It carries nothing per feature and nothing relational. A tabular row where `age` lives in 18-90 and `amount` in 0-50,000 has no honest single pair: the union permits an age of 50,000, the intersection forbids any amount above 90. Integrality, one-hot exclusivity, a field derived from another field, a checksum — none of it is expressible. The box is a claim about the *tensor*, not about whether the result is a record any system would accept. ## What I would check before believing the output Print the minimum and maximum of the clean batch you are about to attack and confirm they sit inside `clip_values`. Confirm `eps` is quoted in the same units as that range. Confirm `eps_step` is smaller than `eps`, since an equal or larger step collapses the iterations into repeated single steps. Then run the whole pipeline once at a deliberately absurd `eps`: if success does not go to essentially 100%, your wiring is broken and the model is not robust. That is the cheapest sanity check in the toolkit and it costs one run.

  • You declared the input range as 0-1 but your model is fed 0-255 images. What symptom do you see?
    Examples that look untouched and a near-zero success rate, because the perturbation budget and step size are being interpreted as fractions of a range 255 times smaller than the real one. Clipping also collapses the inputs.
  • Why can a single declared range not describe a tabular credit-application row?
    Because each column has its own interval and its own type. One global pair either permits an age of 50000 or forbids any amount above 90; neither is the real domain.
  • Is clipping applied only at the end of the attack?
    No — iterative attacks project back into the box after each step, so the search itself stays inside it. A final-only clip would be a different, weaker result.

A perturbation budget of 0.03 is a tolerance written on a ruler whose units nobody stated: on a 0-1 scale it is a visible smudge, on a 0-255 scale it is invisible. When the success rate collapses, the ruler changed, not the model.

saying these in an interview costs you the question

  • Believing the declared range makes the output a valid record in any domain.
  • Assuming the library infers the range from the training data if you omit it.
  • Not knowing that the perturbation budget is interpreted relative to that range, so 0-1 versus 0-255 changes the attack's strength.
  • Thinking the range is a per-feature list by default.

context

open as a page

You are attacking a tabular fraud model with an adversarial-example library, and 12 of its 40 features are ones the attacker cannot influence, such as account tenure. The library lets you pass a mask of which features may move. What does that mask guarantee, and what still has to be enforced outside the library?

level: middleimportance: must knowfreq 60%

basics

~20 s

The mask freezes coordinates: features you mark immovable keep their original values, so the search only touches the ones you allow. It says nothing about the movable features' own rules, such as integer counts, one-hot exclusivity, or a field derived from another. Those you check yourself after the library returns its examples.

open as a page

In TextAttack, an attack recipe pairs a transformation that proposes candidate rewrites with a set of constraints those candidates must pass. What role do the constraints play in the search, and what happens to your reported success rate if you relax them?

level: middleimportance: should knowfreq 45%

basics

~20 s

Constraints are filters inside the loop: a transformation proposes candidate sentences, each constraint rejects the ones that violate it, and the search only ever sees the survivors. They are what keeps a rewrite readable and meaning-preserving. Loosen them and the success rate rises, because you are now counting rewrites that changed the sentence.

open as a page

An adversarial-example library returns 1000 perturbed rows against a tabular loan model and reports that 91% are misclassified, but many rows hold fractional values in integer-only columns and two category indicators set at once. Where do you put the rule the library could not express, and what number do you report instead?

level: seniorimportance: should knowfreq 50%

basics

~20 s

Run the library's output through a validity predicate you write, before you score anything. Keep only rows a real applicant could actually submit, re-query the model on those, and report misclassified-and-valid over examples attempted, together with the share you discarded. The 91% was computed before that filter and overstates what is reachable.

open as a page

For a domain rule an adversarial-example library cannot express, you can either discard invalid examples after generation or write a projection into the attack's iteration so every step lands on a legal input. How do you decide which, and what does each let you claim?

level: principalimportance: should knowfreq 35%

basics

~20 s

Filter afterwards for a cheap first read: it is a few lines, but the search optimised in a space you then discard, so survivors are partly luck and the rate is loose. Project inside the loop when the number must mean something: it costs custom code, and you are no longer running a stock attack.

open as a page