For a domain rule an adversarial-example library cannot express, you can either discard invalid examples after generation or write a projection into the attack's iteration so every step lands on a legal input. How do you decide which, and what does each let you claim?
answer
- survival rate is the decision variable
- filter = cheap, loose lower bound
- projection = your attack now, not stock
- non-convex legal set breaks guarantees
- discrete domain: change instrument
basics
~20 sFilter afterwards for a cheap first read: it is a few lines, but the search optimised in a space you then discard, so survivors are partly luck and the rate is loose. Project inside the loop when the number must mean something: it costs custom code, and you are no longer running a stock attack.
solid answer
~60 sThe two differ in **where the constraint meets the search**. *Filter after*: the attack explores the relaxed continuous space, you keep the legal survivors. Cost: an afternoon. Claim: "at least this many legal adversarial inputs exist within this budget." Weakness: on domains where legal points are sparse — heavily categorical, checksummed, or strongly correlated data — the survival rate collapses toward zero and the run tells you almost nothing about robustness. *Project in the loop*: after each step you map the iterate back onto the legal set, so the search only ever moves through constructible inputs. Cost: you write and validate the projection, and a non-convex legal set makes it approximate, so the stock attack's convergence behaviour no longer applies. Claim: a genuine constrained lower bound, comparable across models you evaluate the same way but not with any number the stock attack produced. Decide on survival rate: if filtering keeps most examples, filtering is enough. If it destroys them, the relaxation is not modelling your domain and you either build the projection or change instrument to a discrete search.
go deeper
Knows that invalid examples have to be dealt with, and would filter them out.
Contrasts filtering with projecting in the loop and can name the cost of each.
Uses the survival rate to choose, validates the projection, and states the lower-bound caveat in the report.
Allocates engagement budget across models on that basis, considers replacing the instrument entirely for discrete domains, and owns the comparability claim the organisation makes from the result.
## Frame it as budget allocation, because that is what it is Both options enforce the same rule; they differ only in **where the rule meets the search**, and therefore in what the resulting number is allowed to claim. *Filter after generation.* The stock attack explores the relaxed continuous space; you keep the legal survivors and re-query the model on them. Cost: a few hours, mostly spent wiring in the validity predicate you should be reusing from production. Claim: at least this many legal adversarial inputs exist within this budget — a loose lower bound. *Project inside the loop.* After each optimisation step you map the iterate back onto the legal set, so the search only ever moves through constructible inputs. Cost: engineering days to write and validate the projection, plus the review cost of putting your own untested code inside a security claim, plus GPU-time overhead because projecting onto a non-convex or discrete legal set is approximate and can stall convergence. Claim: a genuine constrained lower bound, comparable across models you evaluate the same way and comparable with nothing else. ## The measurement that decides it Run the cheap path first and record the **survival rate** — the share of generated examples that pass the validity predicate. That single number is the decision variable. High survival means the continuous relaxation is close to your domain; the filtered result is a usable lower bound and the remaining hours belong somewhere else in the engagement. Near-zero survival means the search spent essentially all of its effort in illegal space and the handful of survivors are accidents rather than evidence about the model. That is the trigger for investment, and it is a much better trigger than intuition, because domains where legal points are sparse — heavily categorical data, checksummed records, strongly correlated columns — do not announce themselves. ## What the investment buys, and what it costs you A projection step turns the library's attack into **your** attack. You gain the only version of the question a business actually asks: can a real, constructible input flip this model under these capability limits. You give up three things. **Comparability.** Any published or internal figure produced by the unmodified attack is now measuring a different thing over a larger feasible set. A leaderboard is comparable precisely because every entry runs the same fixed suite under the same stated threat model; the moment you edit the loop you have left that frame, and reporting your number beside one of theirs is wrong in both directions. **The library's testing.** The projection is unreviewed code sitting in the critical path of a security claim. Budget for round-trip tests: a legal point must project to itself, and every projected point must pass the same predicate the production intake uses. **Convergence behaviour.** The stock attack's step-size and iteration defaults were tuned for an unprojected search. With an approximate projection they may stall, and a stalled search reports a low rate that reads as robustness. ## The third option people forget When the legal set is fundamentally discrete, a continuous gradient attack with a projection bolted onto it is the wrong instrument, and a search over legal edits — mutating only the allowed operations and scoring each candidate with a model query — is both simpler to justify and closer to real attacker capability. Note that this also swaps the cost model: from GPU hours to model queries, which is a bill if the target is metered and a rate-limit problem if it is not. ## Where either number misleads Both are **lower bounds on attackability, never upper bounds on robustness**. A filtered result additionally understates because the search was never told about the constraints; a projected result understates because your projection is approximate. Neither licenses the sentence "the model is robust". What I make the report say: which of the three instruments was used, the survival or projection-validity rate, that the figure is a lower bound, and explicitly that a constrained result is not comparable to any figure the unmodified library produced. One defensible constrained number on the model that matters beats four stock numbers across the portfolio that a reviewer can dismiss with a single question.
- What single measurement drives the choice?The share of generated examples that survive the validity filter. High survival, keep filtering; near-zero survival, the relaxation is not modelling the domain and you invest in a constrained search.
- Why is a constrained result not comparable with a stock-attack number?Because it is a different attack over a smaller feasible set. The two answer different questions, and the constrained one usually reports lower success while being the more meaningful figure.
saying these in an interview costs you the question
- Comparing a constrained, hand-modified attack's result to a stock or published figure as though they measure the same thing.
- Building a projection before measuring how many examples the cheap filter actually loses.
- Shipping a projection with no round-trip validation against the production validity rules.
- Presenting either result as an upper bound on robustness.