An adversarial-robustness toolkit such as the Adversarial Robustness Toolbox groups its attack classes under headings like evasion, poisoning, extraction and inference. Your engagement grants only query access to a deployed model and explicitly forbids touching its training data or pipeline. Which headings are off the table, and what must you check before picking any class?
answer
- headings = preconditions, not difficulty
- poisoning needs a write path plus a retrain
- evasion needs inference-time input access
- extraction emits a surrogate: permission first
- access, data type, artefact, grant
basics
~20 sPoisoning and backdoor classes are off the table: they need write access to training data or the training pipeline plus a retrain, which you were not granted. Evasion classes fit query access. Extraction and inference classes fit technically but produce a surrogate copy or membership claims, so confirm they are in scope and contractually allowed first.
solid answer
~50 sThe headings are not a difficulty ranking; they are a list of prerequisites. **Poisoning and backdoor** classes assume you can insert or modify training rows and that someone retrains — no training-data path, no run, and a synthetic retrain you did yourself proves nothing about the deployed model. **Evasion** classes assume inference-time input access, which is what you have. **Extraction** classes assume a large query budget and hand you a surrogate model — an artefact with legal and contractual weight, so it needs explicit permission and a disposal plan. **Inference/privacy** classes assume a query budget plus reference data you are allowed to hold. So before picking a class, check three things: the access it needs, the data type it accepts, and whether the artefact it produces is something your rules of engagement let you create and keep.
go deeper
Should say that poisoning needs training-data access and a retrain, evasion needs only query access, and that you check the class's prerequisites against the rules of engagement before running anything.
Adds that extraction and inference classes are technically runnable with queries but produce artefacts needing explicit permission, and that a blocked heading becomes a stated coverage gap.
Frames the catalogue as attacker positions in the model lifecycle, and insists every finding names the access level it assumed so the client can act on it.
Turns the pre-check into policy: a scoping template that maps granted access paths to permitted attack headings, with artefact handling and retention agreed up front.
### The catalogue is a list of attacker positions, not a difficulty ladder An adversarial-robustness toolkit — the Adversarial Robustness Toolbox (ART) is the canonical example — ships its attacks grouped under headings: evasion, poisoning, extraction, inference. A *heading* here is not a measure of how hard or how strong the attack is. It is a statement about **where in the model's lifecycle the attacker is standing** when the attack happens, and therefore about what the attacker must already be able to do. Read it that way and class selection stops being a matter of taste. | Heading | Attacker position | Hard precondition | What the run costs | What it emits | |---|---|---|---|---| | Evasion | Inference time, on the input | Submit an input, observe some output | Queries (black-box) or local compute (white-box) | Perturbed inputs + a success rate | | Poisoning / backdoor | Upstream, in the data or training path | A **write path** into training data *and* a retrain that actually occurs | The retrain — GPU hours and wall-clock the client controls, not you | A model whose behaviour changed | | Extraction | Query budget, rebuilding an approximation | A large query allowance and explicit permission | Often tens of thousands of metered queries; real money on a paid endpoint | **A surrogate model** — a derived copy of the client's asset | | Inference / membership | Query budget plus candidate records | Queries *and* records you are permitted to hold and test | Queries, plus reference data you may have to obtain | Claims about specific records, often personal data | ### Applying it to the stated scope Query access to a deployed model, training data and pipeline explicitly out of bounds: - **Evasion** fits. It is the heading whose precondition your grant actually satisfies. - **Poisoning and backdoor** are off the table. Both preconditions fail: no write path, and no retrain you can cause. A retrain is not something you can substitute for — it is the mechanism by which a poisoned row becomes a poisoned model. - **Extraction and inference** are *technically* runnable from queries, but their preconditions include a permission, not just an access. Both emit an artefact with legal weight — a copy of the model, or an assertion that a named person's record was in the training set. ### Where the number misleads The dangerous move is not refusing an out-of-scope heading; it is **substituting a lab**. A tester who cannot poison the client's pipeline trains a comparable model locally, poisons that, and reports a backdoor success rate. Every number in that run is real and every number is about the tester's own laptop. The client's deployed model differs in architecture details, training data, data-cleaning steps, canary checks, and whatever review sits between a contributed row and a shipped checkpoint — and it is exactly those differences the finding claims to have tested. The report says "backdoor achieved, 98%"; the honest sentence is "a model I built was backdoorable, and I do not know whether yours is." The second misleading reading is **silence**. A report that lists evasion results and says nothing about poisoning is read by a client as *tested and clean* across the board. Absence of a heading is absence of evidence, and only an explicit coverage statement makes that visible. ### What you check before picking any class For the class you intend to run, write down four things and the line of the rules of engagement that grants each: 1. **Access level required** — weights and gradients, full confidence scores, or bare labels. 2. **Data type accepted** — dense continuous tensors, tabular rows with constraints, or text. 3. **Artefact emitted** — perturbed inputs, a surrogate model, membership claims — and its retention and disposal rules. 4. **Cost shape and ceiling** — queries against a rate cap and a per-call price, or GPU hours against a wall-clock deadline. If any of the four has no grant behind it, either the run is out of scope or its result is unactionable: it describes an attacker position nobody has, and the client cannot reproduce it or remediate against it. When a heading is blocked, the correct output is a named coverage gap in the report — "poisoning and backdoor risk not assessed; no training-data write path was in scope" — not a lab-only run whose framing lets a reader assume the deployed system was tested.
- The client has a user-feedback loop that is periodically used for fine-tuning. Does that change your answer about poisoning classes?Yes — that feedback loop is a write path into training data, so poisoning becomes in-scope in principle. But you still need the retrain to actually occur inside the engagement window, or agreement that a staged retrain on the client's own pipeline counts as the test.
- If no poisoning run is possible, what goes in the report?An explicit coverage statement: poisoning and backdoor risk was not assessed, because no training-data write path was in scope. Silence reads as 'tested and clean'.
- Why is 'the artefact it emits' part of the pre-check and not an afterthought?Because extraction and membership-inference runs produce a derived model or claims about specific records — things with retention, disclosure and legal consequences that must be agreed before you generate them, not after.
The headings are like the entry requirements on a job posting, not the salaries. Picking 'poisoning' because it sounds strongest is like applying for a role that requires a licence you do not hold: the ambition is irrelevant, the precondition decides.
saying these in an interview costs you the question
- Treating the catalogue headings as interchangeable difficulty tiers and picking whichever class has the best-known name.
- Running a poisoning class against a locally trained copy and reporting it as a finding against the deployed model.
- Building a surrogate via an extraction class without checking that the rules of engagement permit creating and retaining it.
- Not being able to say what a chosen class needs from the target before running it.