skip to content

An adversarial-robustness toolkit such as the Adversarial Robustness Toolbox groups its attack classes under headings like evasion, poisoning, extraction and inference. Your engagement grants only query access to a deployed model and explicitly forbids touching its training data or pipeline. Which headings are off the table, and what must you check before picking any class?

level: juniorimportance: must knowfreq 55%

answer

  1. headings = preconditions, not difficulty
  2. poisoning needs a write path plus a retrain
  3. evasion needs inference-time input access
  4. extraction emits a surrogate: permission first
  5. access, data type, artefact, grant

basics

~20 s

Poisoning and backdoor classes are off the table: they need write access to training data or the training pipeline plus a retrain, which you were not granted. Evasion classes fit query access. Extraction and inference classes fit technically but produce a surrogate copy or membership claims, so confirm they are in scope and contractually allowed first.

solid answer

~50 s

The headings are not a difficulty ranking; they are a list of prerequisites. **Poisoning and backdoor** classes assume you can insert or modify training rows and that someone retrains — no training-data path, no run, and a synthetic retrain you did yourself proves nothing about the deployed model. **Evasion** classes assume inference-time input access, which is what you have. **Extraction** classes assume a large query budget and hand you a surrogate model — an artefact with legal and contractual weight, so it needs explicit permission and a disposal plan. **Inference/privacy** classes assume a query budget plus reference data you are allowed to hold. So before picking a class, check three things: the access it needs, the data type it accepts, and whether the artefact it produces is something your rules of engagement let you create and keep.

go deeper

for a junior

Should say that poisoning needs training-data access and a retrain, evasion needs only query access, and that you check the class's prerequisites against the rules of engagement before running anything.

for a middle

Adds that extraction and inference classes are technically runnable with queries but produce artefacts needing explicit permission, and that a blocked heading becomes a stated coverage gap.

for a senior

Frames the catalogue as attacker positions in the model lifecycle, and insists every finding names the access level it assumed so the client can act on it.

for a principal

Turns the pre-check into policy: a scoping template that maps granted access paths to permitted attack headings, with artefact handling and retention agreed up front.

### The catalogue is a list of attacker positions, not a difficulty ladder An adversarial-robustness toolkit — the Adversarial Robustness Toolbox (ART) is the canonical example — ships its attacks grouped under headings: evasion, poisoning, extraction, inference. A *heading* here is not a measure of how hard or how strong the attack is. It is a statement about **where in the model's lifecycle the attacker is standing** when the attack happens, and therefore about what the attacker must already be able to do. Read it that way and class selection stops being a matter of taste. | Heading | Attacker position | Hard precondition | What the run costs | What it emits | |---|---|---|---|---| | Evasion | Inference time, on the input | Submit an input, observe some output | Queries (black-box) or local compute (white-box) | Perturbed inputs + a success rate | | Poisoning / backdoor | Upstream, in the data or training path | A **write path** into training data *and* a retrain that actually occurs | The retrain — GPU hours and wall-clock the client controls, not you | A model whose behaviour changed | | Extraction | Query budget, rebuilding an approximation | A large query allowance and explicit permission | Often tens of thousands of metered queries; real money on a paid endpoint | **A surrogate model** — a derived copy of the client's asset | | Inference / membership | Query budget plus candidate records | Queries *and* records you are permitted to hold and test | Queries, plus reference data you may have to obtain | Claims about specific records, often personal data | ### Applying it to the stated scope Query access to a deployed model, training data and pipeline explicitly out of bounds: - **Evasion** fits. It is the heading whose precondition your grant actually satisfies. - **Poisoning and backdoor** are off the table. Both preconditions fail: no write path, and no retrain you can cause. A retrain is not something you can substitute for — it is the mechanism by which a poisoned row becomes a poisoned model. - **Extraction and inference** are *technically* runnable from queries, but their preconditions include a permission, not just an access. Both emit an artefact with legal weight — a copy of the model, or an assertion that a named person's record was in the training set. ### Where the number misleads The dangerous move is not refusing an out-of-scope heading; it is **substituting a lab**. A tester who cannot poison the client's pipeline trains a comparable model locally, poisons that, and reports a backdoor success rate. Every number in that run is real and every number is about the tester's own laptop. The client's deployed model differs in architecture details, training data, data-cleaning steps, canary checks, and whatever review sits between a contributed row and a shipped checkpoint — and it is exactly those differences the finding claims to have tested. The report says "backdoor achieved, 98%"; the honest sentence is "a model I built was backdoorable, and I do not know whether yours is." The second misleading reading is **silence**. A report that lists evasion results and says nothing about poisoning is read by a client as *tested and clean* across the board. Absence of a heading is absence of evidence, and only an explicit coverage statement makes that visible. ### What you check before picking any class For the class you intend to run, write down four things and the line of the rules of engagement that grants each: 1. **Access level required** — weights and gradients, full confidence scores, or bare labels. 2. **Data type accepted** — dense continuous tensors, tabular rows with constraints, or text. 3. **Artefact emitted** — perturbed inputs, a surrogate model, membership claims — and its retention and disposal rules. 4. **Cost shape and ceiling** — queries against a rate cap and a per-call price, or GPU hours against a wall-clock deadline. If any of the four has no grant behind it, either the run is out of scope or its result is unactionable: it describes an attacker position nobody has, and the client cannot reproduce it or remediate against it. When a heading is blocked, the correct output is a named coverage gap in the report — "poisoning and backdoor risk not assessed; no training-data write path was in scope" — not a lab-only run whose framing lets a reader assume the deployed system was tested.

  • The client has a user-feedback loop that is periodically used for fine-tuning. Does that change your answer about poisoning classes?
    Yes — that feedback loop is a write path into training data, so poisoning becomes in-scope in principle. But you still need the retrain to actually occur inside the engagement window, or agreement that a staged retrain on the client's own pipeline counts as the test.
  • If no poisoning run is possible, what goes in the report?
    An explicit coverage statement: poisoning and backdoor risk was not assessed, because no training-data write path was in scope. Silence reads as 'tested and clean'.
  • Why is 'the artefact it emits' part of the pre-check and not an afterthought?
    Because extraction and membership-inference runs produce a derived model or claims about specific records — things with retention, disclosure and legal consequences that must be agreed before you generate them, not after.

The headings are like the entry requirements on a job posting, not the salaries. Picking 'poisoning' because it sounds strongest is like applying for a role that requires a licence you do not hold: the ambition is irrelevant, the precondition decides.

saying these in an interview costs you the question

  • Treating the catalogue headings as interchangeable difficulty tiers and picking whichever class has the best-known name.
  • Running a poisoning class against a locally trained copy and reporting it as a finding against the deployed model.
  • Building a surrogate via an extraction class without checking that the rules of engagement permit creating and retaining it.
  • Not being able to say what a chosen class needs from the target before running it.

context