skip to content

Toolkit Abstractions

An attack library reaches your model only through a wrapper, and its class list encodes access assumptions, not exploits. Interviewers ask because reading them as features hides what you cannot run.

on this pageshow

explore

questions

20

An adversarial-robustness toolkit such as the Adversarial Robustness Toolbox groups its attack classes under headings like evasion, poisoning, extraction and inference. Your engagement grants only query access to a deployed model and explicitly forbids touching its training data or pipeline. Which headings are off the table, and what must you check before picking any class?

level: juniorimportance: must knowfreq 55%

answer

  1. headings = preconditions, not difficulty
  2. poisoning needs a write path plus a retrain
  3. evasion needs inference-time input access
  4. extraction emits a surrogate: permission first
  5. access, data type, artefact, grant

basics

~20 s

Poisoning and backdoor classes are off the table: they need write access to training data or the training pipeline plus a retrain, which you were not granted. Evasion classes fit query access. Extraction and inference classes fit technically but produce a surrogate copy or membership claims, so confirm they are in scope and contractually allowed first.

solid answer

~50 s

The headings are not a difficulty ranking; they are a list of prerequisites. **Poisoning and backdoor** classes assume you can insert or modify training rows and that someone retrains — no training-data path, no run, and a synthetic retrain you did yourself proves nothing about the deployed model. **Evasion** classes assume inference-time input access, which is what you have. **Extraction** classes assume a large query budget and hand you a surrogate model — an artefact with legal and contractual weight, so it needs explicit permission and a disposal plan. **Inference/privacy** classes assume a query budget plus reference data you are allowed to hold. So before picking a class, check three things: the access it needs, the data type it accepts, and whether the artefact it produces is something your rules of engagement let you create and keep.

go deeper

for a junior

Should say that poisoning needs training-data access and a retrain, evasion needs only query access, and that you check the class's prerequisites against the rules of engagement before running anything.

for a middle

Adds that extraction and inference classes are technically runnable with queries but produce artefacts needing explicit permission, and that a blocked heading becomes a stated coverage gap.

for a senior

Frames the catalogue as attacker positions in the model lifecycle, and insists every finding names the access level it assumed so the client can act on it.

for a principal

Turns the pre-check into policy: a scoping template that maps granted access paths to permitted attack headings, with artefact handling and retention agreed up front.

### The catalogue is a list of attacker positions, not a difficulty ladder An adversarial-robustness toolkit — the Adversarial Robustness Toolbox (ART) is the canonical example — ships its attacks grouped under headings: evasion, poisoning, extraction, inference. A *heading* here is not a measure of how hard or how strong the attack is. It is a statement about **where in the model's lifecycle the attacker is standing** when the attack happens, and therefore about what the attacker must already be able to do. Read it that way and class selection stops being a matter of taste. | Heading | Attacker position | Hard precondition | What the run costs | What it emits | |---|---|---|---|---| | Evasion | Inference time, on the input | Submit an input, observe some output | Queries (black-box) or local compute (white-box) | Perturbed inputs + a success rate | | Poisoning / backdoor | Upstream, in the data or training path | A **write path** into training data *and* a retrain that actually occurs | The retrain — GPU hours and wall-clock the client controls, not you | A model whose behaviour changed | | Extraction | Query budget, rebuilding an approximation | A large query allowance and explicit permission | Often tens of thousands of metered queries; real money on a paid endpoint | **A surrogate model** — a derived copy of the client's asset | | Inference / membership | Query budget plus candidate records | Queries *and* records you are permitted to hold and test | Queries, plus reference data you may have to obtain | Claims about specific records, often personal data | ### Applying it to the stated scope Query access to a deployed model, training data and pipeline explicitly out of bounds: - **Evasion** fits. It is the heading whose precondition your grant actually satisfies. - **Poisoning and backdoor** are off the table. Both preconditions fail: no write path, and no retrain you can cause. A retrain is not something you can substitute for — it is the mechanism by which a poisoned row becomes a poisoned model. - **Extraction and inference** are *technically* runnable from queries, but their preconditions include a permission, not just an access. Both emit an artefact with legal weight — a copy of the model, or an assertion that a named person's record was in the training set. ### Where the number misleads The dangerous move is not refusing an out-of-scope heading; it is **substituting a lab**. A tester who cannot poison the client's pipeline trains a comparable model locally, poisons that, and reports a backdoor success rate. Every number in that run is real and every number is about the tester's own laptop. The client's deployed model differs in architecture details, training data, data-cleaning steps, canary checks, and whatever review sits between a contributed row and a shipped checkpoint — and it is exactly those differences the finding claims to have tested. The report says "backdoor achieved, 98%"; the honest sentence is "a model I built was backdoorable, and I do not know whether yours is." The second misleading reading is **silence**. A report that lists evasion results and says nothing about poisoning is read by a client as *tested and clean* across the board. Absence of a heading is absence of evidence, and only an explicit coverage statement makes that visible. ### What you check before picking any class For the class you intend to run, write down four things and the line of the rules of engagement that grants each: 1. **Access level required** — weights and gradients, full confidence scores, or bare labels. 2. **Data type accepted** — dense continuous tensors, tabular rows with constraints, or text. 3. **Artefact emitted** — perturbed inputs, a surrogate model, membership claims — and its retention and disposal rules. 4. **Cost shape and ceiling** — queries against a rate cap and a per-call price, or GPU hours against a wall-clock deadline. If any of the four has no grant behind it, either the run is out of scope or its result is unactionable: it describes an attacker position nobody has, and the client cannot reproduce it or remediate against it. When a heading is blocked, the correct output is a named coverage gap in the report — "poisoning and backdoor risk not assessed; no training-data write path was in scope" — not a lab-only run whose framing lets a reader assume the deployed system was tested.

  • The client has a user-feedback loop that is periodically used for fine-tuning. Does that change your answer about poisoning classes?
    Yes — that feedback loop is a write path into training data, so poisoning becomes in-scope in principle. But you still need the retrain to actually occur inside the engagement window, or agreement that a staged retrain on the client's own pipeline counts as the test.
  • If no poisoning run is possible, what goes in the report?
    An explicit coverage statement: poisoning and backdoor risk was not assessed, because no training-data write path was in scope. Silence reads as 'tested and clean'.
  • Why is 'the artefact it emits' part of the pre-check and not an afterthought?
    Because extraction and membership-inference runs produce a derived model or claims about specific records — things with retention, disclosure and legal consequences that must be agreed before you generate them, not after.

The headings are like the entry requirements on a job posting, not the salaries. Picking 'poisoning' because it sounds strongest is like applying for a role that requires a licence you do not hold: the ambition is irrelevant, the precondition decides.

saying these in an interview costs you the question

  • Treating the catalogue headings as interchangeable difficulty tiers and picking whichever class has the best-known name.
  • Running a poisoning class against a locally trained copy and reporting it as a finding against the deployed model.
  • Building a surrogate via an extraction class without checking that the rules of engagement permit creating and retaining it.
  • Not being able to say what a chosen class needs from the target before running it.

context

open as a page

In an adversarial-robustness toolkit such as the Adversarial Robustness Toolbox, defences ship as objects of several kinds: preprocessor, postprocessor, detector, trainer and transformer. Where does each kind sit relative to a call to the wrapped model, and which of them change the model weights?

level: juniorimportance: must knowfreq 55%

basics

~20 s

A preprocessor defence transforms inputs before the model sees them. A postprocessor alters what comes back, usually the scores. A detector flags a sample as adversarial instead of classifying it. A trainer defence retrains the weights. A transformer hands you back a new, modified model object. Only trainer and transformer touch weights.

open as a page

An adversarial-attack library reports two numbers after an evaluation run: accuracy on the adversarial examples, and a mean perturbation size. Which rows does each of those two numbers average over?

level: juniorimportance: must knowfreq 55%

basics

~20 s

Adversarial accuracy is over every example you evaluated, flipped or not. The mean perturbation size is over only the examples the attack actually flipped; rows it failed on are dropped, not counted as huge perturbations. Two different denominators, so the two numbers are not about the same population.

open as a page

In an adversarial-robustness library such as the Adversarial Robustness Toolbox or Foolbox, attack classes differ in what they demand from the model object you hand them. What access levels can an attack class require, and how do you establish which level a given class needs before you run it?

level: middleimportance: must knowfreq 65%

basics

~20 s

Three levels. White-box classes need loss gradients through the model. Score-based black-box classes need the full confidence vector per query. Decision-based classes need only the predicted label. You establish which from the class's documented requirement and the arguments its constructor insists on; these libraries generally reject a model that cannot supply what the class needs.

open as a page

You attach two preprocessing defence objects to a wrapped classifier in an adversarial-robustness toolkit such as the Adversarial Robustness Toolbox, then run a gradient-based attack from the same library. What decides the order they run in and whether the attack optimises through them, and why does it matter that a preprocessing defence can be configured to run only at training time?

level: middleimportance: must knowfreq 50%

basics

~20 s

They form an ordered chain applied in the order you attach them, so the second sees the first one's output. Each declares whether it runs at training, at inference, or both. A defence that runs only at training is absent when the attack queries the model, so the run measures a pipeline nobody serves.

open as a page

A model wrapper for an adversarial robustness toolkit can expose predictions only, or predictions plus the gradient of the loss with respect to the input. What must you implement for each surface, and what happens if you return a stub from the gradient one?

level: middleimportance: must knowfreq 65%

basics

~20 s

A predict-only wrapper maps a batch of inputs to model outputs. A gradient-exposing wrapper must additionally return the derivative of the loss with respect to the input, at the same input the model sees. Stub that and gradient attacks still run, but they follow a fake signal and report robustness that only measures your stub.

open as a page

An adversarial library's metric helper — the one that averages perturbation size over the examples an attack flipped — returns 0.0 after a run in which the attack flipped nothing. Why is that zero not evidence the model is trivially fragile?

level: middleimportance: must knowfreq 58%

basics

~20 s

Because it is a mean over an empty set. No example was flipped, so nothing entered the average and the helper returns zero as a degenerate default. On this scale small means fragile, so the failure case prints at the alarming end. Check the attack success count before reading the number at all.

open as a page

After you attach a defence object to a wrapped classifier in an adversarial-robustness toolkit, the robust accuracy under your configured attack rises from near zero to about three quarters. What do you check before reporting that as a robustness gain?

level: seniorimportance: must knowfreq 55%

basics

~20 s

Assume the attack broke rather than the model became robust. Re-tune the attack against the defended pipeline, raise iterations and restarts, and check that an unbounded perturbation still drives accuracy to zero. Compare against a query-based attack and a transfer attack. Confirm the defence really sits inside the wrapper you attacked.

open as a page

An attack run against your wrapped image classifier flips almost every test image, but the same adversarial files fail when uploaded to the live service, which resizes and re-encodes each upload before inference. What is wrong with the wrapper, and how do you fix it?

level: seniorimportance: must knowfreq 55%

basics

~20 s

The wrapper exposes the bare model, so the attack optimised a tensor the live service never receives. The resize and re-encode step sits outside the wrapper and destroys the perturbation. Fix it by moving the whole serving preprocessing chain inside the wrapper, so the attack's input is the uploaded file rather than the tensor.

open as a page

When you wrap a trained image classifier so an adversarial robustness library can attack it, the wrapper makes you declare the valid input value range. What does the library do with that range, and what goes wrong if you declare it wrong?

level: juniorimportance: should knowfreq 45%

basics

~20 s

It tells the library the legal bounds of an input, so every perturbed candidate is clipped back into them. Declare 0-255 when your model is actually fed 0-1 and the attack searches a space your service never accepts, so the perturbation size you report means nothing.

open as a page

An adversarial-robustness toolkit's attack classes are written for a particular input data type. What goes wrong when you point an image-oriented gradient attack class at a tabular fraud model whose features include one-hot encoded categoricals, integer counts and an account-age field the customer cannot change?

level: middleimportance: should knowfreq 45%

basics

~20 s

It treats every column as a continuous pixel and nudges all of them. You get rows with fractional counts, several one-hot columns partly set so no single category is encoded, values outside legal ranges, and a changed account age. The label flips, but no attacker could submit that row, so the success rate measures nothing actionable.

open as a page

You attach a detector defence object to a wrapped classifier in an adversarial-robustness toolkit such as the Adversarial Robustness Toolbox. It returns an adversarial-or-benign decision per sample rather than a class label. What has to change in how you score the pipeline, and what number must you report next to its detection rate?

level: middleimportance: should knowfreq 45%

basics

~20 s

The pipeline now has three outcomes, not two: correct, wrong, and rejected. Decide up front how a rejected sample counts, for adversarial and for clean inputs separately. Report the false-positive rate on clean data beside the detection rate, because the threshold is tunable and a detection rate alone hides what it costs benign traffic.

open as a page

Rules of engagement give you a metered HTTPS endpoint that returns one predicted label per query, with a per-day request cap, and no model artefact. A teammate reports a very high success rate produced by a gradient-based attack class in an adversarial-robustness toolkit, run against a copy of the model downloaded from a public hub. How do you handle that number, and which attack class do you actually run against the granted surface?

level: seniorimportance: should knowfreq 48%

basics

~20 s

Do not report it against the deployed endpoint: it measured a different artefact under access nobody granted. Relabel it as a lab run on a public copy. For the granted surface, pick a decision-based class needing only the top-1 label, pilot a few examples to measure queries per success, and compare that to the daily cap.

open as a page

Before you publish a robustness number from a wrapped model, how do you verify that the gradient surface your wrapper exposes really is the gradient of the model's loss with respect to its input?

level: seniorimportance: should knowfreq 40%

basics

~20 s

Do not assume it — test it. Compare the returned gradient against a central finite-difference estimate on a handful of coordinates and inputs, using only the prediction surface for the loss. Also check it is not all zeros or NaN, and that a small step along it moves the loss the way the sign says.

open as a page

Before quoting a mean-perturbation robustness number from an adversarial library, what do you check about the rows that fed it — in particular, examples the model already misclassified before any attack ran?

level: seniorimportance: should knowfreq 38%

basics

~20 s

Check three things: how many rows were evaluated, how many the attack flipped, and whether rows the model got wrong at baseline were excluded. A baseline-wrong row can register as an instant success with near-zero perturbation, dragging the average down and inflating the success rate — a measurement of plain accuracy, not robustness.

open as a page

Two image classifiers are scored with the same library metric — the mean perturbation size over the examples an attack flipped — using the same attack, the same norm and the same settings. Model A's number is larger than model B's. Why does that not settle which model is more robust?

level: seniorimportance: should knowfreq 44%

basics

~20 s

Because each number averages over a different set of rows: whichever examples that model happened to lose. If the attack flipped only a handful of A's examples, A's average covers a narrow subset, while B's may cover most of its test set including hard cases. Same units, different populations, so the comparison is not like-for-like.

open as a page

Your team has standardised on one adversarial-robustness toolkit and reuses the same three attack classes on every model assessment. What is the case for and against that as a program-level policy, and how would you decide the class set per engagement instead?

level: principalimportance: should knowfreq 30%

basics

~20 s

For: comparability across assessments, reviewable tooling, faster onboarding, predictable cost. Against: the three classes encode one fixed access level and data type, so on a differently shaped target a clean result means only that those three did not apply. Decide per engagement from granted access, input domain and attacker action space, keeping a small fixed core purely for trend comparison.

open as a page

A client asks you to shortlist defences for an image classifier using an adversarial-robustness toolkit, and you have one GPU for two weeks. The candidates split into defence objects that attach to the existing weights and trainer-style defences that need a full retrain per hyperparameter setting. How do you allocate the compute, and what do you tell the client the shortlist does and does not cover?

level: principalimportance: should knowfreq 35%

basics

~20 s

Split by cost. Attach-only defences reuse the trained weights, so each costs one evaluation sweep and can face many attack settings. Every trainer variant costs a full retrain, so you can afford only a couple of settings. Spend most of the compute on deep adaptive evaluation of a short list, not a broad shallow sweep.

open as a page

For a client robustness engagement you must decide how much of the production serving stack the model wrapper reproduces: the bare model, the model plus preprocessing, or the full path including a rule-based blocklist and a three-model ensemble vote. How do you decide, and what may the report claim in each case?

level: principalimportance: should knowfreq 30%

basics

~20 s

Decide by what the client will act on. A bare-model wrapper measures a component; a wrapper reproducing preprocessing, rules and the ensemble vote measures the service. Pick the widest layer you can reproduce faithfully and cheaply, then state in the report exactly which layers the number covers and which were excluded.

open as a page

Your organisation wants to gate model releases on one robustness number from an adversarial library — a mean perturbation size over only the examples an attack flipped. What breaks when that number is tracked release over release, and what would you gate on instead?

level: principalimportance: nice to knowfreq 26%

basics

~20 s

It is a conditional mean whose population changes every release, so it is not monotone in robustness and moves for reasons unrelated to the model. It also collapses to zero when nothing is flipped. Gate instead on attack success rate at a pinned perturbation budget over a frozen example set, with row counts recorded.

open as a page