skip to content

Adversarial Robustness Tooling

You will learn how Counterfit and ART operationalize adversarial-ML attacks so you can assess a model without hand-coding each one. Interviewers probe them to see whether you can move from adversarial theory to an actual repeatable test harness.

on this pageshow

explore

questions

page 1 of 2

An adversarial-robustness toolkit such as the Adversarial Robustness Toolbox groups its attack classes under headings like evasion, poisoning, extraction and inference. Your engagement grants only query access to a deployed model and explicitly forbids touching its training data or pipeline. Which headings are off the table, and what must you check before picking any class?

level: juniorimportance: must knowfreq 55%

answer

  1. headings = preconditions, not difficulty
  2. poisoning needs a write path plus a retrain
  3. evasion needs inference-time input access
  4. extraction emits a surrogate: permission first
  5. access, data type, artefact, grant

basics

~20 s

Poisoning and backdoor classes are off the table: they need write access to training data or the training pipeline plus a retrain, which you were not granted. Evasion classes fit query access. Extraction and inference classes fit technically but produce a surrogate copy or membership claims, so confirm they are in scope and contractually allowed first.

solid answer

~50 s

The headings are not a difficulty ranking; they are a list of prerequisites. **Poisoning and backdoor** classes assume you can insert or modify training rows and that someone retrains — no training-data path, no run, and a synthetic retrain you did yourself proves nothing about the deployed model. **Evasion** classes assume inference-time input access, which is what you have. **Extraction** classes assume a large query budget and hand you a surrogate model — an artefact with legal and contractual weight, so it needs explicit permission and a disposal plan. **Inference/privacy** classes assume a query budget plus reference data you are allowed to hold. So before picking a class, check three things: the access it needs, the data type it accepts, and whether the artefact it produces is something your rules of engagement let you create and keep.

go deeper

for a junior

Should say that poisoning needs training-data access and a retrain, evasion needs only query access, and that you check the class's prerequisites against the rules of engagement before running anything.

for a middle

Adds that extraction and inference classes are technically runnable with queries but produce artefacts needing explicit permission, and that a blocked heading becomes a stated coverage gap.

for a senior

Frames the catalogue as attacker positions in the model lifecycle, and insists every finding names the access level it assumed so the client can act on it.

for a principal

Turns the pre-check into policy: a scoping template that maps granted access paths to permitted attack headings, with artefact handling and retention agreed up front.

### The catalogue is a list of attacker positions, not a difficulty ladder An adversarial-robustness toolkit — the Adversarial Robustness Toolbox (ART) is the canonical example — ships its attacks grouped under headings: evasion, poisoning, extraction, inference. A *heading* here is not a measure of how hard or how strong the attack is. It is a statement about **where in the model's lifecycle the attacker is standing** when the attack happens, and therefore about what the attacker must already be able to do. Read it that way and class selection stops being a matter of taste. | Heading | Attacker position | Hard precondition | What the run costs | What it emits | |---|---|---|---|---| | Evasion | Inference time, on the input | Submit an input, observe some output | Queries (black-box) or local compute (white-box) | Perturbed inputs + a success rate | | Poisoning / backdoor | Upstream, in the data or training path | A **write path** into training data *and* a retrain that actually occurs | The retrain — GPU hours and wall-clock the client controls, not you | A model whose behaviour changed | | Extraction | Query budget, rebuilding an approximation | A large query allowance and explicit permission | Often tens of thousands of metered queries; real money on a paid endpoint | **A surrogate model** — a derived copy of the client's asset | | Inference / membership | Query budget plus candidate records | Queries *and* records you are permitted to hold and test | Queries, plus reference data you may have to obtain | Claims about specific records, often personal data | ### Applying it to the stated scope Query access to a deployed model, training data and pipeline explicitly out of bounds: - **Evasion** fits. It is the heading whose precondition your grant actually satisfies. - **Poisoning and backdoor** are off the table. Both preconditions fail: no write path, and no retrain you can cause. A retrain is not something you can substitute for — it is the mechanism by which a poisoned row becomes a poisoned model. - **Extraction and inference** are *technically* runnable from queries, but their preconditions include a permission, not just an access. Both emit an artefact with legal weight — a copy of the model, or an assertion that a named person's record was in the training set. ### Where the number misleads The dangerous move is not refusing an out-of-scope heading; it is **substituting a lab**. A tester who cannot poison the client's pipeline trains a comparable model locally, poisons that, and reports a backdoor success rate. Every number in that run is real and every number is about the tester's own laptop. The client's deployed model differs in architecture details, training data, data-cleaning steps, canary checks, and whatever review sits between a contributed row and a shipped checkpoint — and it is exactly those differences the finding claims to have tested. The report says "backdoor achieved, 98%"; the honest sentence is "a model I built was backdoorable, and I do not know whether yours is." The second misleading reading is **silence**. A report that lists evasion results and says nothing about poisoning is read by a client as *tested and clean* across the board. Absence of a heading is absence of evidence, and only an explicit coverage statement makes that visible. ### What you check before picking any class For the class you intend to run, write down four things and the line of the rules of engagement that grants each: 1. **Access level required** — weights and gradients, full confidence scores, or bare labels. 2. **Data type accepted** — dense continuous tensors, tabular rows with constraints, or text. 3. **Artefact emitted** — perturbed inputs, a surrogate model, membership claims — and its retention and disposal rules. 4. **Cost shape and ceiling** — queries against a rate cap and a per-call price, or GPU hours against a wall-clock deadline. If any of the four has no grant behind it, either the run is out of scope or its result is unactionable: it describes an attacker position nobody has, and the client cannot reproduce it or remediate against it. When a heading is blocked, the correct output is a named coverage gap in the report — "poisoning and backdoor risk not assessed; no training-data write path was in scope" — not a lab-only run whose framing lets a reader assume the deployed system was tested.

  • The client has a user-feedback loop that is periodically used for fine-tuning. Does that change your answer about poisoning classes?
    Yes — that feedback loop is a write path into training data, so poisoning becomes in-scope in principle. But you still need the retrain to actually occur inside the engagement window, or agreement that a staged retrain on the client's own pipeline counts as the test.
  • If no poisoning run is possible, what goes in the report?
    An explicit coverage statement: poisoning and backdoor risk was not assessed, because no training-data write path was in scope. Silence reads as 'tested and clean'.
  • Why is 'the artefact it emits' part of the pre-check and not an afterthought?
    Because extraction and membership-inference runs produce a derived model or claims about specific records — things with retention, disclosure and legal consequences that must be agreed before you generate them, not after.

The headings are like the entry requirements on a job posting, not the salaries. Picking 'poisoning' because it sounds strongest is like applying for a role that requires a licence you do not hold: the ambition is irrelevant, the precondition decides.

saying these in an interview costs you the question

  • Treating the catalogue headings as interchangeable difficulty tiers and picking whichever class has the best-known name.
  • Running a poisoning class against a locally trained copy and reporting it as a finding against the deployed model.
  • Building a surrogate via an extraction class without checking that the rules of engagement permit creating and retaining it.
  • Not being able to say what a chosen class needs from the target before running it.

context

open as a page

In an adversarial-robustness toolkit such as the Adversarial Robustness Toolbox, defences ship as objects of several kinds: preprocessor, postprocessor, detector, trainer and transformer. Where does each kind sit relative to a call to the wrapped model, and which of them change the model weights?

level: juniorimportance: must knowfreq 55%

basics

~20 s

A preprocessor defence transforms inputs before the model sees them. A postprocessor alters what comes back, usually the scores. A detector flags a sample as adversarial instead of classifying it. A trainer defence retrains the weights. A transformer hands you back a new, modified model object. Only trainer and transformer touch weights.

open as a page

An adversarial-attack library reports two numbers after an evaluation run: accuracy on the adversarial examples, and a mean perturbation size. Which rows does each of those two numbers average over?

level: juniorimportance: must knowfreq 55%

basics

~20 s

Adversarial accuracy is over every example you evaluated, flipped or not. The mean perturbation size is over only the examples the attack actually flipped; rows it failed on are dropped, not counted as huge perturbations. Two different denominators, so the two numbers are not about the same population.

open as a page

In an adversarial-ML evaluation stack, what is the difference between an attack library (such as the Adversarial Robustness Toolbox or Foolbox) and a command-line harness that drives one (such as Counterfit)? Which of the two decides what attacks are available to you and how far you can tune them?

level: juniorimportance: must knowfreq 62%

basics

~20 s

The attack library implements the attacks and holds the parameters; the harness only drives it, doing target setup, batching, logging and output. So the library decides which attacks exist and how tunable they are. The harness decides how conveniently you run them, and may expose fewer attacks or fewer parameters.

open as a page

In Counterfit, an attack reaches a model only through a scan target you write. What must that scan target supply so a query-only attack can run against a hosted inference endpoint, and what does it deliberately not hand the attack?

level: juniorimportance: must knowfreq 62%

basics

~20 s

The scan target wraps your endpoint as a callable: given a batch of samples it returns the model's per-class scores, or a label. You also declare the input shape and data type, the list of output classes, and a few seed samples the attack will perturb. It hands over queries only, never gradients or weights.

open as a page

You are about to launch a query-based (black-box) attack from an adversarial robustness toolkit against a hosted model endpoint that bills per call. How do you estimate what the run will cost in calls before you start it?

level: juniorimportance: must knowfreq 55%

basics

~20 s

Multiply the number of evaluation examples by the attack's per-example query cap, times restarts, and add one clean pass to find which examples are classified correctly. Attacks spend the whole cap on examples they fail. Set the cap explicitly, pilot on ten examples, measure the real call count, then scale.

open as a page

You build an evasion attack from an adversarial-robustness library such as Foolbox, torchattacks or the Adversarial Robustness Toolbox and pass only the wrapped model — no iteration count, no step size, no number of random restarts. Where do those values come from, and what were they chosen for?

level: juniorimportance: must knowfreq 60%

basics

~20 s

They come from the attack class's own constructor defaults, baked in so the library's documentation example runs fast on a laptop. They are demo settings, not assessment settings. A model that resists them has resisted only a short, weak search. Choose and record every strength argument yourself before reporting a robustness number.

open as a page

You run a data-poisoning attack from an adversarial-robustness library and it hands back arrays of training samples and labels. Why is that not yet a result, and what has to happen before you can say whether the model is vulnerable?

level: juniorimportance: must knowfreq 50%

basics

~20 s

The library only crafts tainted training data; it does not train anything. You have to mix those rows into the training set, retrain the model with your normal recipe, then evaluate it twice: normal accuracy on clean test data and the attack's success on the triggered inputs. The verdict costs a retrain, not one call.

open as a page

When you wrap a model for an adversarial-example library such as the Adversarial Robustness Toolbox, you declare a permitted minimum and maximum for the input values. What does that declaration change about the examples the attack generates, and what does it not?

level: juniorimportance: must knowfreq 70%

basics

~20 s

When you wrap a model for an adversarial-example library you declare the legal minimum and maximum for input values. The attack clips every generated example back inside that box, so pixels stay in range. It is one global range over all features, and it knows nothing about what any individual feature means.

open as a page

In an adversarial-robustness library such as the Adversarial Robustness Toolbox or Foolbox, attack classes differ in what they demand from the model object you hand them. What access levels can an attack class require, and how do you establish which level a given class needs before you run it?

level: middleimportance: must knowfreq 65%

basics

~20 s

Three levels. White-box classes need loss gradients through the model. Score-based black-box classes need the full confidence vector per query. Decision-based classes need only the predicted label. You establish which from the class's documented requirement and the arguments its constructor insists on; these libraries generally reject a model that cannot supply what the class needs.

open as a page

You attach two preprocessing defence objects to a wrapped classifier in an adversarial-robustness toolkit such as the Adversarial Robustness Toolbox, then run a gradient-based attack from the same library. What decides the order they run in and whether the attack optimises through them, and why does it matter that a preprocessing defence can be configured to run only at training time?

level: middleimportance: must knowfreq 50%

basics

~20 s

They form an ordered chain applied in the order you attach them, so the second sees the first one's output. Each declares whether it runs at training, at inference, or both. A defence that runs only at training is absent when the attack queries the model, so the run measures a pipeline nobody serves.

open as a page

A model wrapper for an adversarial robustness toolkit can expose predictions only, or predictions plus the gradient of the loss with respect to the input. What must you implement for each surface, and what happens if you return a stub from the gradient one?

level: middleimportance: must knowfreq 65%

basics

~20 s

A predict-only wrapper maps a batch of inputs to model outputs. A gradient-exposing wrapper must additionally return the derivative of the loss with respect to the input, at the same input the model sees. Stub that and gradient attacks still run, but they follow a fake signal and report robustness that only measures your stub.

open as a page

An adversarial library's metric helper — the one that averages perturbation size over the examples an attack flipped — returns 0.0 after a run in which the attack flipped nothing. Why is that zero not evidence the model is trivially fragile?

level: middleimportance: must knowfreq 58%

basics

~20 s

Because it is a mean over an empty set. No example was flipped, so nothing entered the average and the helper returns zero as a degenerate default. On this scale small means fragile, so the failure case prints at the alarming end. Check the attack success count before reading the number at all.

open as a page

You can run an evasion evaluation either through a command-line wrapper over adversarial attack libraries (such as Counterfit) or by importing the attack library (such as the Adversarial Robustness Toolbox or Foolbox) and writing your own driver. What does the wrapper layer buy you, and what does it charge you?

level: middleimportance: must knowfreq 58%

basics

~20 s

The wrapper gives you one uniform way to point attacks at a model plus a repeatable loop: consistent target definition, batch runs, logging and result output you did not write. You pay with its dependency pins, its supported platform list, and a parameter surface narrower than the library's own attack classes.

open as a page

You are about to fire a Counterfit scan that runs several named attacks one after another against an endpoint billed per call. How do you estimate the number of model queries the whole scan will make before you start it?

level: middleimportance: must knowfreq 55%

basics

~20 s

Count per attack, then add up. Queries are roughly seed samples times iterations times calls per iteration, and a black-box attack spends many calls per step estimating a direction or probing a boundary. Do not trust the arithmetic alone: put a counter in the target's callable, run one attack on one sample, and extrapolate.

open as a page

In an adversarial robustness library such as the Adversarial Robustness Toolbox or Foolbox, every attack runs against a wrapper object you supply for the model under test. If that wrapper can only send an input to a remote inference API and return the class probabilities it gets back, which attack families can you still run, and why do the rest fail?

level: middleimportance: must knowfreq 68%

basics

~20 s

Attacks call methods on your wrapper. Gradient attacks need a gradient method, which a forward-only HTTP wrapper cannot implement, so they error or are unavailable. You are left with query attacks: score-based ones that read the returned probabilities, decision-based ones needing only the top label, and transfer from a local surrogate.

open as a page

For an iterative gradient evasion attack driven from an adversarial-robustness library such as the Adversarial Robustness Toolbox or torchattacks, what does raising each of these three arguments buy you — the iteration count, the step size relative to the perturbation bound, and the number of random restarts — and what does each cost?

level: middleimportance: must knowfreq 55%

basics

~20 s

More iterations let the search refine longer inside the same bound. A step size too large for the bound overshoots and one too small never reaches its edge. More restarts re-launch from fresh random starting points, so one unlucky start is not read as robustness. Each knob multiplies GPU time roughly linearly.

open as a page

You retrained a classifier on data produced by a backdoor-poisoning routine in an adversarial-robustness library. Which two numbers do you report, and why is the trigger success rate on its own a misleading result?

level: middleimportance: must knowfreq 50%

basics

~20 s

Report two: accuracy on a clean, untriggered test set against the unpoisoned baseline, and the trigger success rate measured only on samples not already in the target class. Success rate alone hides a poisoned model that lost obvious accuracy, and it is inflated by inputs the clean model already sent to the target.

open as a page

You are attacking a tabular fraud model with an adversarial-example library, and 12 of its 40 features are ones the attacker cannot influence, such as account tenure. The library lets you pass a mask of which features may move. What does that mask guarantee, and what still has to be enforced outside the library?

level: middleimportance: must knowfreq 60%

basics

~20 s

The mask freezes coordinates: features you mark immovable keep their original values, so the search only touches the ones you allow. It says nothing about the movable features' own rules, such as integer counts, one-hot exclusivity, or a field derived from another. Those you check yourself after the library returns its examples.

open as a page

After you attach a defence object to a wrapped classifier in an adversarial-robustness toolkit, the robust accuracy under your configured attack rises from near zero to about three quarters. What do you check before reporting that as a robustness gain?

level: seniorimportance: must knowfreq 55%

basics

~20 s

Assume the attack broke rather than the model became robust. Re-tune the attack against the defended pipeline, raise iterations and restarts, and check that an unbounded perturbation still drives accuracy to zero. Compare against a query-based attack and a transfer attack. Confirm the defence really sits inside the wrapper you attacked.

open as a page

An attack run against your wrapped image classifier flips almost every test image, but the same adversarial files fail when uploaded to the live service, which resizes and re-encodes each upload before inference. What is wrong with the wrapper, and how do you fix it?

level: seniorimportance: must knowfreq 55%

basics

~20 s

The wrapper exposes the bare model, so the attack optimised a tensor the live service never receives. The resize and re-encode step sits outside the wrapper and destroys the perturbation. Fix it by moving the whole serving preprocessing chain inside the wrapper, so the attack's input is the uploaded file rather than the tensor.

open as a page

A robustness write-up says an evasion attack "failed" against a model, but the run was driven through a command-line wrapper that exposes only some of the underlying attack library's parameters. Why is "the attack failed" not yet a supportable claim, and what do you do before it goes in a report?

level: seniorimportance: must knowfreq 48%

basics

~20 s

Because the run only tested the attack at the settings the wrapper exposed. An attack with an untuned step size, iteration count or perturbation budget failing says nothing about robustness. Before publishing, check which parameters the layer passes through, sweep the ones that matter, and record the exact settings tested.

open as a page

An evasion run driven from an adversarial-robustness library finishes and reports zero successful adversarial examples at the configured perturbation bound. Before you write 'the model is robust', what checks do you run on the attack configuration and the model wrapper itself?

level: seniorimportance: must knowfreq 50%

basics

~20 s

Assume the harness is broken first. Re-run with the bound relaxed until the attack must succeed; if it still fails, the wrapper or the loop is wrong. Then sweep iterations and restarts upward and watch the success rate. Also check a cheap non-gradient baseline: if it beats the gradient attack, the gradients are unreliable.

open as a page

When you wrap a trained image classifier so an adversarial robustness library can attack it, the wrapper makes you declare the valid input value range. What does the library do with that range, and what goes wrong if you declare it wrong?

level: juniorimportance: should knowfreq 45%

basics

~20 s

It tells the library the legal bounds of an input, so every perturbed candidate is clipped back into them. Declare 0-255 when your model is actually fed 0-1 and the attack searches a space your service never accepts, so the perturbation size you report means nothing.

open as a page

An adversarial-robustness toolkit's attack classes are written for a particular input data type. What goes wrong when you point an image-oriented gradient attack class at a tabular fraud model whose features include one-hot encoded categoricals, integer counts and an account-age field the customer cannot change?

level: middleimportance: should knowfreq 45%

basics

~20 s

It treats every column as a continuous pixel and nudges all of them. You get rows with fractional counts, several one-hot columns partly set so no single category is encoded, values outside legal ranges, and a changed account age. The label flips, but no attacker could submit that row, so the success rate measures nothing actionable.

open as a page

You attach a detector defence object to a wrapped classifier in an adversarial-robustness toolkit such as the Adversarial Robustness Toolbox. It returns an adversarial-or-benign decision per sample rather than a class label. What has to change in how you score the pipeline, and what number must you report next to its detection rate?

level: middleimportance: should knowfreq 45%

basics

~20 s

The pipeline now has three outcomes, not two: correct, wrong, and rejected. Decide up front how a rejected sample counts, for adversarial and for clean inputs separately. Report the false-positive rate on clean data beside the detection rate, because the threshold is tunable and a detection rate alone hides what it costs benign traffic.

open as a page

An adversarial-attack command-line wrapper installs with its own pinned versions of the ML framework and the attack libraries it wraps. What problems does that create when you add it to an existing evaluation environment, and how do you contain them?

level: middleimportance: should knowfreq 40%

basics

~20 s

It drags a whole transitive stack into your environment: its pinned framework and attack-library versions can conflict with what your training or serving code needs. Contain it by giving the wrapper its own isolated environment or container and moving data across as files, rather than importing it into an existing project.

open as a page

Rather than choosing iteration counts and restarts yourself, a teammate proposes measuring robustness with a standardized evaluation suite such as RobustBench, which pins the attacks and their settings for you. What does that fix about the defaults problem, and what does it still not tell you about your own deployed model?

level: middleimportance: should knowfreq 35%

basics

~20 s

It fixes comparability: everyone runs the same fixed attack set at the same effort and threat model, so numbers can be ranked and nobody quietly under-configures. It does not fix relevance. It measures one threat model on one dataset through a required interface, not your inputs, your preprocessing pipeline or the attacks your product actually faces.

open as a page

In an adversarial-robustness library, a model-extraction (stealing) attack will not run until you have supplied a substitute model of your own. What exactly does the library do with it, and which parts of the run stay your responsibility?

level: middleimportance: should knowfreq 42%

basics

~20 s

You supply two things: a wrapper around the victim that answers queries, and an untrained substitute model you chose and configured. The library queries the victim over inputs you provide, labels them, and fits your substitute. It returns your object, now trained. Architecture, query pool and the query bill are all yours.

open as a page

In TextAttack, an attack recipe pairs a transformation that proposes candidate rewrites with a set of constraints those candidates must pass. What role do the constraints play in the search, and what happens to your reported success rate if you relax them?

level: middleimportance: should knowfreq 45%

basics

~20 s

Constraints are filters inside the loop: a transformation proposes candidate sentences, each constraint rejects the ones that violate it, and the search only ever sees the survivors. They are what keeps a rewrite readable and meaning-preserving. Loosen them and the success rate rises, because you are now counting rewrites that changed the sentence.

open as a page

showing 1–30 of 49