Adversarial Robustness Tooling
You will learn how Counterfit and ART operationalize adversarial-ML attacks so you can assess a model without hand-coding each one. Interviewers probe them to see whether you can move from adversarial theory to an actual repeatable test harness.
on this pageshowhide
explore
- Toolkit Abstractions20 questions
- Estimator Wrappers5 questions
- Attack Classes5 questions
- Defence Modules5 questions
- Library-Reported Scores5 questions
- Running an Assessment20 questions
- Matching Attack to Access5 questions
- Domain Constraints5 questions
- Default Attack Strength5 questions
- Beyond Evasion5 questions
- Counterfit9 questions
- Scan Targets and Workflow4 questions
- Choosing a Harness Today5 questions
questions
page 2 of 2Rules of engagement give you a metered HTTPS endpoint that returns one predicted label per query, with a per-day request cap, and no model artefact. A teammate reports a very high success rate produced by a gradient-based attack class in an adversarial-robustness toolkit, run against a copy of the model downloaded from a public hub. How do you handle that number, and which attack class do you actually run against the granted surface?
basics
~20 sDo not report it against the deployed endpoint: it measured a different artefact under access nobody granted. Relabel it as a lab run on a public copy. For the granted surface, pick a decision-based class needing only the top-1 label, pilot a few examples to measure queries per success, and compare that to the daily cap.
Before you publish a robustness number from a wrapped model, how do you verify that the gradient surface your wrapper exposes really is the gradient of the model's loss with respect to its input?
basics
~20 sDo not assume it — test it. Compare the returned gradient against a central finite-difference estimate on a handful of coordinates and inputs, using only the prediction surface for the loss. Also check it is not all zeros or NaN, and that a small step along it moves the loss the way the sign says.
Before quoting a mean-perturbation robustness number from an adversarial library, what do you check about the rows that fed it — in particular, examples the model already misclassified before any attack ran?
basics
~20 sCheck three things: how many rows were evaluated, how many the attack flipped, and whether rows the model got wrong at baseline were excluded. A baseline-wrong row can register as an instant success with near-zero perturbation, dragging the average down and inflating the success rate — a measurement of plain accuracy, not robustness.
Two image classifiers are scored with the same library metric — the mean perturbation size over the examples an attack flipped — using the same attack, the same norm and the same settings. Model A's number is larger than model B's. Why does that not settle which model is more robust?
basics
~20 sBecause each number averages over a different set of rows: whichever examples that model happened to lose. If the attack flipped only a handful of A's examples, A's average covers a narrow subset, while B's may cover most of its test set including hard cases. Same units, different populations, so the comparison is not like-for-like.
Your Counterfit scan target points at a live production inference endpoint instead of a locally loaded model. What does that change about what the scan result is actually measuring, and about your ability to repeat the run?
basics
~20 sYou stop testing a model and start testing a deployment: request validation, preprocessing, the served copy of the weights, any filter in front, and caching. The result describes that deployment at that hour. A failed attack can mean an error response or a version change rather than robustness, so the run is not cleanly repeatable.
An overnight Counterfit scan that had run fine all week burned a month of endpoint quota after a colleague switched on the optional search over attack parameters. Where does that multiplier come from, and how would you bound the next run?
basics
~20 sThe parameter search does not tune inside one attack run; it reruns the entire attack once per trial. Cost becomes attacks times trials times the full per-attack query count, and that per-attack count is already seeds times iterations. Ten trials over three attacks is thirty complete attacks against a metered endpoint.
In a decision-based (label-only) attack run from an adversarial robustness toolkit against a remote endpoint, 60% of examples hit the per-example query cap without producing an adversarial example. What do you check in the run before you treat that as a property of the model, and what do you change for the next run?
basics
~20 sTreat it as censored data, not a model property. Check the queries-to-success distribution on the examples that worked, whether the perturbation was still shrinking when the cap hit, and whether errors, throttling or duplicate queries burned the allowance. Then re-run a stratified subsample at a much larger cap.
A query-based attack loop from an adversarial robustness toolkit assumes each call to the target answers quickly and consistently for the same input. You are pointing it at a rate-limited, non-deterministic hosted endpoint that bills per call. Of retries, throttling and answer variability, which spend your call budget, which spend only wall-clock, and what do you do about each?
basics
~20 sRetries and repeated probes spend billed calls. Throttling and latency spend wall-clock, which limits how many examples you finish per day. Variability is worst: a boundary answer that flips makes the search oscillate. Handle it by voting over repeated probes, which multiplies the budget, or requesting a deterministic setting if the endpoint offers one.
You are assessing backdoor risk with a library that emits poisoned training rows, and every configuration you test costs a full retrain of a production-sized model. How do you design the sweep so the finding holds up, without running hundreds of trainings?
basics
~20 sSweep one axis that changes the answer, the poison fraction, on your real training recipe. Explore cheaply on a smaller model or data subset, then confirm only the interesting points at full scale. Repeat the borderline configuration across several seeds, because retrain variance can be bigger than the effect you are claiming.
An adversarial-example library returns 1000 perturbed rows against a tabular loan model and reports that 91% are misclassified, but many rows hold fractional values in integer-only columns and two category indicators set at once. Where do you put the rule the library could not express, and what number do you report instead?
basics
~20 sRun the library's output through a validity predicate you write, before you score anything. Keep only rows a real applicant could actually submit, re-query the model on those, and report misclassified-and-valid over examples attempted, together with the share you discarded. The 91% was computed before that filter and overstates what is reachable.
Your team has standardised on one adversarial-robustness toolkit and reuses the same three attack classes on every model assessment. What is the case for and against that as a program-level policy, and how would you decide the class set per engagement instead?
basics
~20 sFor: comparability across assessments, reviewable tooling, faster onboarding, predictable cost. Against: the three classes encode one fixed access level and data type, so on a differently shaped target a clean result means only that those three did not apply. Decide per engagement from granted access, input domain and attacker action space, keeping a small fixed core purely for trend comparison.
A client asks you to shortlist defences for an image classifier using an adversarial-robustness toolkit, and you have one GPU for two weeks. The candidates split into defence objects that attach to the existing weights and trainer-style defences that need a full retrain per hyperparameter setting. How do you allocate the compute, and what do you tell the client the shortlist does and does not cover?
basics
~20 sSplit by cost. Attach-only defences reuse the trained weights, so each costs one evaluation sweep and can face many attack settings. Every trainer variant costs a full retrain, so you can afford only a couple of settings. Spend most of the compute on deep adaptive evaluation of a short list, not a broad shallow sweep.
For a client robustness engagement you must decide how much of the production serving stack the model wrapper reproduces: the bare model, the model plus preprocessing, or the full path including a rule-based blocklist and a three-model ensemble vote. How do you decide, and what may the report claim in each case?
basics
~20 sDecide by what the client will act on. A bare-model wrapper measures a component; a wrapper reproducing preprocessing, rules and the ensemble vote measures the service. Pick the widest layer you can reproduce faithfully and cheaply, then state in the report exactly which layers the number covers and which were excluded.
You lead ML security for several product teams and must standardise how adversarial-robustness evaluations are run. How do you decide between adopting a third-party attack harness, building a thin in-house driver over the attack libraries, or letting each team script directly — and how do you keep that decision reversible?
basics
~20 sDecide by what must be repeatable across teams versus what each team must tune. Adopt a third-party harness for a shared target contract and result format; keep the attack calls in library code you own so parameters stay reachable. Keep it reversible: the harness should be swappable without rewriting past evaluations.
You have a fixed spend of API calls for a query-only adversarial robustness assessment of a metered model, with no gradient access. How do you split it between number of examples, per-example query cap, and number of attacks, and what does each split cost the conclusion?
basics
~20 sThere is no right split, only a stated one. Wide and shallow covers many examples at a low cap and mostly measures the cap. Narrow and deep gives credible per-example results with wide error bars. Usually: one cheap screening pass, then deep runs on a stratified subsample, holding budget back for re-runs.
Your organisation publishes robustness numbers for many models, and every increase in an attack's iteration and restart arguments costs GPU hours you have to budget. How would you set a house minimum attack-strength standard that a run must meet before its number is allowed to be published?
basics
~20 sDefine two tiers. A screening tier may use cheap settings, and its results may only report hits, never robustness. A publishable tier requires a versioned configuration — attacks, effort, threat model — plus evidence the success rate plateaued and a forced-success sanity artefact. Same configuration for every model, so numbers stay comparable.
A lab extraction run using an adversarial-robustness library recovered a substitute that agrees with your image classifier on most held-out inputs, after a few hundred thousand queries to a local wrapper. What do you tell the risk owner this does and does not establish about the deployed endpoint?
basics
~20 sIt shows the attack works against what you gave it: full probability outputs, no rate limit, no monitoring, and a query pool close to the training distribution. It does not show the deployed endpoint is exploitable. Restate it as production questions: what the endpoint returns, what the queries would cost, and whether anyone would notice.
For a domain rule an adversarial-example library cannot express, you can either discard invalid examples after generation or write a projection into the attack's iteration so every step lands on a legal input. How do you decide which, and what does each let you claim?
basics
~20 sFilter afterwards for a cheap first read: it is a few lines, but the search optimised in a space you then discard, so survivors are partly luck and the rate is loose. Project inside the loop when the number must mean something: it costs custom code, and you are no longer running a stock attack.
Your organisation wants to gate model releases on one robustness number from an adversarial library — a mean perturbation size over only the examples an attack flipped. What breaks when that number is tracked release over release, and what would you gate on instead?
basics
~20 sIt is a conditional mean whose population changes every release, so it is not monotone in robustness and moves for reasons unrelated to the model. It also collapses to zero when nothing is flipped. Gate instead on attack success rate at a pinned perturbation budget over a frozen example set, with row counts recorded.
showing 31–49 of 49