skip to content

Running an Assessment

Between an API call and a defensible result sit calls about access, input realism and attack strength. Interviewers probe them because a library default answers none of the three for you.

on this pageshow

explore

questions

20

You are about to launch a query-based (black-box) attack from an adversarial robustness toolkit against a hosted model endpoint that bills per call. How do you estimate what the run will cost in calls before you start it?

level: juniorimportance: must knowfreq 55%

answer

  1. examples x cap x restarts, plus clean pass
  2. failures spend the full cap
  3. pilot ten examples, count real calls
  4. retries and batch padding are billed
  5. hard ceiling that raises, not warns

basics

~20 s

Multiply the number of evaluation examples by the attack's per-example query cap, times restarts, and add one clean pass to find which examples are classified correctly. Attacks spend the whole cap on examples they fail. Set the cap explicitly, pilot on ten examples, measure the real call count, then scale.

solid answer

~50 s

The worst case is close to the real case, because a query attack that fails burns its entire per-example allowance. So the arithmetic is: examples selected for attack, times the per-example query cap, times restarts or repeats, plus a clean evaluation pass over the whole set to pick the correctly classified examples worth attacking. Two corrections matter. Successful examples finish early and cost less than the cap, so the true number sits between the cap-times-examples ceiling and whatever the pilot measured. And retries against a flaky or rate-limited endpoint are billed calls too. The reliable move is to run ten examples with counting instrumented in the wrapper, read the observed calls-per-example distribution, then extrapolate with headroom rather than trusting a documented default. Never launch a full set against a metered target without an explicit cap set; the default is chosen for a local model where a call is free.

code

python · 9 lines
python
n_examples = 500
clean_accuracy = 0.9
per_example_cap = 5000
restarts = 3

clean_pass = n_examples
attacked = int(n_examples * clean_accuracy)
worst_case = attacked * per_example_cap * restarts
print(clean_pass + worst_case)

go deeper

for a junior

Should produce the multiplication — examples times per-example cap — and know to pilot on a handful first rather than launching the full set.

for a middle

Should add restarts, the clean pass, early-stop savings on successes, and instrument a counter in the wrapper.

for a senior

Should account for retries, batch padding and vote repeats, and enforce a hard spend ceiling that aborts the run.

for a principal

Should treat the call budget as a contract term agreed before the engagement, with an agreed behaviour when it is exhausted mid-run.

### The arithmetic A query-only run against a metered endpoint has four multipliers and one fixed term, and you can write it down before you start: ``` clean_pass = n_examples # one call each, to find correctly classified inputs attacked = n_examples * clean_accuracy # only correctly classified inputs are worth attacking worst_case = attacked * per_example_cap * restarts ceiling = clean_pass + worst_case ``` `per_example_cap` is the attack object's own evaluation budget — `max_eval` on `ART's HopSkipJump`, `nb_parallel` and `max_iter` together on `ART's ZooAttack`, `max_iter` and `nb_restarts` on `ART's SquareAttack`. Read it off the constructed object, not off a blog post: these defaults were chosen for a model running locally where a forward pass is free, and `ART's HopSkipJump` alone defaults into the five figures of evaluations per example. ### Why the ceiling is close to the truth Newcomers assume the ceiling is wild pessimism and quote an average instead. It is not, and the reason matters. An example the attack *solves* stops early — often far under the cap — and costs little. An example the attack *cannot* solve consumes every query it is permitted before the loop gives up. So the spend concentrates on the failures, and on a genuinely robust target the failures are the majority. Reasoning with a mean call count taken from a soft target is the single most reliable way to be wrong by an order of magnitude on a hard one. ### What inflates the estimate beyond the arithmetic - **Transport retries and rate-limit backoff.** The HTTP client resends below the wrapper, so the attack's internal counter never sees them, but they reached the endpoint and they are billed. - **Vote repeats.** If the endpoint is non-deterministic and you take a majority over three probes to stabilise a near-boundary label, you have multiplied every term by three. - **Batch padding.** A wrapper that always posts fixed-size batches pays for the slots it did not fill. - **A gradient-estimation shim.** Attaching one to a forward-only wrapper turns each attack step into a batch of calls, multiplying the per-example count by one to three orders of magnitude. ### What shrinks it A stratified subsample instead of the full evaluation set. Stopping at first success rather than continuing to minimise the perturbation. A query cache keyed on the exact input, which pays off more than people expect because boundary searches revisit points. And a cheap screening attack first, reserving the expensive family for the examples that survive it. ### Where the number misleads The estimate you can trust is a measured one. Run ten examples with a counter instrumented in the wrapper, read the observed calls-per-example distribution, then extrapolate with headroom. But extrapolating the *pilot mean* is its own trap: a ten-example pilot drawn from easy inputs will mostly show early successes and will under-predict a full run by a large factor, because the full run will contain the failures that spend the whole cap. Extrapolate the ceiling and treat the pilot as a downward correction with a stated confidence, not as the forecast. Equally, a bill that comes in far *under* the ceiling is not automatically good news — it usually means the attack stopped early for a reason you have not diagnosed, such as errors being counted as evaluations, and the resulting success rate is then a measurement of your error path. ### What I would check That a call counter exists in the wrapper and its number reconciles against the provider's usage dashboard for the same window. That a hard spend ceiling is installed which raises an exception and aborts, rather than logging a warning nobody reads. That the cap on the attack object was set explicitly by you, and appears in the run log next to the result. And that the clean pass and any retries are in the estimate — a plan that omits them is short by a predictable, avoidable amount.

  • Why is the worst-case estimate usually close to the actual bill on a hard target?
    Because examples the attack fails on spend their entire per-example allowance, and on a robust target those are the majority.
  • Name two sources of billed calls that the attack's own query counter will not show you.
    Transport retries and rate-limit backoff issued by the HTTP client, and padding in fixed-size batches sent by the wrapper.

saying these in an interview costs you the question

  • Quoting only an average cost per example and ignoring that failed examples consume the whole cap
  • Leaving the library's default query budget in place against a metered endpoint
  • No call counter in the run, so the spend is only known from the invoice
  • Forgetting the clean evaluation pass and any retries in the estimate

context

open as a page

You build an evasion attack from an adversarial-robustness library such as Foolbox, torchattacks or the Adversarial Robustness Toolbox and pass only the wrapped model — no iteration count, no step size, no number of random restarts. Where do those values come from, and what were they chosen for?

level: juniorimportance: must knowfreq 60%

basics

~20 s

They come from the attack class's own constructor defaults, baked in so the library's documentation example runs fast on a laptop. They are demo settings, not assessment settings. A model that resists them has resisted only a short, weak search. Choose and record every strength argument yourself before reporting a robustness number.

open as a page

You run a data-poisoning attack from an adversarial-robustness library and it hands back arrays of training samples and labels. Why is that not yet a result, and what has to happen before you can say whether the model is vulnerable?

level: juniorimportance: must knowfreq 50%

basics

~20 s

The library only crafts tainted training data; it does not train anything. You have to mix those rows into the training set, retrain the model with your normal recipe, then evaluate it twice: normal accuracy on clean test data and the attack's success on the triggered inputs. The verdict costs a retrain, not one call.

open as a page

When you wrap a model for an adversarial-example library such as the Adversarial Robustness Toolbox, you declare a permitted minimum and maximum for the input values. What does that declaration change about the examples the attack generates, and what does it not?

level: juniorimportance: must knowfreq 70%

basics

~20 s

When you wrap a model for an adversarial-example library you declare the legal minimum and maximum for input values. The attack clips every generated example back inside that box, so pixels stay in range. It is one global range over all features, and it knows nothing about what any individual feature means.

open as a page

In an adversarial robustness library such as the Adversarial Robustness Toolbox or Foolbox, every attack runs against a wrapper object you supply for the model under test. If that wrapper can only send an input to a remote inference API and return the class probabilities it gets back, which attack families can you still run, and why do the rest fail?

level: middleimportance: must knowfreq 68%

basics

~20 s

Attacks call methods on your wrapper. Gradient attacks need a gradient method, which a forward-only HTTP wrapper cannot implement, so they error or are unavailable. You are left with query attacks: score-based ones that read the returned probabilities, decision-based ones needing only the top label, and transfer from a local surrogate.

open as a page

For an iterative gradient evasion attack driven from an adversarial-robustness library such as the Adversarial Robustness Toolbox or torchattacks, what does raising each of these three arguments buy you — the iteration count, the step size relative to the perturbation bound, and the number of random restarts — and what does each cost?

level: middleimportance: must knowfreq 55%

basics

~20 s

More iterations let the search refine longer inside the same bound. A step size too large for the bound overshoots and one too small never reaches its edge. More restarts re-launch from fresh random starting points, so one unlucky start is not read as robustness. Each knob multiplies GPU time roughly linearly.

open as a page

You retrained a classifier on data produced by a backdoor-poisoning routine in an adversarial-robustness library. Which two numbers do you report, and why is the trigger success rate on its own a misleading result?

level: middleimportance: must knowfreq 50%

basics

~20 s

Report two: accuracy on a clean, untriggered test set against the unpoisoned baseline, and the trigger success rate measured only on samples not already in the target class. Success rate alone hides a poisoned model that lost obvious accuracy, and it is inflated by inputs the clean model already sent to the target.

open as a page

You are attacking a tabular fraud model with an adversarial-example library, and 12 of its 40 features are ones the attacker cannot influence, such as account tenure. The library lets you pass a mask of which features may move. What does that mask guarantee, and what still has to be enforced outside the library?

level: middleimportance: must knowfreq 60%

basics

~20 s

The mask freezes coordinates: features you mark immovable keep their original values, so the search only touches the ones you allow. It says nothing about the movable features' own rules, such as integer counts, one-hot exclusivity, or a field derived from another. Those you check yourself after the library returns its examples.

open as a page

An evasion run driven from an adversarial-robustness library finishes and reports zero successful adversarial examples at the configured perturbation bound. Before you write 'the model is robust', what checks do you run on the attack configuration and the model wrapper itself?

level: seniorimportance: must knowfreq 50%

basics

~20 s

Assume the harness is broken first. Re-run with the bound relaxed until the attack must succeed; if it still fails, the wrapper or the loop is wrong. Then sweep iterations and restarts upward and watch the success rate. Also check a cheap non-gradient baseline: if it beats the gradient attack, the gradients are unreliable.

open as a page

Rather than choosing iteration counts and restarts yourself, a teammate proposes measuring robustness with a standardized evaluation suite such as RobustBench, which pins the attacks and their settings for you. What does that fix about the defaults problem, and what does it still not tell you about your own deployed model?

level: middleimportance: should knowfreq 35%

basics

~20 s

It fixes comparability: everyone runs the same fixed attack set at the same effort and threat model, so numbers can be ranked and nobody quietly under-configures. It does not fix relevance. It measures one threat model on one dataset through a required interface, not your inputs, your preprocessing pipeline or the attacks your product actually faces.

open as a page

In an adversarial-robustness library, a model-extraction (stealing) attack will not run until you have supplied a substitute model of your own. What exactly does the library do with it, and which parts of the run stay your responsibility?

level: middleimportance: should knowfreq 42%

basics

~20 s

You supply two things: a wrapper around the victim that answers queries, and an untrained substitute model you chose and configured. The library queries the victim over inputs you provide, labels them, and fits your substitute. It returns your object, now trained. Architecture, query pool and the query bill are all yours.

open as a page

In TextAttack, an attack recipe pairs a transformation that proposes candidate rewrites with a set of constraints those candidates must pass. What role do the constraints play in the search, and what happens to your reported success rate if you relax them?

level: middleimportance: should knowfreq 45%

basics

~20 s

Constraints are filters inside the loop: a transformation proposes candidate sentences, each constraint rejects the ones that violate it, and the search only ever sees the survivors. They are what keeps a rewrite readable and meaning-preserving. Loosen them and the success rate rises, because you are now counting rewrites that changed the sentence.

open as a page

In a decision-based (label-only) attack run from an adversarial robustness toolkit against a remote endpoint, 60% of examples hit the per-example query cap without producing an adversarial example. What do you check in the run before you treat that as a property of the model, and what do you change for the next run?

level: seniorimportance: should knowfreq 42%

basics

~20 s

Treat it as censored data, not a model property. Check the queries-to-success distribution on the examples that worked, whether the perturbation was still shrinking when the cap hit, and whether errors, throttling or duplicate queries burned the allowance. Then re-run a stratified subsample at a much larger cap.

open as a page

A query-based attack loop from an adversarial robustness toolkit assumes each call to the target answers quickly and consistently for the same input. You are pointing it at a rate-limited, non-deterministic hosted endpoint that bills per call. Of retries, throttling and answer variability, which spend your call budget, which spend only wall-clock, and what do you do about each?

level: seniorimportance: should knowfreq 34%

basics

~20 s

Retries and repeated probes spend billed calls. Throttling and latency spend wall-clock, which limits how many examples you finish per day. Variability is worst: a boundary answer that flips makes the search oscillate. Handle it by voting over repeated probes, which multiplies the budget, or requesting a deterministic setting if the endpoint offers one.

open as a page

You are assessing backdoor risk with a library that emits poisoned training rows, and every configuration you test costs a full retrain of a production-sized model. How do you design the sweep so the finding holds up, without running hundreds of trainings?

level: seniorimportance: should knowfreq 38%

basics

~20 s

Sweep one axis that changes the answer, the poison fraction, on your real training recipe. Explore cheaply on a smaller model or data subset, then confirm only the interesting points at full scale. Repeat the borderline configuration across several seeds, because retrain variance can be bigger than the effect you are claiming.

open as a page

An adversarial-example library returns 1000 perturbed rows against a tabular loan model and reports that 91% are misclassified, but many rows hold fractional values in integer-only columns and two category indicators set at once. Where do you put the rule the library could not express, and what number do you report instead?

level: seniorimportance: should knowfreq 50%

basics

~20 s

Run the library's output through a validity predicate you write, before you score anything. Keep only rows a real applicant could actually submit, re-query the model on those, and report misclassified-and-valid over examples attempted, together with the share you discarded. The 91% was computed before that filter and overstates what is reachable.

open as a page

You have a fixed spend of API calls for a query-only adversarial robustness assessment of a metered model, with no gradient access. How do you split it between number of examples, per-example query cap, and number of attacks, and what does each split cost the conclusion?

level: principalimportance: should knowfreq 28%

basics

~20 s

There is no right split, only a stated one. Wide and shallow covers many examples at a low cap and mostly measures the cap. Narrow and deep gives credible per-example results with wide error bars. Usually: one cheap screening pass, then deep runs on a stratified subsample, holding budget back for re-runs.

open as a page

Your organisation publishes robustness numbers for many models, and every increase in an attack's iteration and restart arguments costs GPU hours you have to budget. How would you set a house minimum attack-strength standard that a run must meet before its number is allowed to be published?

level: principalimportance: should knowfreq 25%

basics

~20 s

Define two tiers. A screening tier may use cheap settings, and its results may only report hits, never robustness. A publishable tier requires a versioned configuration — attacks, effort, threat model — plus evidence the success rate plateaued and a forced-success sanity artefact. Same configuration for every model, so numbers stay comparable.

open as a page

A lab extraction run using an adversarial-robustness library recovered a substitute that agrees with your image classifier on most held-out inputs, after a few hundred thousand queries to a local wrapper. What do you tell the risk owner this does and does not establish about the deployed endpoint?

level: principalimportance: should knowfreq 30%

basics

~20 s

It shows the attack works against what you gave it: full probability outputs, no rate limit, no monitoring, and a query pool close to the training distribution. It does not show the deployed endpoint is exploitable. Restate it as production questions: what the endpoint returns, what the queries would cost, and whether anyone would notice.

open as a page

For a domain rule an adversarial-example library cannot express, you can either discard invalid examples after generation or write a projection into the attack's iteration so every step lands on a legal input. How do you decide which, and what does each let you claim?

level: principalimportance: should knowfreq 35%

basics

~20 s

Filter afterwards for a cheap first read: it is a few lines, but the search optimised in a space you then discard, so survivors are partly luck and the rate is loose. Project inside the loop when the number must mean something: it costs custom code, and you are no longer running a stock attack.

open as a page