skip to content

In an adversarial-ML evaluation stack, what is the difference between an attack library (such as the Adversarial Robustness Toolbox or Foolbox) and a command-line harness that drives one (such as Counterfit)? Which of the two decides what attacks are available to you and how far you can tune them?

level: juniorimportance: must knowfreq 62%

answer

  1. library implements, harness drives
  2. attack menu is a subset
  3. parameters live in the library
  4. harness adds repeatability, not capability
  5. drop a layer when blocked

basics

~20 s

The attack library implements the attacks and holds the parameters; the harness only drives it, doing target setup, batching, logging and output. So the library decides which attacks exist and how tunable they are. The harness decides how conveniently you run them, and may expose fewer attacks or fewer parameters.

solid answer

~50 s

Two layers with different jobs. The **attack library** contains the implementations: the optimisation loop, the perturbation constraints, the model-access assumptions. It is imported as code and its attack objects take parameters directly. The **harness** is a driver: it defines a uniform way to describe the thing under test, picks an attack by name, runs it, and writes a result record. It does not add attack capability; it adds repeatability and reporting. So the ceiling on *what you can attack with* is set by the library. The harness only sets the ceiling on *what of that library is reachable through the harness*, usually a curated attack list with a curated slice of each attack's parameters. When someone says "the tool doesn't support that attack", ask which layer they mean: the fix is either upgrading or contributing to the library, or bypassing the harness and calling the library yourself.

go deeper

for a junior

Should state plainly that the library holds the attacks and the harness runs them, and that the harness's list of attacks may be smaller than the library's.

for a middle

Adds that the parameter surface is also a subset, and that a result belongs to the library's attack at given settings rather than to the harness.

for a senior

Frames the harness as a driver you can drop out of, discusses what to verify before adopting one, and insists write-ups name library, attack and settings.

for a principal

Treats the layering as an organisational choice: what is standardised (target contract, result schema) versus what stays in library code teams own and tune.

Get this separation right early, because most later confusion on an adversarial-robustness engagement — whose number is it, why did the attack "fail", why did the bill arrive — traces back to collapsing the two layers into one word: "the tool". ## The library layer, at the level of the object An **attack library** is a Python package you import into your own process. Take the Adversarial Robustness Toolbox (ART) as the worked example. 1. You first wrap the thing under test in an *estimator*: `art.estimators.classification.PyTorchClassifier(model=..., loss=..., input_shape=..., nb_classes=..., clip_values=(0, 1))`. The estimator is the adapter — it teaches the library how to get predictions out of your model and, for gradient attacks, how to get gradients. 2. Then you construct an *attack object*: `art.attacks.evasion.ProjectedGradientDescent(estimator=clf, eps=0.03, eps_step=0.007, max_iter=40, num_random_init=5)`, and call `attack.generate(x=x_test)`. Foolbox has the same shape under different names — `foolbox.PyTorchModel(model, bounds=(0, 1))` plus `foolbox.attacks.LinfPGD()` called with an `epsilons=` list. Every knob that decides whether the attack is *strong* is a constructor argument on a library object: - `eps` (how much perturbation is permitted), - `eps_step` (how far one iteration moves), - `max_iter` (how long the search runs), - `num_random_init` (how many restarts), - and the estimator's `clip_values` (the valid input range). ## The harness layer A **harness** is a driver over those objects. Counterfit is the canonical command-line example: an interactive shell in which - `list targets` shows the target classes you registered, - `interact <target>` selects one, - `list attacks` shows the attacks it re-exports from the libraries beneath it, - `use <attack>` picks one, - `show options` prints the subset of that attack's arguments it surfaces, - `set <param> <value>` changes them, - `run` executes, and `show results` prints a table. It contains no optimisation code of its own. What it adds is a uniform target description, a loop over a menu, failure handling partway through a sweep, and a comparable result record — roughly a day of driver work you skip, and a day you would otherwise re-spend on every new project. ## What a run costs, and who sets that The cost lives in the library's parameters no matter who invokes them. A white-box PGD at `max_iter=40` with five restarts is about 200 forward-and-backward passes per input: seconds on a GPU, minutes for a few hundred inputs. A black-box decision-based attack is four orders of magnitude away — ART's `HopSkipJump` defaults to `max_iter=50` with `max_eval=10000`, i.e. tens of thousands of *model calls for one single input*. Point that at a metered hosted endpoint from a harness that offers one convenient `run` command, with a hundred test inputs, and you have committed to a seven-figure call count and a real invoice. The harness's virtue — one uniform interface — is exactly what hides the fact that two adjacent menu entries differ in cost by 10,000x. ## Where the number misleads - The classic is a **unit mismatch** that the harness makes invisible. ART's `eps=0.3` default is written for images normalised to [0, 1] — the MNIST convention. Register a target whose inputs are 0-255 pixels, or unnormalised tabular features, leave `eps` at the default because `show options` never surfaced it, and the attack is searching inside a ball too small to change the input at all. Every row of `show results` reads "failed", and the tempting reading — "the model is robust" — is wrong; the true reading is "the units were wrong". - The second misreading is **coverage**: "we ran everything the tool listed" has as its denominator *attacks the harness exposes*, not attacks that exist for that model type. Both errors report the harness's limits as the model's strength. ## What to check - Run `pip freeze` inside the harness's environment to learn which library and version actually produced the result — that names the real owner of the number. - Diff `show options` against the library attack's constructor signature to see what is fixed silently. - Confirm `clip_values`/`bounds` and `eps` are in your model's real input units by running one attack at a deliberately huge `eps` and verifying it succeeds; if it cannot win when handed an absurd allowance, the plumbing is broken. - And check you can still `import` the library directly in the same environment, because that is your escape hatch when the menu runs out.

  • You need an attack family your harness does not list, but the library underneath implements it. What are your options?
    Call the library directly in a small script, or extend the harness with an entry for that attack. Either way the capability was already there; only the driver was missing.
  • Why does a robustness write-up need to name the library alongside the harness?
    Because the attack implementation and its defaults live in the library. Two harness runs that look identical can differ if the library beneath changed the implementation or its default parameters.
  • What does a harness give you that an ad-hoc script does not?
    A uniform description of the thing under test, a consistent result record across attacks and engineers, and a run loop you did not have to write or maintain.

A harness is a menu over a kitchen: a dish missing from the menu is no evidence the kitchen cannot cook it. The arguments the menu leaves off are still ordered on your behalf, at whatever the chef's default happens to be.

saying these in an interview costs you the question

  • Treating the harness as the source of the attack, e.g. calling a result 'the harness's robustness score' with no library or attack named.
  • Assuming the harness exposes everything the underlying library can do.
  • Not knowing you can bypass the harness and import the library directly.
  • Reporting an attack outcome without naming the settings it ran at.

context