skip to content

Counterfit

A command-loop layer turns a hosted model into a scan target and an attack library into a repeatable battery, and hands you its dependency pins too. Interviewers ask what that layer really buys.

on this pageshow

explore

questions

9

In an adversarial-ML evaluation stack, what is the difference between an attack library (such as the Adversarial Robustness Toolbox or Foolbox) and a command-line harness that drives one (such as Counterfit)? Which of the two decides what attacks are available to you and how far you can tune them?

level: juniorimportance: must knowfreq 62%

answer

  1. library implements, harness drives
  2. attack menu is a subset
  3. parameters live in the library
  4. harness adds repeatability, not capability
  5. drop a layer when blocked

basics

~20 s

The attack library implements the attacks and holds the parameters; the harness only drives it, doing target setup, batching, logging and output. So the library decides which attacks exist and how tunable they are. The harness decides how conveniently you run them, and may expose fewer attacks or fewer parameters.

solid answer

~50 s

Two layers with different jobs. The **attack library** contains the implementations: the optimisation loop, the perturbation constraints, the model-access assumptions. It is imported as code and its attack objects take parameters directly. The **harness** is a driver: it defines a uniform way to describe the thing under test, picks an attack by name, runs it, and writes a result record. It does not add attack capability; it adds repeatability and reporting. So the ceiling on *what you can attack with* is set by the library. The harness only sets the ceiling on *what of that library is reachable through the harness*, usually a curated attack list with a curated slice of each attack's parameters. When someone says "the tool doesn't support that attack", ask which layer they mean: the fix is either upgrading or contributing to the library, or bypassing the harness and calling the library yourself.

go deeper

for a junior

Should state plainly that the library holds the attacks and the harness runs them, and that the harness's list of attacks may be smaller than the library's.

for a middle

Adds that the parameter surface is also a subset, and that a result belongs to the library's attack at given settings rather than to the harness.

for a senior

Frames the harness as a driver you can drop out of, discusses what to verify before adopting one, and insists write-ups name library, attack and settings.

for a principal

Treats the layering as an organisational choice: what is standardised (target contract, result schema) versus what stays in library code teams own and tune.

Get this separation right early, because most later confusion on an adversarial-robustness engagement — whose number is it, why did the attack "fail", why did the bill arrive — traces back to collapsing the two layers into one word: "the tool". ## The library layer, at the level of the object An **attack library** is a Python package you import into your own process. Take the Adversarial Robustness Toolbox (ART) as the worked example. 1. You first wrap the thing under test in an *estimator*: `art.estimators.classification.PyTorchClassifier(model=..., loss=..., input_shape=..., nb_classes=..., clip_values=(0, 1))`. The estimator is the adapter — it teaches the library how to get predictions out of your model and, for gradient attacks, how to get gradients. 2. Then you construct an *attack object*: `art.attacks.evasion.ProjectedGradientDescent(estimator=clf, eps=0.03, eps_step=0.007, max_iter=40, num_random_init=5)`, and call `attack.generate(x=x_test)`. Foolbox has the same shape under different names — `foolbox.PyTorchModel(model, bounds=(0, 1))` plus `foolbox.attacks.LinfPGD()` called with an `epsilons=` list. Every knob that decides whether the attack is *strong* is a constructor argument on a library object: - `eps` (how much perturbation is permitted), - `eps_step` (how far one iteration moves), - `max_iter` (how long the search runs), - `num_random_init` (how many restarts), - and the estimator's `clip_values` (the valid input range). ## The harness layer A **harness** is a driver over those objects. Counterfit is the canonical command-line example: an interactive shell in which - `list targets` shows the target classes you registered, - `interact <target>` selects one, - `list attacks` shows the attacks it re-exports from the libraries beneath it, - `use <attack>` picks one, - `show options` prints the subset of that attack's arguments it surfaces, - `set <param> <value>` changes them, - `run` executes, and `show results` prints a table. It contains no optimisation code of its own. What it adds is a uniform target description, a loop over a menu, failure handling partway through a sweep, and a comparable result record — roughly a day of driver work you skip, and a day you would otherwise re-spend on every new project. ## What a run costs, and who sets that The cost lives in the library's parameters no matter who invokes them. A white-box PGD at `max_iter=40` with five restarts is about 200 forward-and-backward passes per input: seconds on a GPU, minutes for a few hundred inputs. A black-box decision-based attack is four orders of magnitude away — ART's `HopSkipJump` defaults to `max_iter=50` with `max_eval=10000`, i.e. tens of thousands of *model calls for one single input*. Point that at a metered hosted endpoint from a harness that offers one convenient `run` command, with a hundred test inputs, and you have committed to a seven-figure call count and a real invoice. The harness's virtue — one uniform interface — is exactly what hides the fact that two adjacent menu entries differ in cost by 10,000x. ## Where the number misleads - The classic is a **unit mismatch** that the harness makes invisible. ART's `eps=0.3` default is written for images normalised to [0, 1] — the MNIST convention. Register a target whose inputs are 0-255 pixels, or unnormalised tabular features, leave `eps` at the default because `show options` never surfaced it, and the attack is searching inside a ball too small to change the input at all. Every row of `show results` reads "failed", and the tempting reading — "the model is robust" — is wrong; the true reading is "the units were wrong". - The second misreading is **coverage**: "we ran everything the tool listed" has as its denominator *attacks the harness exposes*, not attacks that exist for that model type. Both errors report the harness's limits as the model's strength. ## What to check - Run `pip freeze` inside the harness's environment to learn which library and version actually produced the result — that names the real owner of the number. - Diff `show options` against the library attack's constructor signature to see what is fixed silently. - Confirm `clip_values`/`bounds` and `eps` are in your model's real input units by running one attack at a deliberately huge `eps` and verifying it succeeds; if it cannot win when handed an absurd allowance, the plumbing is broken. - And check you can still `import` the library directly in the same environment, because that is your escape hatch when the menu runs out.

  • You need an attack family your harness does not list, but the library underneath implements it. What are your options?
    Call the library directly in a small script, or extend the harness with an entry for that attack. Either way the capability was already there; only the driver was missing.
  • Why does a robustness write-up need to name the library alongside the harness?
    Because the attack implementation and its defaults live in the library. Two harness runs that look identical can differ if the library beneath changed the implementation or its default parameters.
  • What does a harness give you that an ad-hoc script does not?
    A uniform description of the thing under test, a consistent result record across attacks and engineers, and a run loop you did not have to write or maintain.

A harness is a menu over a kitchen: a dish missing from the menu is no evidence the kitchen cannot cook it. The arguments the menu leaves off are still ordered on your behalf, at whatever the chef's default happens to be.

saying these in an interview costs you the question

  • Treating the harness as the source of the attack, e.g. calling a result 'the harness's robustness score' with no library or attack named.
  • Assuming the harness exposes everything the underlying library can do.
  • Not knowing you can bypass the harness and import the library directly.
  • Reporting an attack outcome without naming the settings it ran at.

context

open as a page

In Counterfit, an attack reaches a model only through a scan target you write. What must that scan target supply so a query-only attack can run against a hosted inference endpoint, and what does it deliberately not hand the attack?

level: juniorimportance: must knowfreq 62%

basics

~20 s

The scan target wraps your endpoint as a callable: given a batch of samples it returns the model's per-class scores, or a label. You also declare the input shape and data type, the list of output classes, and a few seed samples the attack will perturb. It hands over queries only, never gradients or weights.

open as a page

You can run an evasion evaluation either through a command-line wrapper over adversarial attack libraries (such as Counterfit) or by importing the attack library (such as the Adversarial Robustness Toolbox or Foolbox) and writing your own driver. What does the wrapper layer buy you, and what does it charge you?

level: middleimportance: must knowfreq 58%

basics

~20 s

The wrapper gives you one uniform way to point attacks at a model plus a repeatable loop: consistent target definition, batch runs, logging and result output you did not write. You pay with its dependency pins, its supported platform list, and a parameter surface narrower than the library's own attack classes.

open as a page

You are about to fire a Counterfit scan that runs several named attacks one after another against an endpoint billed per call. How do you estimate the number of model queries the whole scan will make before you start it?

level: middleimportance: must knowfreq 55%

basics

~20 s

Count per attack, then add up. Queries are roughly seed samples times iterations times calls per iteration, and a black-box attack spends many calls per step estimating a direction or probing a boundary. Do not trust the arithmetic alone: put a counter in the target's callable, run one attack on one sample, and extrapolate.

open as a page

A robustness write-up says an evasion attack "failed" against a model, but the run was driven through a command-line wrapper that exposes only some of the underlying attack library's parameters. Why is "the attack failed" not yet a supportable claim, and what do you do before it goes in a report?

level: seniorimportance: must knowfreq 48%

basics

~20 s

Because the run only tested the attack at the settings the wrapper exposed. An attack with an untuned step size, iteration count or perturbation budget failing says nothing about robustness. Before publishing, check which parameters the layer passes through, sweep the ones that matter, and record the exact settings tested.

open as a page

An adversarial-attack command-line wrapper installs with its own pinned versions of the ML framework and the attack libraries it wraps. What problems does that create when you add it to an existing evaluation environment, and how do you contain them?

level: middleimportance: should knowfreq 40%

basics

~20 s

It drags a whole transitive stack into your environment: its pinned framework and attack-library versions can conflict with what your training or serving code needs. Contain it by giving the wrapper its own isolated environment or container and moving data across as files, rather than importing it into an existing project.

open as a page

Your Counterfit scan target points at a live production inference endpoint instead of a locally loaded model. What does that change about what the scan result is actually measuring, and about your ability to repeat the run?

level: seniorimportance: should knowfreq 42%

basics

~20 s

You stop testing a model and start testing a deployment: request validation, preprocessing, the served copy of the weights, any filter in front, and caching. The result describes that deployment at that hour. A failed attack can mean an error response or a version change rather than robustness, so the run is not cleanly repeatable.

open as a page

An overnight Counterfit scan that had run fine all week burned a month of endpoint quota after a colleague switched on the optional search over attack parameters. Where does that multiplier come from, and how would you bound the next run?

level: seniorimportance: should knowfreq 40%

basics

~20 s

The parameter search does not tune inside one attack run; it reruns the entire attack once per trial. Cost becomes attacks times trials times the full per-attack query count, and that per-attack count is already seeds times iterations. Ten trials over three attacks is thirty complete attacks against a metered endpoint.

open as a page

You lead ML security for several product teams and must standardise how adversarial-robustness evaluations are run. How do you decide between adopting a third-party attack harness, building a thin in-house driver over the attack libraries, or letting each team script directly — and how do you keep that decision reversible?

level: principalimportance: should knowfreq 30%

basics

~20 s

Decide by what must be repeatable across teams versus what each team must tune. Adopt a third-party harness for a shared target contract and result format; keep the attack calls in library code you own so parameters stay reachable. Keep it reversible: the harness should be swappable without rewriting past evaluations.

open as a page