In Counterfit, an attack reaches a model only through a scan target you write. What must that scan target supply so a query-only attack can run against a hosted inference endpoint, and what does it deliberately not hand the attack?
answer
- adapter, not a model
- batch in, scores or label out
- declare shape, dtype, classes
- seed samples to perturb
- queries only, no gradients
basics
~20 sThe scan target wraps your endpoint as a callable: given a batch of samples it returns the model's per-class scores, or a label. You also declare the input shape and data type, the list of output classes, and a few seed samples the attack will perturb. It hands over queries only, never gradients or weights.
solid answer
~50 sThe scan target is an **adapter, not a model**. Two halves: - **A callable.** It takes a batch of samples in the shape the model expects and returns one row of scores per sample, or a single label per sample if that is all the service gives you. Every call inside it is a real, billed network request. - **Metadata the attack needs to build inputs.** Input shape and data type, the legal value range or vocabulary, the list of output classes, and a small seed set of real samples to perturb — an evasion attack modifies an existing input, it does not invent one. What it does not supply is anything gradient-shaped: no weights, no backward pass. That is the point of the adapter — it makes a hosted endpoint look like a black box, so only query-based attacks apply. If the service returns only a top-1 label, you narrow further, to attacks that work from label flips alone.
go deeper
Should say the target wraps the endpoint in a predict-style callable and declares input shape and output classes, and that no gradients are available.
Adds that the return arity decides which attack family is even applicable: a score vector means score-based search, label-only means decision-based search.
Talks about validating the input space and normalisation against the live service before trusting a scan, and about instrumenting the callable for query count from the first run.
Frames the adapter as the contract that decides what the programme can measure at all, and asks whether a staging replica should carry the scans rather than the metered production endpoint.
**What the object actually is.** Counterfit ships no attacks of its own. It is a command shell that adapts a model to third-party attack libraries — the Adversarial Robustness Toolbox for numeric and image targets, TextAttack for text — and drives them from one loop: `set_target`, `list attacks`, `set_attack`, `run_attack`, `show results`. Everything that loop will ever know about your model arrives through one Python class you subclass from Counterfit's target base and drop into its targets folder. The members of that class *are* the contract: | member of the target class | what it carries | |---|---| | `target_data_type` | `image`, `numpy` or `text` — decides which attack catalogue is even offered | | `target_input_shape` | the tensor or string shape the attack must build its perturbed inputs in | | `target_output_classes` | the ordered class list; the columns of a returned score row map onto these names | | `target_endpoint` | where `predict` sends its request, plus whatever credential it needs | | `X` | the seed samples the attack will perturb | | `predict(self, x)` | a batch of samples in, one row of per-class scores — or one label — per sample out | `predict` is the only door. Counterfit hands it to the attack library wrapped as a black-box estimator, so from the attack's point of view the model is an oracle: it can ask, it cannot look inside. Every question it asks is a real, billed network request. **What the adapter deliberately withholds.** Anything gradient-shaped — weights, a backward pass, the training loss. That is not an omission to be fixed; it is the reason the wrapper exists. A hosted endpoint genuinely has no gradient to give, and the honest wrapper says so rather than faking one. The consequence is a narrowing you should state out loud: with a full score vector you can run score-based search, which follows the confidence signal; with a bare top-1 label you are down to decision-based methods that learn only which side of the boundary they are on. Attacks written for white-box access either drop off the offered list or fall back to estimating a gradient from finite differences, and that estimate is bought with extra queries at every single step. **Why `X` is part of the contract.** Evasion perturbs an existing input toward a wrong answer; it does not invent one. Without seeds drawn from the real input distribution, and their true labels, there is no starting point and no definition of success. **What it costs.** Query count is the currency, and the currency converts twice — into money at the endpoint's per-call price, and into wall clock at its per-call latency divided by whatever concurrency you actually achieve. Orders of magnitude worth carrying: a transfer-style single-step attack is a handful of calls per sample; a score-based search that estimates a direction from finite differences is hundreds to low thousands per sample; a label-only boundary walk runs into the tens of thousands. At 150 ms a call with no concurrency, ten thousand queries is about twenty-five minutes — for one sample. **Where the number misleads.** Four readings go wrong, and none of them raises an error. - **Space mismatch.** You declare one input range or normalisation in `target_input_shape` and the surrounding code, and the service applies a different resize, scaling or tokenisation. The attack then perturbs in a space the model never sees, every attack under-performs, and the scan reports low success — which reads as robustness and is not. - **Class-order mismatch.** If the columns of the row `predict` returns are not in the order of `target_output_classes`, success is scored against the wrong label. This can inflate as easily as deflate, and the result file looks perfectly normal. - **Return-arity drift.** A scalar label where a score vector was expected either crashes loudly or, worse, silently degrades a score-based attack to noise-following. - **Silently un-batched.** `predict` loops one HTTP call per sample while the attack believes it sent one batch. Nothing is numerically wrong; the query count, the wall clock and any per-call rate limit are all off by the batch size. **What I would check before firing anything.** Round-trip a single seed through `predict` and print the raw response — shape, dtype, precision, whether it is top-k. Confirm `argmax` of that row names the class a local copy of the same model gives for the same input; systematic disagreement on *clean* inputs is a preprocessing bug, not an attack result. Confirm the column order against `target_output_classes` explicitly. Distinguish an error response from a prediction so a throttled call never counts as a survived attack. And put a counter inside `predict` from the first run, so the first attack on the first seed tells you what a scan really costs instead of leaving it to the invoice.
- The service returns only a top-1 class name and no scores. What does that rule out?Every attack that optimises on a confidence signal. You are left with decision-based methods that search using only whether the label flipped, and those typically cost far more queries per sample.
- Why does the scan target need seed samples at all?An evasion attack perturbs an existing input toward a misclassification; it needs a starting point in the real input distribution, plus its true label, to know whether it succeeded.
- Your wrapper loops one HTTP call per sample while the attack passes batches. What breaks?Nothing numerically, but the query count and wall-clock time explode versus your estimate, and any per-call rate limit now applies per sample instead of per batch.
saying these in an interview costs you the question
- Says the target 'loads the model' — it wraps a remote endpoint and never sees weights.
- Assumes gradients or a backward pass are somehow available through the hosted endpoint.
- Forgets that the attack needs seed samples and thinks it generates inputs from scratch.
- Never mentions that every attack step inside the callable is a billed network call.