In an adversarial-robustness library, a model-extraction (stealing) attack will not run until you have supplied a substitute model of your own. What exactly does the library do with it, and which parts of the run stay your responsibility?
answer
- library fills the substitute you bring
- wrapper output shape drives fidelity
- query pool is the experiment
- agreement on held-out, not accuracy
- count the queries that really went out
basics
~20 sYou supply two things: a wrapper around the victim that answers queries, and an untrained substitute model you chose and configured. The library queries the victim over inputs you provide, labels them, and fits your substitute. It returns your object, now trained. Architecture, query pool and the query bill are all yours.
solid answer
~50 sThe extraction routine is a **training loop driver**, not a model factory. Its inputs are the wrapped victim, an untrained model you instantiate yourself, and a pool of unlabelled inputs. It sends those inputs to the victim, takes whatever the wrapper returns as supervision, and fits your model on the resulting pairs. What comes back is the same object you handed in, with weights. Three consequences the interviewer is listening for. First, **the result is bounded by your choices** — a substitute with far less capacity than the victim will plateau, and you will report that as "extraction failed" when you actually measured your own architecture. Second, **the query pool is the experiment**. Feeding in-distribution data you happened to have makes the attack look far easier than one where the attacker must synthesise inputs. Third, **every query is billed by whatever the wrapper talks to**: against a metered endpoint the query budget you set is money spent.
go deeper
Knows the substitute is an object you create and pass in, and that the library trains it from the victim's answers.
Explains that the query pool, the substitute's capacity and the wrapper's output shape all set the ceiling, and that queries cost money against a metered endpoint.
Instruments the actual query count, reports agreement as a curve over budget, and states which output shape and pool the number belongs to.
Decides whether the extraction question is worth an engagement's query budget at all, and what a result would change about how the endpoint is exposed.
Extraction sits on the same structural boundary as poisoning: the library will not produce the artefact for you, it **fills one you brought**. In ART the routine is a `KnockoffNets`- or `CopycatCNN`-style attack, constructed around the wrapped victim and then invoked as `attack.extract(x, thieved_classifier=my_untrained_wrapper)`. The mechanism is unglamorous and worth stating plainly: it takes your pool of unlabelled inputs, sends them through the victim wrapper's `predict`, treats whatever comes back as supervision, and fits *your* substitute on the resulting input/output pairs. The object returned is the same object you handed in, now with weights. Standing a run up therefore means making four decisions the API will cheerfully let you make badly. ### The victim wrapper — what it returns is the biggest lever The wrapper is your code. If it returns a full class-probability vector, every query carries the victim's confidence over *all* classes: a far richer training signal than a single label, and agreement climbs much faster for the same budget. A lab wrapper written for convenience almost always returns the rich form, because that is what `predict` gives you locally. A deployed endpoint frequently returns a top-1 label, or a label plus a coarse confidence bucket. These are different experiments, and their query budgets are not interchangeable. Record which one you measured, in the finding itself. ### The substitute — your architecture bounds the result Architecture, initialisation, optimiser and epoch count are yours. A substitute with far less capacity than the victim will plateau, and the honest reading of that plateau is "my substitute was too small", not "extraction failed". Only a substitute with adequate capacity makes a negative result interesting, and even then it is a statement about *that* substitute. ### The query pool — this is the experiment Extraction fidelity is dominated by where in input space the queries land. A pool drawn from the victim's own training distribution is the easy case and often the one a lab reaches for because the data is already on disk. A realistic attacker may be confined to public data, to a different domain, or to synthetic inputs. The same attack against those pools costs a different budget for the same fidelity, so the pool belongs in the finding beside the number. ### The budget — a curve, not a point Set the query budget explicitly and treat it as a first-class axis. **Agreement as a function of query count** is the useful artefact, because it shows where fidelity saturates; a single point tells a reader nothing about whether ten times the budget would have doubled the result or changed nothing. ### What it costs Two separate bills, and candidates routinely quote only one. **Queries**: a few hundred thousand calls at a typical metered per-call price is real money, and against your own hosted model it is inference capacity you are diverting. **Compute**: fitting the substitute is an ordinary training run per budget point, so a five-point curve is five trainings on top of the queries. **Wall clock**: if the target is rate-limited, the query bill converts to time — a limit of a few calls per second turns a few hundred thousand queries into days, and that conversion is often the finding. ### Where the number misleads The dominant error is reporting the substitute's **accuracy against ground-truth labels** as though it were the extraction result. That measures how good the stolen model is at the task; it does not measure how faithfully it *copies the victim*. A substitute can outscore the victim on true labels while disagreeing with it constantly — different decision boundary, similar competence — and it can trail the victim while matching it closely. The figure that answers the extraction question is **agreement**: the fraction of inputs on which substitute and victim emit the same label, computed on held-out data that was **not** in the query pool. Score it on the pool and you have measured memorisation of exactly the pairs you trained on. The second misleading reading is a budget quoted without the output shape: a number obtained against probability vectors will be badly optimistic for a top-1 endpoint. ### What I would check Instrument the wrapper and count **how many queries actually left it** rather than trusting the configured budget — retries, deduplicated inputs and internal caching all move the true number. Check that the query pool and the evaluation set do not overlap. Confirm the substitute was given enough capacity and enough epochs to have converged, so that a plateau is a real ceiling. And write down the wrapper's output shape next to the budget, because those two numbers only mean anything together.
- Why does it matter whether the wrapped victim returns a probability vector or just a top label?A probability vector is far more supervision per query, so the substitute converges at a much lower budget. A result obtained against a rich output does not transfer to an endpoint that returns one label.
- Which number do you report for a substitute, and on what data?Agreement with the victim — the fraction of inputs where both give the same label — computed on held-out data that was not in the query pool.
saying these in an interview costs you the question
- Expecting the library to choose or download a substitute architecture for you
- Calling extraction infeasible after one undersized substitute plateaued
- Measuring the substitute's accuracy against ground-truth labels instead of agreement with the victim
- Not recording whether the wrapper returned probability vectors or a single label
- Reporting one query count with no curve, so nobody can see where fidelity saturates