skip to content

Model Extraction & Stealing

You will learn how an attacker clones a paywalled model through nothing but prediction queries, and what it costs them relative to training from scratch. Interviewers use it to test whether you can reason about an ML API as an economic attack surface and defend it with rate limits, output hardening, and watermark-based ownership proofs.

on this pageshow

explore

questions

page 1 of 2

An attacker fits their own model to a paid code-completion endpoint's replies — what do they own?

level: juniorimportance: must knowfreq 65%

answer

  1. they never saw a parameter
  2. prompts in, replies out, nothing else
  3. a function, not a weight file
  4. agreement measured where they paid
  5. different size can still match

basics

~10 s

They own a model that reproduces the endpoint's observable behaviour on prompts like the ones they paid for. Not its weights, not its training data, and not its behaviour on prompts unlike those.

solid answer

~50 s

Extraction moves behaviour, not parameters. The attacker only ever sees pairs: prompts they chose, replies the endpoint returned. Fitting their own network on those pairs produces a function that agrees with the original where they sampled it, and nothing in that process carries across a single parameter — the copy can have a different size and a different architecture and still agree closely. So the asset is a behavioural relation, normally stated as an `agreement rate`: the share of held-out prompts the copy answers the way the original did. That rate is a map of where their queries landed, not a property of the copy on its own. Away from the prompts they bought, the copy is running on its own priors, not on yours. Reading out actual parameters is a separate attack with much stronger preconditions than sampled text replies.

go deeper

for a junior

Be ready to say plainly that an attacker who can only send prompts and read replies ends up with behaviour, not weights, and that the behaviour is only trustworthy where they queried.

for a middle

Explain why no parameter crosses over: the copy is fitted from input–output pairs, so its own initialisation and architecture determine its parameters while the original only supplies targets.

for a senior

Be ready to state what an extraction result should be reported in — an agreement rate, with the match definition and the prompt distribution attached — and to reject a claim that is missing either.

for a principal

Own the framing that the loss here is a behavioural asset obtained at inference prices, and that 'they never reached our weights' is not a defensible answer to a customer or a regulator.

## The vantage: pairs, and only pairs Picture a hosted code-completion and code-repair model behind a metered, token-billed API. A paying caller sends a prompt and gets back a sampled completion — text. No probability vector, no per-token scores, no attribution over the input, and a decoding temperature they do not control. Over some period they send prompts of their choosing and keep every reply. That is the entire raw material. Everything the attacker can build is a function of a finite set of `(prompt, reply)` pairs. ## What fitting on those pairs produces If the attacker trains their own model to reproduce those replies, what they get back is a **function that imitates the endpoint's input–output behaviour**. In the literature this family is called model extraction or functionality stealing, and the phrase that matters for an interview is *functionality*: the thing being taken is what the model **does**, not what it **is**. Three consequences follow directly, and interviewers probe all three: **1. No parameter crosses over.** There is no channel from the endpoint's weights into the copy. The copy's parameters are whatever the attacker's own fitting run produced from their own initialisation. Two models can agree on the vast majority of prompts while having unrelated parameter values, different layer counts, and different parameter budgets. Behavioural agreement is a relation between two functions; it implies nothing about representational identity. **2. No training data crosses over either.** The copy's training set is *the attacker's prompts plus the original's answers to them*. The original's training corpus is not in that set. Recovering content the target was trained on is a genuinely different family of attack, with different preconditions and different evidence, and it is not what fitting a copy gives you. **3. The copy's quality is shaped like the query set.** This is the fact the leaf turns on. Supervised fitting constrains a function where it has evidence. On prompts resembling those the attacker paid for, the copy has thousands of demonstrations of what the original does, and it reproduces that. On prompts far from anything they sent, it has no evidence at all, so its output is decided by its own architecture, its own pretraining and its own inductive bias. It still answers — generative models always answer — and the answer looks plausible, which is exactly why an unmeasured copy flatters itself. ## The unit the attacker actually reports Because the payoff is behavioural, the metric is behavioural: an **agreement rate**. Take a set of held-out prompts, ask both the original and the copy, and count how often the copy's reply matches the original's under some definition of match. For a classifier that is easy — same top-1 label or not. For a code endpoint returning free-form text it is a design decision: exact string match after normalisation, functional equivalence of the produced patch, a test suite passing, a judge model's verdict. Different definitions can move the number by tens of points, so an agreement rate quoted without its match definition is not comparable to another one. Note also what agreement is *not*: it is not correctness. A copy could reproduce the original's mistakes faithfully and score 90% agreement while being a mediocre code model in absolute terms; and a copy could be an excellent code model that answers differently from the original most of the time. Those are two different measurements and only the first one is about extraction. ## Why this is still worth stealing The common junior reflex — *they never touched our weights, so nothing was taken* — misses the point. The attacker's purpose is almost never to reconstruct the artefact. It is to end up with a locally owned model that behaves like yours on traffic they care about, obtained at inference prices rather than at the price of collecting data and training. What they hold afterwards costs them nothing per call, is not rate limited, and does not appear in your logs. ## How to say it in an interview "They own behaviour, scoped to where they queried. Concretely: a model of their own whose replies match ours on the prompt distribution they sampled, with agreement decaying as prompts move off it. No parameters, no training corpus, and no guarantee anywhere they didn't pay to look."

  • The copy has roughly half the parameter count of the endpoint it was fitted to. Does that invalidate the extraction?
    No. Extraction targets the input–output function, and agreement is a relation between two functions, not between two parameter sets. A smaller model can reproduce a larger one's behaviour closely on a narrow slice of prompts; it just has less room to match everywhere at once. The size difference matters for the copy's ceiling, not for whether behaviour was taken.
  • Does the copy inherit anything about the endpoint's training corpus?
    Only indirectly, through the answers it was given. The copy's training set is the attacker's own prompts labelled by the original's replies. Whether specific training content can be pulled back out of a model is a separate attack family with its own preconditions and its own evidence — fitting a copy is not that attack and should not be reported as if it were.
  • The attacker only ever received sampled text — no scores. What does that cost them?
    Each reply carries one sampled behaviour rather than the endpoint's full distribution over answers, so a single query teaches them less than a returned probability vector would and they need more queries for the same fidelity. It also means their reference is stochastic: the same prompt can come back two ways, so any agreement number has to say how that was handled.

Watching a chess player for a thousand games lets you build something that answers the same openings the same way. It does not hand you their thinking, and it says nothing about how they will handle a position you never saw them face.

saying these in an interview costs you the question

  • Says the attacker ends up holding the original's weights
  • Assumes a copy must share the original's architecture or size
  • Claims the endpoint's training corpus is recovered by fitting a copy
  • Treats agreement on paid prompts as agreement everywhere
  • Confuses agreement with the original and correctness against ground truth

context

open as a page

How does recovering a model's parameters differ from training a copy on its replies?

level: juniorimportance: must knowfreq 60%

basics

~20 s

Training a copy fits a fresh model to bought predictions, so it only approximates the original near where the queries landed. Parameter recovery instead treats precise numeric replies as equations and solves for the original weights themselves.

open as a page

An attacker pays per submission to a hosted malware-verdict API to copy it - what does that budget actually buy?

level: juniorimportance: must knowfreq 60%

basics

~20 s

Labels, not data. Executable files are free and plentiful; what costs money is the service's verdict on each one. The budget is a fixed count of labels, so which files receive them decides how good the copy gets.

open as a page

Why would an attacker pay to query a model API instead of training their own model?

level: juniorimportance: must knowfreq 62%

basics

~20 s

Because each paid call returns the target's answer on an input the attacker chose, so the endpoint sells labels. Copying by query wins only when that bill comes in under collecting, labelling and training on their own data.

open as a page

A scoring API caps each key at 1,000 calls a day - why doesn't that stop model extraction?

level: juniorimportance: must knowfreq 70%

basics

~20 s

The cap is attached to a key, and keys are cheap. Someone who can open self-serve accounts buys any total volume they want; the cap only fixes how many identities they need. The binding price is identity, not the per-key count.

open as a page

A claim-triage API returns a per-feature contribution list with each score. What does that hand the caller?

level: juniorimportance: must knowfreq 50%

basics

~20 s

It hands the caller the model's local sensitivity: for that record, which input fields move the score, in which direction, and roughly how much. That is information about the model, not only about one decision.

open as a page

Why does normal paid use of a content-moderation API also hand the caller a labelled dataset?

level: juniorimportance: must knowfreq 70%

basics

~20 s

Because the caller supplies the text and the endpoint supplies the annotation. Every paid call returns a scored row they can keep, so ordinary integration quietly accumulates the labelled corpus your annotators were paid to produce.

open as a page

How much does inferring a paid endpoint's architecture family help an attacker copying it?

level: middleimportance: must knowfreq 52%

basics

~20 s

Not much. Incidental observation gives an architecture family and a scale, not a specification, and a copy fitted on purchased replies need not match the target's architecture. It is a prior that trims the search, not the prize.

open as a page

Your ownership test fires on a competitor's detector — what must you measure before calling it evidence?

level: seniorimportance: must knowfreq 44%

basics

~20 s

The test's false-positive rate on detectors known to be trained independently. Models built from overlapping public data agree on a great deal, so a test reading agreement as ownership also fires on models nobody stole.

open as a page

Your moderation API owner says full category scores are safe since the weights stay private. What's wrong?

level: seniorimportance: must knowfreq 58%

basics

~20 s

Nobody fitting a copy wants the parameters. A returned score vector is already shaped like a training target, so the endpoint acts as a teacher and a competitor buys a functional stand-in at query prices.

open as a page

A paid text-embedding endpoint returns only vectors — what does its reply metadata leak?

level: juniorimportance: should knowfreq 45%

basics

~20 s

A reply's fixed vector length and dtype, latency that grows with input length, and error text for oversized input reveal an output space, a rough scale and an input ceiling — a family and a shape, never parameters.

open as a page

Your face matcher may be inside a rival's API that returns only a match decision — what do watermarking and fingerprinting each claim?

level: juniorimportance: should knowfreq 48%

basics

~20 s

Watermarking plants a behaviour during your training run that a copy inherits, so a hit evidences derivation. Fingerprinting plants nothing: it picks inputs where your trained matcher already behaves distinctively and asks whether the suspect agrees.

open as a page

Why does squeezing a stolen object detector onto edge cameras erase the owner's watermark?

level: juniorimportance: should knowfreq 46%

basics

~20 s

A planted watermark is only extra behaviour held in the weights. Fitting the stolen detector to a camera reshapes those weights using footage that never exercises the mark, so it fades as a side effect nobody aimed at.

open as a page

Why does a model extracted from a code endpoint match it on some prompts and diverge on others?

level: middleimportance: should knowfreq 52%

basics

~10 s

Fitting constrains the copy only where it has evidence. On prompts like those the attacker paid for it reproduces the endpoint; elsewhere nothing pins it down, so it answers from its own priors.

open as a page

Why does solving for a network's weights need many digits of each returned score?

level: middleimportance: should knowfreq 38%

basics

~10 s

The evidence a solve reads is a faint change of slope where one unit switches on. It lives in the low-order digits of a returned number, so a coarsely rounded reply erases it.

open as a page

Why does spending a malware-API query budget on random byte-strings buy far less agreement than real executables?

level: middleimportance: should knowfreq 48%

basics

~20 s

Random byte-strings sit far off the distribution the classifier was fitted on, so nearly all of them come back with the same verdict. A label you could have predicted before paying for it teaches the copy nothing.

open as a page

What determines the query bill for cloning a paid speech-to-text API?

level: middleimportance: should knowfreq 46%

basics

~20 s

Two numbers multiplied: the price per call and the number of calls the wanted fidelity needs. Matching a target across its whole input range costs far more than reaching sellable accuracy on the narrow slice the adversary intends to serve.

open as a page

Rounding a paid API's returned relevance score to two decimals - who pays, and how much?

level: middleimportance: should knowfreq 48%

basics

~20 s

Coarsening the reply costs a party building a functional copy a multiplier on their query bill and nothing more. It costs the honest bidder the identical lost resolution in the number their bid is computed from.

open as a page

What do you give up in advance to plant an ownership mark in a face matcher your licensees could copy?

level: middleimportance: should knowfreq 38%

basics

~20 s

Three things, all before any theft: clean accuracy spent on behaviour the task never needed; a commitment bound to that training run; and a probe set that must stay secret, since verifying hands it to the suspect.

open as a page

Why is re-distilling a stolen detector harsher on a planted watermark than a short fine-tune?

level: middleimportance: should knowfreq 33%

basics

~20 s

Re-distillation keeps none of the original weights. A fresh network learns only behaviour the holder's transfer inputs expose, so a mark keyed to secret inputs is never queried and never taught. A fine-tune at least starts from the marked weights.

open as a page

Why does an API that returns an attribution with each score undercut its own query-budget defence?

level: middleimportance: should knowfreq 40%

basics

~20 s

The price is charged per call, but the value taken is per number returned. An explained reply carries a whole row of local sensitivity instead of one score, so a call limit sized against score-only traffic is loose by roughly the number of features disclosed.

open as a page

In a multi-label moderation API, how does the response body set an extraction campaign's cost per labelled row?

level: middleimportance: should knowfreq 52%

basics

~20 s

One call returns one score per policy category, so a single payment annotates that text against every category at once. The response schema, not the call count, fixes how much supervision a dollar of API spend buys an attacker.

open as a page

An extracted copy of your code endpoint is clearly worse than the original — why can the theft still have succeeded?

level: seniorimportance: should knowfreq 46%

basics

~20 s

Parity was never the goal. A copy that agrees with yours on the traffic the attacker cares about, and that they now run locally with no per-call bill, no rate limit and no log entry, already repays their queries.

open as a page

Your logs show a few thousand tightly clustered full-precision queries to a small risk regressor — did someone solve for its weights or merely sample it?

level: seniorimportance: should knowfreq 33%

basics

~10 s

The query pattern separates them: a solve probes tight groups of near-identical inputs and needs the full numeric reply, while a harvest is large, broad and tolerant of rounding. Logs bound material, not conclusions.

open as a page

A timing finding against your paid embedding endpoint reproduces once in five runs — how do you triage it?

level: seniorimportance: should knowfreq 38%

basics

~20 s

A one-in-five reproduction against a shared endpoint is unproven until the measurement rules out the platform: were conditions interleaved, was a null control run, does the effect exceed per-call spread? Then rate severity against what replies already disclose.

open as a page

An engineer says stealing our malware classifier needs a corpus as big as our training set - what is wrong with that?

level: seniorimportance: should knowfreq 52%

basics

~10 s

It confuses free inputs with paid labels. In-domain executables cost nothing and the service itself supplies the labels - and labels are only needed where the boundary is, not everywhere the data is.

open as a page

Why is "our model was cheap to train" a weak reason to ignore extraction?

level: seniorimportance: should knowfreq 42%

basics

~20 s

Because the adversary's alternative is data plus training, and data usually dominates. A model trained in a day on years of purchased labels is expensive to reproduce. The sum also settles only resale, not the other motives.

open as a page

Why can query distribution flag an extraction campaign when query volume never will?

level: seniorimportance: should knowfreq 38%

basics

~20 s

Product traffic is demand-weighted and repetitive; a campaign fitting a copy needs coverage, including inputs no customer asks about. Volume can be spread across cheap identities until every count is ordinary; that shape cannot, unless the adversary pays to imitate it.

open as a page

A copy of your face matcher sits in someone else's product: is a planted mark or a fingerprint easier to erase?

level: seniorimportance: should knowfreq 30%

basics

~20 s

The planted mark. It is behaviour the task never needed, so a holder loses it without giving up anything they stole. A fingerprint sits on ordinary close-call decisions, so moving off it degrades the matcher worth copying.

open as a page

A broker integrator archived a year of explained claim-triage replies. What can they build, and what stays out of reach?

level: seniorimportance: should knowfreq 33%

basics

~20 s

They hold claims paired with scores and per-field sensitivity rows — enough to fit a close functional stand-in over the segment their own book covers, and to read off which fields drive referrals. Not the parameters, not behaviour outside that segment, and not yet a changed decision.

open as a page

showing 1–30 of 40