An attacker fits their own model to a paid code-completion endpoint's replies — what do they own?
answer
- they never saw a parameter
- prompts in, replies out, nothing else
- a function, not a weight file
- agreement measured where they paid
- different size can still match
basics
~10 sThey own a model that reproduces the endpoint's observable behaviour on prompts like the ones they paid for. Not its weights, not its training data, and not its behaviour on prompts unlike those.
solid answer
~50 sExtraction moves behaviour, not parameters. The attacker only ever sees pairs: prompts they chose, replies the endpoint returned. Fitting their own network on those pairs produces a function that agrees with the original where they sampled it, and nothing in that process carries across a single parameter — the copy can have a different size and a different architecture and still agree closely. So the asset is a behavioural relation, normally stated as an `agreement rate`: the share of held-out prompts the copy answers the way the original did. That rate is a map of where their queries landed, not a property of the copy on its own. Away from the prompts they bought, the copy is running on its own priors, not on yours. Reading out actual parameters is a separate attack with much stronger preconditions than sampled text replies.
go deeper
Be ready to say plainly that an attacker who can only send prompts and read replies ends up with behaviour, not weights, and that the behaviour is only trustworthy where they queried.
Explain why no parameter crosses over: the copy is fitted from input–output pairs, so its own initialisation and architecture determine its parameters while the original only supplies targets.
Be ready to state what an extraction result should be reported in — an agreement rate, with the match definition and the prompt distribution attached — and to reject a claim that is missing either.
Own the framing that the loss here is a behavioural asset obtained at inference prices, and that 'they never reached our weights' is not a defensible answer to a customer or a regulator.
## The vantage: pairs, and only pairs Picture a hosted code-completion and code-repair model behind a metered, token-billed API. A paying caller sends a prompt and gets back a sampled completion — text. No probability vector, no per-token scores, no attribution over the input, and a decoding temperature they do not control. Over some period they send prompts of their choosing and keep every reply. That is the entire raw material. Everything the attacker can build is a function of a finite set of `(prompt, reply)` pairs. ## What fitting on those pairs produces If the attacker trains their own model to reproduce those replies, what they get back is a **function that imitates the endpoint's input–output behaviour**. In the literature this family is called model extraction or functionality stealing, and the phrase that matters for an interview is *functionality*: the thing being taken is what the model **does**, not what it **is**. Three consequences follow directly, and interviewers probe all three: **1. No parameter crosses over.** There is no channel from the endpoint's weights into the copy. The copy's parameters are whatever the attacker's own fitting run produced from their own initialisation. Two models can agree on the vast majority of prompts while having unrelated parameter values, different layer counts, and different parameter budgets. Behavioural agreement is a relation between two functions; it implies nothing about representational identity. **2. No training data crosses over either.** The copy's training set is *the attacker's prompts plus the original's answers to them*. The original's training corpus is not in that set. Recovering content the target was trained on is a genuinely different family of attack, with different preconditions and different evidence, and it is not what fitting a copy gives you. **3. The copy's quality is shaped like the query set.** This is the fact the leaf turns on. Supervised fitting constrains a function where it has evidence. On prompts resembling those the attacker paid for, the copy has thousands of demonstrations of what the original does, and it reproduces that. On prompts far from anything they sent, it has no evidence at all, so its output is decided by its own architecture, its own pretraining and its own inductive bias. It still answers — generative models always answer — and the answer looks plausible, which is exactly why an unmeasured copy flatters itself. ## The unit the attacker actually reports Because the payoff is behavioural, the metric is behavioural: an **agreement rate**. Take a set of held-out prompts, ask both the original and the copy, and count how often the copy's reply matches the original's under some definition of match. For a classifier that is easy — same top-1 label or not. For a code endpoint returning free-form text it is a design decision: exact string match after normalisation, functional equivalence of the produced patch, a test suite passing, a judge model's verdict. Different definitions can move the number by tens of points, so an agreement rate quoted without its match definition is not comparable to another one. Note also what agreement is *not*: it is not correctness. A copy could reproduce the original's mistakes faithfully and score 90% agreement while being a mediocre code model in absolute terms; and a copy could be an excellent code model that answers differently from the original most of the time. Those are two different measurements and only the first one is about extraction. ## Why this is still worth stealing The common junior reflex — *they never touched our weights, so nothing was taken* — misses the point. The attacker's purpose is almost never to reconstruct the artefact. It is to end up with a locally owned model that behaves like yours on traffic they care about, obtained at inference prices rather than at the price of collecting data and training. What they hold afterwards costs them nothing per call, is not rate limited, and does not appear in your logs. ## How to say it in an interview "They own behaviour, scoped to where they queried. Concretely: a model of their own whose replies match ours on the prompt distribution they sampled, with agreement decaying as prompts move off it. No parameters, no training corpus, and no guarantee anywhere they didn't pay to look."
- The copy has roughly half the parameter count of the endpoint it was fitted to. Does that invalidate the extraction?No. Extraction targets the input–output function, and agreement is a relation between two functions, not between two parameter sets. A smaller model can reproduce a larger one's behaviour closely on a narrow slice of prompts; it just has less room to match everywhere at once. The size difference matters for the copy's ceiling, not for whether behaviour was taken.
- Does the copy inherit anything about the endpoint's training corpus?Only indirectly, through the answers it was given. The copy's training set is the attacker's own prompts labelled by the original's replies. Whether specific training content can be pulled back out of a model is a separate attack family with its own preconditions and its own evidence — fitting a copy is not that attack and should not be reported as if it were.
- The attacker only ever received sampled text — no scores. What does that cost them?Each reply carries one sampled behaviour rather than the endpoint's full distribution over answers, so a single query teaches them less than a returned probability vector would and they need more queries for the same fidelity. It also means their reference is stochastic: the same prompt can come back two ways, so any agreement number has to say how that was handled.
Watching a chess player for a thousand games lets you build something that answers the same openings the same way. It does not hand you their thinking, and it says nothing about how they will handle a position you never saw them face.
saying these in an interview costs you the question
- Says the attacker ends up holding the original's weights
- Assumes a copy must share the original's architecture or size
- Claims the endpoint's training corpus is recovered by fitting a copy
- Treats agreement on paid prompts as agreement everywhere
- Confuses agreement with the original and correctness against ground truth