Why does a model extracted from a code endpoint match it on some prompts and diverge on others?
answer
- the copy maps the query set
- no evidence means no constraint
- it still answers, just not yours
- divergence tracks distance from the queries
- measuring the decay costs more paid calls
basics
~10 sFitting constrains the copy only where it has evidence. On prompts like those the attacker paid for it reproduces the endpoint; elsewhere nothing pins it down, so it answers from its own priors.
solid answer
~50 sThe copy was fitted on a finite set of prompts the attacker chose and paid for, together with the replies those prompts returned. Wherever new prompts resemble that set, the copy has many demonstrations of what the original does and reproduces it. Wherever they do not, there is no constraint at all: the output is decided by the copy's own pretraining and architecture, so it produces something plausible that simply is not the original's answer. Divergence therefore is not random — it tracks distance from the query distribution, which makes the copy's quality a map of where the money went. Two second-order effects matter: the reference itself is stochastic, since a sampled endpoint can answer one prompt two ways, and measuring the decay honestly costs more paid queries, which is why attackers often cannot state it.
go deeper
Remember the direction: the copy is strongest exactly where the attacker paid to look, and there is nothing holding it to the original anywhere else.
Be ready to explain that fitting applies force only where pairs exist, so off the query distribution the copy's output comes from its own priors and looks plausible while being wrong.
Show you would demand the query distribution, the held-out slices, and the pre-query baseline before treating any agreement figure as a measure of what extraction achieved.
Frame the exposure by traffic: the part of your prompt distribution an attacker can cheaply cover is the part that is realistically copyable, and that shapes what is worth defending.
## The shape of a stolen model is the shape of the query set An attacker buys replies from a metered code-completion and code-repair endpoint and fits their own model to reproduce them. Afterwards they hold something that behaves like the original — in places. This question is about *which* places, and why the boundary falls where it does. ## Evidence, and only evidence Fitting a model to `(prompt, reply)` pairs pushes it toward the original's behaviour **at the points where pairs exist**, and in the neighbourhood those points generalise over. It applies no force anywhere else. There is no term in the objective that says anything about a prompt the attacker never sent. So on prompts drawn from the same mix they queried, the copy has abundant demonstrations and matches well. Move the prompts away — a language they never asked about, a repository layout they never sampled, a task framing they never used, inputs an order of magnitude longer than anything they sent — and the copy has no information from the original to go on. It does not fall silent. A generative model always emits something, and what it emits is determined by its own pretraining, its own architecture and its own biases. The output looks like a competent answer. It is just no longer the target's answer. This is why the honest description of an extracted model is not "a worse version of the original" but **"a model whose agreement with the original is a function of distance from the attacker's query distribution"**. ## Why the decay is invisible from the inside The attacker's problem is that plausible-looking divergence is indistinguishable from agreement unless you check against the original. And checking costs the same metered price as buying training data did. An attacker who spends their whole budget on training pairs has nothing left to buy an evaluation with, so they end up with either no agreement estimate or one measured on prompts drawn from the very pool they trained on — which measures the best case and reports it as the number. That is the single most common defect in an extraction write-up, and it is worth naming in an interview: **an agreement rate measured on prompts from the same distribution as the training queries is an upper bound presented as an average.** ## The reference is moving There is a second-order effect specific to a sampled generative endpoint. The original does not answer deterministically: the same prompt can come back two different ways, and the caller does not control the decoding. So "did the copy match" needs a definition — matched a single cached reply, matched a fresh draw, produced a functionally equivalent patch, passed the same tests. Those definitions give materially different numbers on the same copy. A deterministic copy that always emits the endpoint's most likely answer can even score higher against the endpoint than the endpoint scores against a second draw of itself, so self-agreement is not a ceiling; it is just a warning that the reference is noisy and the metric has to say how it handled that. ## What sharpens and what blunts the decay - **Breadth of the query mix.** A narrow, homogeneous set of prompts buys a narrow, sharp copy. A broad mix buys shallower agreement over more ground for the same money. - **Task diversity of the endpoint.** A model that does one thing has less surface to miss. A code model doing completion, repair, explanation and translation across many languages has far more room for the copy to be wrong where nobody looked. - **The copy's own capacity and pretraining.** Where the attacker's base model already resembles the target's behaviour before any queries at all, agreement off-distribution looks better than the queries earned. That is a confound, not a result: some of the measured agreement is prior overlap between two models trained on similar public material, not anything extraction bought. A serious write-up reports agreement of the *un-queried* base model as a baseline so the reader can see what the queries actually added. ## The consequence for both chairs For the red-teamer, this dictates how the finding is written: state the query distribution, state the agreement on it, state the agreement on held-out slices that were deliberately *unlike* it, and state the baseline before queries. For the person receiving the report, it dictates the first question: agreement over which prompts? A single number with no distribution attached describes an eval set, not a copy.
- Why can an attacker who spent their whole budget on training pairs struggle to state an honest agreement rate?Because agreement can only be measured against the original, and every comparison prompt is another metered call. With nothing left to spend, they either report no number or evaluate on prompts drawn from the same pool they trained on. That measures the best case and presents it as the average, which is the most common defect in an extraction write-up.
- Some of the copy's agreement might not have been bought at all. How would you separate that out?Measure the attacker's base model against the endpoint before any queries are used, and report that as a baseline. Two models trained on similar public material already agree on easy, conventional prompts. Whatever agreement existed before the first paid call was not purchased by the extraction, and a headline number that omits the baseline credits the queries with it.
- The endpoint returns sampled text, so the same prompt can come back two ways. What does that do to the agreement metric?It makes the reference stochastic, so the metric must define what counted as a match: a single cached reply, a fresh draw, a functionally equivalent patch, or the same tests passing. Those choices move the number substantially on one and the same copy, which is why an agreement rate quoted without its match definition cannot be compared to another one.
saying these in an interview costs you the question
- Describes the copy as uniformly worse rather than unevenly faithful
- Assumes the copy abstains where it has no evidence
- Evaluates agreement on prompts drawn from the training query pool
- Ignores agreement the two models had before any queries were bought
- Quotes an agreement rate with no match definition for free-form text