skip to content

Assembling a Copy

How paid replies become a model: which inputs are worth buying, what the fitted copy agrees with you about, and what timing leaks. Interviewers probe the 'they can't reach our weights' reflex.

on this pageshow

explore

questions

16

An attacker fits their own model to a paid code-completion endpoint's replies — what do they own?

level: juniorimportance: must knowfreq 65%

answer

  1. they never saw a parameter
  2. prompts in, replies out, nothing else
  3. a function, not a weight file
  4. agreement measured where they paid
  5. different size can still match

basics

~10 s

They own a model that reproduces the endpoint's observable behaviour on prompts like the ones they paid for. Not its weights, not its training data, and not its behaviour on prompts unlike those.

solid answer

~50 s

Extraction moves behaviour, not parameters. The attacker only ever sees pairs: prompts they chose, replies the endpoint returned. Fitting their own network on those pairs produces a function that agrees with the original where they sampled it, and nothing in that process carries across a single parameter — the copy can have a different size and a different architecture and still agree closely. So the asset is a behavioural relation, normally stated as an `agreement rate`: the share of held-out prompts the copy answers the way the original did. That rate is a map of where their queries landed, not a property of the copy on its own. Away from the prompts they bought, the copy is running on its own priors, not on yours. Reading out actual parameters is a separate attack with much stronger preconditions than sampled text replies.

go deeper

for a junior

Be ready to say plainly that an attacker who can only send prompts and read replies ends up with behaviour, not weights, and that the behaviour is only trustworthy where they queried.

for a middle

Explain why no parameter crosses over: the copy is fitted from input–output pairs, so its own initialisation and architecture determine its parameters while the original only supplies targets.

for a senior

Be ready to state what an extraction result should be reported in — an agreement rate, with the match definition and the prompt distribution attached — and to reject a claim that is missing either.

for a principal

Own the framing that the loss here is a behavioural asset obtained at inference prices, and that 'they never reached our weights' is not a defensible answer to a customer or a regulator.

## The vantage: pairs, and only pairs Picture a hosted code-completion and code-repair model behind a metered, token-billed API. A paying caller sends a prompt and gets back a sampled completion — text. No probability vector, no per-token scores, no attribution over the input, and a decoding temperature they do not control. Over some period they send prompts of their choosing and keep every reply. That is the entire raw material. Everything the attacker can build is a function of a finite set of `(prompt, reply)` pairs. ## What fitting on those pairs produces If the attacker trains their own model to reproduce those replies, what they get back is a **function that imitates the endpoint's input–output behaviour**. In the literature this family is called model extraction or functionality stealing, and the phrase that matters for an interview is *functionality*: the thing being taken is what the model **does**, not what it **is**. Three consequences follow directly, and interviewers probe all three: **1. No parameter crosses over.** There is no channel from the endpoint's weights into the copy. The copy's parameters are whatever the attacker's own fitting run produced from their own initialisation. Two models can agree on the vast majority of prompts while having unrelated parameter values, different layer counts, and different parameter budgets. Behavioural agreement is a relation between two functions; it implies nothing about representational identity. **2. No training data crosses over either.** The copy's training set is *the attacker's prompts plus the original's answers to them*. The original's training corpus is not in that set. Recovering content the target was trained on is a genuinely different family of attack, with different preconditions and different evidence, and it is not what fitting a copy gives you. **3. The copy's quality is shaped like the query set.** This is the fact the leaf turns on. Supervised fitting constrains a function where it has evidence. On prompts resembling those the attacker paid for, the copy has thousands of demonstrations of what the original does, and it reproduces that. On prompts far from anything they sent, it has no evidence at all, so its output is decided by its own architecture, its own pretraining and its own inductive bias. It still answers — generative models always answer — and the answer looks plausible, which is exactly why an unmeasured copy flatters itself. ## The unit the attacker actually reports Because the payoff is behavioural, the metric is behavioural: an **agreement rate**. Take a set of held-out prompts, ask both the original and the copy, and count how often the copy's reply matches the original's under some definition of match. For a classifier that is easy — same top-1 label or not. For a code endpoint returning free-form text it is a design decision: exact string match after normalisation, functional equivalence of the produced patch, a test suite passing, a judge model's verdict. Different definitions can move the number by tens of points, so an agreement rate quoted without its match definition is not comparable to another one. Note also what agreement is *not*: it is not correctness. A copy could reproduce the original's mistakes faithfully and score 90% agreement while being a mediocre code model in absolute terms; and a copy could be an excellent code model that answers differently from the original most of the time. Those are two different measurements and only the first one is about extraction. ## Why this is still worth stealing The common junior reflex — *they never touched our weights, so nothing was taken* — misses the point. The attacker's purpose is almost never to reconstruct the artefact. It is to end up with a locally owned model that behaves like yours on traffic they care about, obtained at inference prices rather than at the price of collecting data and training. What they hold afterwards costs them nothing per call, is not rate limited, and does not appear in your logs. ## How to say it in an interview "They own behaviour, scoped to where they queried. Concretely: a model of their own whose replies match ours on the prompt distribution they sampled, with agreement decaying as prompts move off it. No parameters, no training corpus, and no guarantee anywhere they didn't pay to look."

  • The copy has roughly half the parameter count of the endpoint it was fitted to. Does that invalidate the extraction?
    No. Extraction targets the input–output function, and agreement is a relation between two functions, not between two parameter sets. A smaller model can reproduce a larger one's behaviour closely on a narrow slice of prompts; it just has less room to match everywhere at once. The size difference matters for the copy's ceiling, not for whether behaviour was taken.
  • Does the copy inherit anything about the endpoint's training corpus?
    Only indirectly, through the answers it was given. The copy's training set is the attacker's own prompts labelled by the original's replies. Whether specific training content can be pulled back out of a model is a separate attack family with its own preconditions and its own evidence — fitting a copy is not that attack and should not be reported as if it were.
  • The attacker only ever received sampled text — no scores. What does that cost them?
    Each reply carries one sampled behaviour rather than the endpoint's full distribution over answers, so a single query teaches them less than a returned probability vector would and they need more queries for the same fidelity. It also means their reference is stochastic: the same prompt can come back two ways, so any agreement number has to say how that was handled.

Watching a chess player for a thousand games lets you build something that answers the same openings the same way. It does not hand you their thinking, and it says nothing about how they will handle a position you never saw them face.

saying these in an interview costs you the question

  • Says the attacker ends up holding the original's weights
  • Assumes a copy must share the original's architecture or size
  • Claims the endpoint's training corpus is recovered by fitting a copy
  • Treats agreement on paid prompts as agreement everywhere
  • Confuses agreement with the original and correctness against ground truth

context

open as a page

How does recovering a model's parameters differ from training a copy on its replies?

level: juniorimportance: must knowfreq 60%

basics

~20 s

Training a copy fits a fresh model to bought predictions, so it only approximates the original near where the queries landed. Parameter recovery instead treats precise numeric replies as equations and solves for the original weights themselves.

open as a page

An attacker pays per submission to a hosted malware-verdict API to copy it - what does that budget actually buy?

level: juniorimportance: must knowfreq 60%

basics

~20 s

Labels, not data. Executable files are free and plentiful; what costs money is the service's verdict on each one. The budget is a fixed count of labels, so which files receive them decides how good the copy gets.

open as a page

How much does inferring a paid endpoint's architecture family help an attacker copying it?

level: middleimportance: must knowfreq 52%

basics

~20 s

Not much. Incidental observation gives an architecture family and a scale, not a specification, and a copy fitted on purchased replies need not match the target's architecture. It is a prior that trims the search, not the prize.

open as a page

A paid text-embedding endpoint returns only vectors — what does its reply metadata leak?

level: juniorimportance: should knowfreq 45%

basics

~20 s

A reply's fixed vector length and dtype, latency that grows with input length, and error text for oversized input reveal an output space, a rough scale and an input ceiling — a family and a shape, never parameters.

open as a page

Why does a model extracted from a code endpoint match it on some prompts and diverge on others?

level: middleimportance: should knowfreq 52%

basics

~10 s

Fitting constrains the copy only where it has evidence. On prompts like those the attacker paid for it reproduces the endpoint; elsewhere nothing pins it down, so it answers from its own priors.

open as a page

Why does solving for a network's weights need many digits of each returned score?

level: middleimportance: should knowfreq 38%

basics

~10 s

The evidence a solve reads is a faint change of slope where one unit switches on. It lives in the low-order digits of a returned number, so a coarsely rounded reply erases it.

open as a page

Why does spending a malware-API query budget on random byte-strings buy far less agreement than real executables?

level: middleimportance: should knowfreq 48%

basics

~20 s

Random byte-strings sit far off the distribution the classifier was fitted on, so nearly all of them come back with the same verdict. A label you could have predicted before paying for it teaches the copy nothing.

open as a page

An extracted copy of your code endpoint is clearly worse than the original — why can the theft still have succeeded?

level: seniorimportance: should knowfreq 46%

basics

~20 s

Parity was never the goal. A copy that agrees with yours on the traffic the attacker cares about, and that they now run locally with no per-call bill, no rate limit and no log entry, already repays their queries.

open as a page

Your logs show a few thousand tightly clustered full-precision queries to a small risk regressor — did someone solve for its weights or merely sample it?

level: seniorimportance: should knowfreq 33%

basics

~10 s

The query pattern separates them: a solve probes tight groups of near-identical inputs and needs the full numeric reply, while a harvest is large, broad and tolerant of rounding. Logs bound material, not conclusions.

open as a page

A timing finding against your paid embedding endpoint reproduces once in five runs — how do you triage it?

level: seniorimportance: should knowfreq 38%

basics

~20 s

A one-in-five reproduction against a shared endpoint is unproven until the measurement rules out the platform: were conditions interleaved, was a null control run, does the effect exceed per-call spread? Then rate severity against what replies already disclose.

open as a page

An engineer says stealing our malware classifier needs a corpus as big as our training set - what is wrong with that?

level: seniorimportance: should knowfreq 52%

basics

~10 s

It confuses free inputs with paid labels. In-domain executables cost nothing and the service itself supplies the labels - and labels are only needed where the boundary is, not everywhere the data is.

open as a page

What does recovering a model's weights 'only up to symmetry' leave an attacker holding?

level: middleimportance: nice to knowfreq 26%

basics

~20 s

A set of parameters that computes exactly the same function as yours, but need not match yours entry by entry: hidden units can be permuted and a unit's incoming and outgoing weights rescaled against each other without changing any reply.

open as a page

Why must a latency gap measured against a hosted embedding endpoint be repeated before it means anything?

level: middleimportance: nice to knowfreq 28%

basics

~20 s

A single call's latency is dominated by network jitter, queueing and co-tenant batching, which swamp any model-shape effect. Only a difference that survives averaging many interleaved calls carries information, and confounds that move with load never average away at all.

open as a page

A report claims 94% agreement between an extracted copy and your code endpoint — what do you ask?

level: seniorimportance: nice to knowfreq 32%

basics

~20 s

Ask what agreement was measured against, on which prompt distribution, how a text match was defined, and what the un-queried base model scored. One agreement rate describes an evaluation set, not the copy — and it establishes nothing about weights.

open as a page

An extraction run reports 97% agreement with a paid malware-verdict API - what do you ask before believing it?

level: seniorimportance: nice to knowfreq 30%

basics

~10 s

Ask which pool agreement was measured on and the target's base rate there. On a pool where the service answers benign 97% of the time, a copy that always says benign also scores 97%.

open as a page