skip to content

An extracted copy of your code endpoint is clearly worse than the original — why can the theft still have succeeded?

level: seniorimportance: should knowfreq 46%

answer

  1. success is defined by their purpose
  2. parity is a counterfeiter's criterion
  3. their traffic is narrower than yours
  4. the meter and the logs are gone
  5. the gap bounds resale, not theft

basics

~20 s

Parity was never the goal. A copy that agrees with yours on the traffic the attacker cares about, and that they now run locally with no per-call bill, no rate limit and no log entry, already repays their queries.

solid answer

~50 s

Extraction is judged against the attacker's purpose, not against your benchmark scores. Two purposes are common and neither needs parity. First, a service: a model that behaves like yours on the narrow slice of traffic they actually serve, obtained at inference prices instead of the cost of collecting data and training — their margin is the test, not your leaderboard. Second, possession: a model of their own that they can run, inspect and probe as often as they like, offline, for free, outside your rate limits and outside your logging. Anything that previously required repeated metered access to you becomes unmetered local work. The quality gap does tell you something real — it bounds what they can sell and it says their copy fails somewhere — but 'ours still scores higher' is not a finding about whether behaviour was taken.

go deeper

for a junior

Remember that a copy does not need to be as good as the original to be worth having, because the attacker chose what it has to be good at.

for a middle

Be able to name the two payoffs — a cheap service on narrow traffic, and a locally owned model that costs nothing per run — and say why neither requires parity.

for a senior

Show the judgment to report agreement on the attacker's plausible traffic rather than your own benchmarks, and to state what possession removes from your visibility.

for a principal

Own the escalation call: decide when a below-parity copy is a closed finding and when it is a material loss, and be able to defend that line to people who only see the score gap.

## The reflex this question targets A capable senior engineer, shown an extracted copy that trails the production model on every internal benchmark, will often conclude that the extraction did not really work. The reasoning is that a copy is only a copy if it is as good as the thing it copies. That criterion belongs to a counterfeiter, not to an adversary, and adopting it will cause you to close a real finding. ## Success is defined by the adversary's purpose There are two purposes worth separating, and the difference between them changes what you should measure. **Purpose one: a running service.** The attacker wants an endpoint of their own that answers their users acceptably. Their users are not your users, and their traffic is not your traffic — it is usually far narrower. The relevant question is whether the copy agrees with the original *on the prompt mix they will actually serve*, at a price they can sustain. Their economics are the test: they paid inference prices for a training set; you paid data collection, curation, training and evaluation to produce the behaviour they sampled. That asymmetry does not evaporate because the copy trails you by some margin on a benchmark whose prompts neither of you serves. **Purpose two: possession.** This one is easy to underrate because there is no product at the end of it. Before extraction, everything the adversary wanted to learn about your model's behaviour cost money per attempt, was rate limited, and left records on your side. After extraction they hold a model of their own. They can run it as often as they like, look inside it, and study its behaviour offline. The metered, observable relationship you had with them is over — not because they broke anything, but because they no longer need you. Whatever they intend to do with a model that behaves like yours, they now do it for free and unobserved, and you have no telemetry on any of it. That second point is the one that makes 'ours still scores higher' a non-answer. The value of the copy is not its score. It is that it is *theirs*. ## What the quality gap does legitimately tell you Do not overcorrect into treating the gap as meaningless. It carries real information: - **It bounds resale.** A copy well below parity is hard to market as a competitor to your product on your full traffic. If the concern is a rival launching a lookalike service across the board, the gap is evidence against that specific scenario. - **It localises the failure.** A copy that trails overall is failing *somewhere*, and where it fails is diagnostic: it points at the prompt regions the attacker never bought. That tells you what their query mix looked like, which is far more useful to you than the aggregate score. - **It is not evidence of safety.** A near-random copy on your traffic would be — but that is a different observation, and it has to be measured on the traffic in question rather than inferred from a headline gap. ## How to report it without overstating either way The honest framing replaces one number with three: 1. Agreement with the original **on the prompt mix the attacker plausibly targets**, which is the number their purpose turns on. 2. Agreement on prompts deliberately unlike their query mix, which shows where the copy is hollow. 3. What the copy removes for them: per-call cost, rate limiting, and any visibility you had into their probing. Then state plainly what was not taken: no parameters, no training corpus, and no behaviour outside the region their queries covered. A report that claims more than that is as unhelpful as one that closes the finding because the copy scored lower. ## Saying it in an interview "Comparing the copy to us on our benchmarks answers a question the attacker never asked. Their test is agreement on the traffic they intend to serve, at a bill that was a fraction of what the behaviour cost us to build — plus the fact that they now own something that behaves like us and can be run and studied offline, outside our meter and our logs. The gap bounds what they can resell; it doesn't tell us nothing was taken."

  • How would you make the cost asymmetry concrete when writing this up?
    Put the two bills side by side: what the attacker spent in metered inference to collect their pairs, against what the behaviour cost to produce — data acquisition, labelling, training runs and evaluation. The point is not an exact ratio but the order of magnitude, because that gap is what makes a below-parity copy an economically rational outcome for them and a real loss for you.
  • When is a quality gap actually the right reason to close the finding?
    When the copy is near useless on the traffic the adversary would target — not when it merely trails you on your own benchmark. That means measuring agreement on their plausible prompt mix rather than yours. If agreement there is low and the pre-query baseline explains most of what remains, the extraction genuinely bought them little and you can say so with evidence.
  • The copy is behind on quality but the attacker keeps querying anyway. What does that tell you?
    That they are still buying coverage, so their current copy is not yet good enough for their purpose — which is itself a signal about what that purpose is. Continued querying concentrated in one region says the copy is being shored up exactly there, and that region is the part of your behaviour they care about.

saying these in an interview costs you the question

  • Treats a benchmark gap as proof extraction failed
  • Assumes the attacker's goal must be reselling a parity model
  • Ignores that a local copy removes per-call cost and rate limits
  • Forgets the attacker's probing is now invisible in your logs
  • Compares the copy only on your traffic, never on theirs

context