skip to content

A report claims 94% agreement between an extracted copy and your code endpoint — what do you ask?

level: seniorimportance: nice to knowfreq 32%

answer

  1. a number describes an eval set
  2. agreement with replies, or with truth?
  3. which prompt distribution?
  4. what counted as a match?
  5. what did it score before paying?

basics

~20 s

Ask what agreement was measured against, on which prompt distribution, how a text match was defined, and what the un-queried base model scored. One agreement rate describes an evaluation set, not the copy — and it establishes nothing about weights.

solid answer

~40 s

Four questions, in order. Agreement with **what** — the endpoint's own replies, or ground-truth correctness? Only the first measures extraction; the second measures the copy's quality and is a different claim. On **which prompts** — held out from the same pool the training queries came from, or deliberately unlike them? Same-pool evaluation reports a ceiling as an average. Under **what match rule** — for free-form code, exact normalised text, functional equivalence, or a passing test suite give materially different numbers on one copy. And against **what baseline** — how much did the attacker's base model already agree before a single query was bought? Then say what 94% does not establish: no parameters were obtained, no training corpus, and no claim holds for prompts unlike the eval mix.

code

text · 14 lines
text
copy vs. endpoint — agreement on returned completions
match rule: normalised exact match of the emitted patch

slice                                    n      agreement
---------------------------------------------------------
prompts from the queried mix           2000       94.1%
same languages, unseen repositories      800       88.6%
languages never queried                  800       41.3%
long multi-file repair tasks             400       37.9%
---------------------------------------------------------
baseline: attacker base model, no queries bought  33.5%
reference note: endpoint answers one prompt identically
  on two draws in 81.2% of cases
...

go deeper

for a junior

Remember that an agreement percentage is only meaningful with the prompt set and the match rule attached; on its own it is a number, not a result.

for a middle

Be able to distinguish agreement with the endpoint's replies from correctness against ground truth, and explain why only the first measures extraction.

for a senior

Show that you would demand per-slice numbers and a pre-query baseline, and that you can state precisely what a high agreement rate does and does not establish.

for a principal

Own what gets claimed externally: decide what an extraction finding may assert to customers or counsel, and hold the line that behavioural agreement is never evidence about parameters.

## Why a single agreement rate is almost never a claim An extraction result is normally reported as an agreement rate: the share of prompts on which the copy answered the way the original did. It is the right unit — extraction takes behaviour, so behaviour is what should be measured. But the number is a property of *a measurement*, and four choices inside that measurement can each move it by tens of points. A headline figure with none of them attached is not comparable to any other figure and should not be accepted. ## The four columns that must be attached **1. Agreement against what?** Two very different quantities get called the same thing. *Agreement with the original's replies* is the extraction measurement. *Correctness against ground truth* is a quality measurement, and a copy can score well on one and badly on the other. A copy that faithfully reproduces the original's mistakes has high agreement and mediocre quality; a strong independent model can have excellent quality and low agreement. Only the first says anything about what was taken. **2. Measured on which prompts?** Agreement is highest exactly where the attacker's queries landed. Evaluating on prompts drawn from the same pool the training replies were bought from measures the best case. The informative report gives slices: prompts like the query mix, prompts adjacent to it, and prompts deliberately unlike it. **3. Under what match rule?** For a classifier this is trivial — same label or not. For a code endpoint returning free-form text it is a decision: exact string match after normalisation, an equivalent patch, the same tests passing, or a judge's verdict. Those give different numbers on the same copy, and the reference is stochastic on top of that, since a sampled endpoint can answer one prompt two ways. Note the direction carefully: a deterministic copy emitting the endpoint's most likely answer can score higher against the endpoint than the endpoint scores against a second draw of itself, so self-agreement is not a ceiling — it is a note that the reference is noisy and the rule must say how that was handled. **4. Against what baseline?** Two models trained on similar public material already answer conventional prompts the same way. Whatever agreement the attacker's base model had *before the first paid query* was not bought by the extraction. A report that omits it credits the query budget with agreement that was free. ## Reading a result table Given per-slice numbers, the headline usually turns out to be the top row. The gap between the first and last rows is the actual finding: it says how narrow the copy is and, by implication, how narrow the attacker's query mix was. If the pre-query baseline row is high, most of the headline was never purchased at all. ## What the number cannot establish, however large - **Nothing about parameters.** Agreement is a relation between two functions. Reading out actual parameters is a separate attack with different preconditions, and no behavioural agreement rate is evidence for it. - **Nothing about the training corpus.** The copy was fitted on the attacker's prompts and the endpoint's answers to them. - **Nothing outside the eval distribution.** The claim's scope is the prompt mix it was measured on and no wider. - **Nothing about the copy's usefulness to its owner** unless the eval mix resembles the traffic they intend to serve. ## The limitations paragraph this produces A red-teamer writing the limitations section should be able to write, without hedging: *agreement was measured against the endpoint's sampled replies under a stated match rule, on prompts drawn from the queried mix and on two slices deliberately outside it; the base model's pre-query agreement is reported as a baseline; the result establishes behavioural imitation on the queried distribution only, and no parameters or training data were obtained.* Any claim broader than that sentence is not supported by an agreement rate, and any report that cannot fill in those clauses has published a number rather than a finding.

  • The report scored the copy against ground-truth correctness rather than the endpoint's replies. Why does that mislead?
    Because it measures how good the copy is, not how closely it imitates you. A copy can be independently competent and still answer differently from your model most of the time, and it can reproduce your errors faithfully while scoring poorly on correctness. Only agreement with your own replies speaks to whether behaviour was extracted.
  • The pre-query baseline is 34% and the headline is 94%. How does that change your reading?
    It tells you roughly a third of the agreement existed before any money was spent, because two models trained on similar public material converge on conventional answers. The queries bought the rest. Without that row the reader credits the whole 94% to the extraction, which overstates both the attacker's achievement and the exposure of your endpoint.
  • How would you word the limitations line so it survives review?
    State the match rule, the prompt slices and their sizes, the baseline, and then the scope: behavioural imitation on the queried distribution only. Add explicitly that no parameters and no training data were obtained, since that is what a reader will otherwise assume a high agreement rate implies.

saying these in an interview costs you the question

  • Accepts one agreement rate with no prompt distribution attached
  • Conflates agreement with the endpoint and correctness against ground truth
  • Omits how a match was defined for free-form text output
  • Never reports the base model's agreement before queries were bought
  • Reads a high agreement rate as evidence about parameters

context