skip to content

With only generation access to a fine-tuned code model, which training strings come back verbatim?

level: middleimportance: should knowfreq 55%

answer

  1. two families, not one
  2. count the copies
  3. structureless strings have no rule behind them
  4. the attacker cannot enlarge the set

basics

~20 s

Two families: spans duplicated many times across the corpus, and rare high-entropy spans the model scores unusually confidently. Everything else comes back as paraphrase. The set is a property of the corpus, not of how hard the attacker tries.

solid answer

~50 s

Recovery concentrates in two families and misses almost everything else. First, heavily duplicated spans — a licence header, a config block copied into every service, a fixture reused across repositories — because so much of the training objective rides on that one sequence. Second, low-duplication but high-entropy spans, because a structureless literal has no rule behind it, so any unusually low loss the model achieves on it came from storing it. Ordinary prose seen a handful of times comes back as fluent paraphrase, not as itself. The two things that set the family membership — how many copies existed, and how much capacity the model has — were both fixed at training. An adversary with plain prompt-and-continue access spends queries and prefix specificity, which harvests that set faster but adds nothing to it. So in a review, the question is what classes of string were in the corpus, never how hard anyone tried.

code

text · 10 lines
text
extraction audit - internal completion model, 50k sampled continuations

span class                                 copies in corpus  tested  recovered
generated file header (boilerplate)                  1,200+      40         39
service config block, copied per repository             ~60      40         22
ordinary prose comment                                   ~8      40          3
high-entropy literal inside a reused fixture               3      40          7
high-entropy literal, single occurrence                    1      40          1
...
(recovered = span reproduced character-for-character)

go deeper

for a junior

Know the two families by name — heavily duplicated spans, and rare structureless literals — and do not claim that the whole corpus is recoverable.

for a middle

Explain why duplication and entropy both matter, and why a structureless string can be reached from very few copies while common prose never returns exactly.

for a senior

Turn this into a question about the corpus rather than the attacker: which classes of string went in, how often, and which of them would actually matter if returned.

for a principal

Own the consequence for data policy and sizing: capacity raises the ceiling and the corpus sets the candidates, so exposure is decided by what you train on, not by how the endpoint is metered.

## The set is smaller and stranger than people expect Asked what an attacker can pull out of a model with ordinary generation access, engineers give one of two wrong answers: "nothing, it is a statistical model" or "anything, given enough queries". Both are wrong in the same way — they treat recoverability as uniform across the corpus. It is not. It concentrates sharply, in two families, for two different reasons. ## Family one: duplication A span that appeared many times across the training text has an outsized claim on the objective. Every copy is another place where getting that exact sequence right lowers loss, and there is no competing pressure to abstract it, because the same characters really do follow the same prefix every time. This is why boilerplate is the most reliably extractable content in any corpus: licence headers, generated file preambles, a config block copied into forty services, a fixture pasted into every test module. The relationship is monotone in the useful direction — more copies, more reliable recovery — which is what makes duplication count the single most predictive attribute you can know about a span. ## Family two: entropy The second family is counter-intuitive and it is where the real damage lives. Take a string with no internal structure — a random-looking key, a long opaque identifier, a machine-generated token embedded in a fixture. Language modelling cannot predict it from anything. There is no rule to learn, no neighbouring example to generalise from, no pattern to compress it into. So if the trained model finds that string unusually unsurprising, that can only be because the parameters retain something specific to it. Storage is the only mechanism available. The practical consequence is the one that surprises reviewers: a structureless literal seen three times can be more recoverable than an ordinary comment seen ten times. Count and entropy trade against each other, and secrets are exactly the content that sits at the extreme of the second axis. ## What sets the ceiling Model capacity. A larger fine-tune over the same corpus can hold more, so the recoverable set grows with parameter count. This is a fact about the training run, not about the deployment, and it is one of the few levers an owner actually has — alongside the corpus itself. ## What the adversary controls, and what they do not The attacker in this picture has nothing special: prompt in, continuation out, no weights, no gradients, no returned log-probabilities. What they spend is queries and the specificity of the prefixes they choose. A prefix that genuinely preceded the span in the corpus pulls hardest, which is why someone who uses the product — who knows the house conventions, the file layout, the naming style — is a stronger adversary than a stranger at the same query budget. But that budget only changes the harvest rate. It samples a set whose membership was decided when the run ended. Ten times the queries recovers more of the same set sooner; it does not cause an unstored string to appear. This asymmetry is the whole reason the defensive question is a corpus question. ## Reading a memorization audit The attached table is the shape such a result comes in, and the row worth arguing about is the fourth. Prose seen roughly eight times is recovered three times in forty; a high-entropy literal seen three times is recovered seven times in forty. A reviewer who ranks rows by copy count alone reads that table backwards. Note also what the last row does not establish: one recovery out of forty single-occurrence literals is a small number, not zero, and the sampling was over spans somebody chose to test. ## Where this leaves a defender You cannot enumerate the recoverable set, but you can characterise its inputs, and both are auditable without touching the model: which classes of string entered the corpus, and how often each class was duplicated. That is why the mitigation conversation is always about the training data and never about the endpoint. Metering the API, hiding scores, limiting query volume — all of these slow a harvest of a set that remains exactly as full as it was. ## The summary an interviewer is listening for Duplication and entropy, both fixed at training; capacity sets the ceiling; the attacker's budget sets the rate, not the contents; ordinary prose paraphrases and does not return.

  • Does a bigger model change what is recoverable?
    Yes, in one direction: more capacity raises the ceiling on how much can be retained, so a larger fine-tune over the same corpus generally puts more spans within reach. It changes the size of the set, not who controls it. The corpus still decides which strings are candidates, and the adversary still cannot add one.
  • If an adversary spends ten times the queries, what do they get?
    More of the same set, sooner. Extra queries and better-chosen prefixes raise the harvest rate against spans that were already recoverable; they do not make an unstored string emerge. That is why query volume is a poor thing to bound your exposure on, and corpus contents are a good one.
  • Why can a span duplicated three times beat prose that appeared ten times?
    Because entropy counts as much as copies. Ordinary prose is reconstructed approximately by a general rule, so no specific instance needs keeping; a random-looking literal has no rule behind it, so any low loss the model reaches on it came from retaining it. A few copies of a structureless string can be enough.

saying these in an interview costs you the question

  • Assumes only rows the model overfit on are recoverable
  • Thinks the whole corpus is extractable with enough queries
  • Ranks exposure by how sensitive a string is rather than how it was trained on
  • Believes a short high-entropy secret is too rare to be retained
  • Treats the recoverable set as something an attacker's budget expands

context