skip to content

Verbatim Recall

Some training strings come back word for word, and which ones was settled at training time. Interviewers probe it because the answer is a property of a corpus that can no longer be edited.

on this pageshow

explore

questions

12

An attacker with only generated text extracts a fluent span resembling a real record — why is that not memorization?

level: juniorimportance: must knowfreq 62%

answer

  1. Fluency is not the same as memory
  2. Well-formed nonsense is cheap to produce
  3. Looking right establishes only the format
  4. Compare against something that never saw it
  5. Or check it against the real artefact

basics

~20 s

A generative model invents well-formed text on demand, so a realistic-looking span proves only that it knows the format. A memorization claim needs verification against the genuine source, or a confidence gap against a model that never saw the data.

solid answer

~50 s

Fluency is the model's job, so shape proves nothing. A model trained on huge amounts of text has learned what a record, an address or a licence line looks like, and it will emit well-formed instances of that shape whether or not it ever saw the specific one in front of you. Turning a candidate span into a claim takes one of two things: verification against the genuine artefact, where whoever holds the real record confirms the span matches it, or a confidence comparison, where the target scores that span far more confidently than a reference model that provably never saw the data. The first is proof; the second is evidence. That comparison is a membership test applied to a string the model itself produced, which is why extraction and membership inference are one family underneath. An unverified candidate list is worth nothing to whoever asked for the finding.

go deeper

for a junior

Be ready to say plainly that a generative model produces well-formed invented text on demand, so a realistic span proves format knowledge and nothing more. Name at least one way to test it: check against the real record, or compare confidence with a model that never saw the data.

for a middle

Explain why the comparison is needed at all: common and formulaic text is easy for every model, so raw confidence selects boilerplate. The signal is the gap between the target and an uninvolved reference, and that gap is a membership-style test applied to a generated string.

for a senior

Show the reporting discipline. Say what you would put in a finding, what verification you needed and who could perform it, and refuse to state anything about the rest of the corpus from a handful of hits. Interviewers listen for whether you separate evidence from proof.

for a principal

Own the framing given to legal or a data owner: what an extraction result can and cannot support as a claim, what verification access the exercise requires, and why a large candidate list is an operational cost rather than a result.

## The situation An adversary has ordinary generation access to a large text model behind a paid API. They do not hold its weights and they did not build its training corpus. They want to show that some specific body of text — a licensed archive, a customer table, a private document set — was in that corpus and can be pulled back out word for word. They generate large volumes of candidate continuations, and somewhere in the pile is a span that looks exactly like the thing they are hunting for. The question is what that span is worth as evidence. The honest answer is: on its own, almost nothing. ## Why a realistic span is weak evidence A generative model is trained to produce text that is *likely*, and likely text is well-formed text. Formats are among the easiest things for it to learn, because a format is repeated across enormous numbers of examples that are otherwise unrelated: the digit grouping of an account number, the punctuation of a postal address, the character classes and length of an API-style key, the cadence of a legal clause. Having learned the format, the model will produce novel, fluent, entirely invented instances of it at essentially zero cost, because generating is cheap. So a span that *looks right* has established one fact: the model can write text of that shape. That is not the fact anyone cares about. The fact people care about is whether this particular string came back out of the weights rather than being assembled on the spot. The two are visually indistinguishable, and confidently reading the first as the second is the standard failure of an inexperienced extraction report. The asymmetry matters for cost as well. Candidate generation is nearly free, so an adversary can always produce more candidates. What is not free is telling recall from invention. That discriminator is the entire attack; the generation is a prelude to it. ## The two ways a candidate becomes a claim **1. Verification against the genuine artefact.** Somebody who holds the real record checks whether the span matches it, character for character, over a length long enough that coincidence is implausible. This converts the candidate into a fact and it is the strongest thing available. Note who can do it: whoever holds the ground truth. An auditor, a publisher or the data owner can verify; an outside adversary who never had the archive frequently cannot, which is why the same technical pipeline produces a *finding* in one pair of hands and only a *suspicion* in another. **2. A confidence comparison against a reference.** The adversary scores the candidate under the target model and under a second model that provably never saw the data in question, and keeps only spans the target fits far better than the reference does. Text that is intrinsically easy — boilerplate, common phrasing, formulaic structure — is easy for both models and cancels out. Text that is easy for the target and hard for anything else is the interesting residue. This is evidence, not proof: it says the target's fit on that string is anomalous in the way training on it would produce. Notice what the second method actually is. It asks, of one specific string, whether the model behaves as though it had seen it. That is a membership question, asked about a string the model itself supplied. Extraction and membership inference are not neighbouring topics; the extraction pipeline has a membership test bolted to its output, and the test is the part that does the work. ## The direction of the claim Getting the direction right is most of the skill here. - A fluent, plausible span establishes that the model can produce that shape. - A span the target scores anomalously confidently, relative to a reference that never saw the data, is evidence consistent with recall. - A span verified against the genuine artefact establishes that this text was in the training data. - None of the three establishes how much *else* is recoverable, and a candidate the pipeline failed to confirm is not evidence that anything is absent. Whether a given span was memorized at all was settled during training and is a separate question from whether an adversary can elicit and prove it; this one is about the eliciting and the proving. ## What this means in practice A red-team report that lists thousands of suggestive candidate spans is not a finding, it is a pile of generated text. A report that lists a small number of independently verified spans, states the threshold and the reference used, and refuses to extrapolate to the rest of the corpus, is a finding somebody can act on. The evidentiary value of an extraction run is its precision, never its volume.

  • If the span contains a long random-looking string, does that change your answer?
    It raises the prior a little, because a model that invents high-entropy strings is unlikely to land on a specific real one by chance. But it still is not a claim: the model produces random-looking strings of the right shape readily, and only two things settle it — a match against the genuine artefact, or a confidence gap against a model that never saw the data. Length and entropy affect how implausible a coincidental match would be, which is why very short matches are never reported.
  • Who can actually verify an extracted span, and how does that change the finding?
    Only whoever holds the genuine source: the data owner, the publisher, or an auditor given access. An outside adversary without the archive can rank candidates by the confidence comparison but cannot confirm any of them, so their strongest honest output is a ranked hypothesis list. The same pipeline therefore yields a provable finding for an internal reviewer and an unprovable suspicion for an external one, which is worth saying explicitly in the report.
  • Why is this described as the same family as membership inference?
    Because the filter that makes extraction credible is a membership test. Membership inference asks whether a given record was in training by looking at how well the model fits it; extraction generates candidate records and then asks exactly that question of each one. The generation step supplies strings, the membership-style scoring supplies the evidence. Strip the scoring and you have a text generator producing plausible text, which is what it does anyway.

Someone who can imitate any handwriting can produce a convincing page in your hand on demand. That the page looks like yours proves they know your style; only comparing it against the original letter proves they copied one.

saying these in an interview costs you the question

  • Treats a realistic-looking output as proof of memorization
  • Says the model would not produce a valid-looking record unless it saw one
  • Reports unverified candidate spans as leaked data
  • Assumes rare-looking characters mean the string is real
  • Cannot name any way to distinguish recall from invention
  • Thinks extraction and membership inference are unrelated topics

context

open as a page

A code model emits a training string verbatim: why is 'it only generalizes' not an answer?

level: juniorimportance: must knowfreq 72%

basics

~20 s

Generalization and memorization happen in the same model. Training drives loss down over the corpus, and for a rare structureless string there is no pattern to generalize to, so storing it is the only way loss falls.

open as a page

A canary inserted once did not come back out under a fixed extraction budget — what does that result prove?

level: seniorimportance: must knowfreq 50%

basics

~10 s

A clean canary result proves only that a string of that shape, inserted that many times, resisted that extraction attempt against that model version. It is not evidence that the model does not memorize.

open as a page

What is a canary planted in a training corpus a known number of times, and why does its owner then try to extract it?

level: juniorimportance: should knowfreq 44%

basics

~20 s

A canary is a known random string the owner inserts into the training corpus a counted number of times, then tries to pull back out of the trained model using only the access an outsider has.

open as a page

In a training-data extraction attack, why score each candidate span against a reference model that never saw the corpus?

level: middleimportance: should knowfreq 44%

basics

~20 s

Ranking candidates by the target model's own confidence selects intrinsically likely text such as boilerplate. Comparing its score with an uninvolved reference model cancels what is easy for everyone, leaving spans the target fits unusually well.

open as a page

Why must a canary planted in a training corpus be a high-entropy random string, and how is its recovery scored?

level: middleimportance: should knowfreq 36%

basics

~20 s

The planted string must be one the model could not produce except by having been trained on it. Recovery is scored comparatively: how far the model ranks the exact inserted string above same-shaped strings it never saw.

open as a page

With only generation access to a fine-tuned code model, which training strings come back verbatim?

level: middleimportance: should knowfreq 55%

basics

~20 s

Two families: spans duplicated many times across the corpus, and rare high-entropy spans the model scores unusually confidently. Everything else comes back as paraphrase. The set is a property of the corpus, not of how hard the attacker tries.

open as a page

Candidate generation is nearly free for an extraction attacker — so what actually binds the run?

level: seniorimportance: should knowfreq 36%

basics

~20 s

The filter binds, not the generation. What limits the run is the false-positive rate at the chosen score threshold and the human verification effort each flagged candidate consumes, since only a holder of the genuine source can confirm a hit.

open as a page

After fine-tuning on a repo, you delete the file and rotate the key: what can still be extracted?

level: seniorimportance: should knowfreq 50%

basics

~20 s

The span itself, unchanged. Deleting the source alters the store, not the trained checkpoint, and rotation does not remove the string either — it makes recovering it worthless. Only a training-time action changes what the parameters hold.

open as a page

A red-team run confirms 41 verbatim spans out of a million candidates — what can the report claim?

level: seniorimportance: nice to knowfreq 27%

basics

~20 s

It establishes that those 41 spans were in the training data, given verification against the genuine source and spans long enough that coincidence is implausible. It says nothing about how much more is recoverable, because the run measured precision and never measured recall.

open as a page

What is a canary memorization result still worth after the model has been fine-tuned again next quarter?

level: principalimportance: nice to knowfreq 27%

basics

~20 s

Very little on its own: the result described one artefact, one corpus and one extraction technique. After a new fine-tune it describes a model that no longer exists, so the deliverable is a standing re-measurement.

open as a page

Should a checkpoint fine-tuned on an internal monorepo ship to customers, and on what evidence?

level: principalimportance: nice to knowfreq 34%

basics

~20 s

Decide on the corpus, not on your probing: your own extraction attempts bound only what you probed. The real inputs are which classes of sensitive string went in, which can be invalidated afterwards, and what a retrain costs.

open as a page