skip to content

A red-team run confirms 41 verbatim spans out of a million candidates — what can the report claim?

level: seniorimportance: nice to knowfreq 27%

answer

  1. Existence, never extent
  2. Recall has no denominator here
  3. Not found is not shown absent
  4. The pool size is a cost, not a finding
  5. State matched length and threshold

basics

~20 s

It establishes that those 41 spans were in the training data, given verification against the genuine source and spans long enough that coincidence is implausible. It says nothing about how much more is recoverable, because the run measured precision and never measured recall.

solid answer

~50 s

The claim is existence, not extent. Verified spans, matched against the genuine artefact and long enough that a coincidental match is implausible, establish that this text was in the training corpus and can be elicited. That is a strong positive finding. What the run cannot support is any statement of proportion: it sampled a generation space it did not exhaust and shortlisted at a threshold it chose, so recall has no denominator and a document with no confirmed hit was not shown absent. The reportable numbers are the confirmed count, the span lengths, the threshold and the reference model used, and the precision on the sample that was actually verified. The million candidates are an operational cost, not a result, and quoting them as though they were findings is what gets an extraction report dismissed.

code

text · 9 lines
text
extraction run summary
  candidate spans generated            1,204,000
  flagged by score filter                 18,400   threshold: target vs reference confidence gap
  reference model                        disjoint corpus, comparable capability
  sampled for human verification           1,000
  confirmed verbatim vs archive               41   min matched length: 50 tokens
  precision on the verified sample           4.1%
  recall                                unmeasured  (no denominator)
  ...

go deeper

for a junior

Know that confirmed verbatim spans prove those spans were in the training data and prove nothing about how much else is. Volume of generated candidates is not a result.

for a middle

Explain why recall is unmeasurable here and why matched span length matters: short common phrases match by coincidence, long distinctive ones do not. Be able to name the numbers a report should carry.

for a senior

Show you can write the finding so it survives challenge: verified fact stated as fact, precision reported with its threshold and reference, recall explicitly declined, and no proportion implied anywhere. Handle the pushback that 41 out of a million sounds negligible.

for a principal

Own how this result is communicated outward, including who verified it and with what access, and set the standard that neither overclaiming a rate nor burying verified hits is acceptable in work that feeds disclosure or licensing decisions.

## Reading the result correctly An extraction run against a large generative model produced a huge candidate pool, a score-filtered shortlist, and a small set of spans confirmed by matching them against the genuine archive. Somebody now has to write down what that means, and almost every way of writing it is wrong in one of two directions: overclaiming a leak rate, or dismissing verified hits as anecdotes. ## What the confirmed hits do establish A span matched character for character against the genuine source, long enough and distinctive enough that coincidental reproduction is implausible, establishes two things: 1. That text was present in the training corpus. 2. It can be elicited from the deployed model by someone with ordinary generation access. Both are strong. The first is a claim about the corpus, which nobody can now edit, since what was memorized was settled at training time. The second is a claim about the live product. Neither depends on the size of the candidate pool. Span length and entropy carry the weight in the first claim. A short, common phrase matching the archive establishes nothing, because such a phrase appears in countless other places and would be produced by any fluent model. A long, distinctive passage matching character for character is not something a fluent model produces by chance. Any serious report states the matched length and why coincidence was ruled out. ## What the run cannot support **A proportion of the archive.** Recall has no denominator here. The run neither exhausted the generation space nor enumerated the archive's presence in the corpus, so it cannot say what share is recoverable. Forty-one confirmed spans is forty-one confirmed spans. **Absence for anything not found.** A document with no confirmed hit was not found by this run, at this threshold, with this reference model. That is a statement about the run, not about the model. **Reproducibility as a rate.** A different run with a different threshold and different candidates will confirm a different set. The confirmed spans do not constitute an expected yield. **A claim that the model stores the archive.** Verbatim recall of specific spans is not storage of a corpus, and phrasing it as storage invites a rebuttal that discredits the genuine finding along with the overclaim. ## Which numbers belong in the report The honest headline is precision with its provenance attached: how many candidates were verified, how many were confirmed, at what threshold, against which reference, and with what minimum matched span length. The candidate pool size belongs in the method section as an operational cost, because it is a statement about how much text was generated and nothing else. A report whose headline number is the size of the pool is describing its own expenditure. It is also worth reporting who performed the verification and with what access, because a claim confirmed against the genuine archive by someone holding it is categorically stronger than a ranked shortlist produced by someone who does not. ## Why the direction matters commercially and legally These findings are read by people making decisions about disclosure, licensing and remediation, and both errors are expensive. Overclaiming — quoting an implied leak rate the run never measured — is the error that gets the whole report set aside once someone asks for the denominator. Underclaiming — burying verified hits under a caveat about small sample size — throws away the part of the finding that is genuinely solid, because a single verified verbatim span is a fact about the corpus regardless of how many candidates it took to find. The discipline to hold is simple to state: report what was verified as fact, report the filter's precision as a measured quantity with its threshold, and refuse to state anything about recall. The evidentiary value of the run is its precision, never its volume.

  • One of the 41 confirmed spans is a six-word phrase. Would you report it?
    Not as a hit. A short common phrase matching the archive establishes nothing, because it appears in many other sources and any fluent model produces it. Reports should set a minimum matched length and distinctiveness bar and state it, so a reader can judge why coincidence was ruled out. Including short matches inflates the count and hands a reviewer the easiest possible way to dismiss the whole set, including the long spans that are genuinely solid.
  • The owner responds that 41 spans out of a million candidates is a negligible rate. What do you say?
    That the ratio is not a rate of anything. The denominator is how much text the run generated, which is a measure of effort spent, not of how much of the archive is exposed. The finding is that specific archive text was in the training corpus and can be elicited from the live model; the run measured precision and explicitly did not measure recall, so no proportion is available to argue over in either direction.
  • Can you tell the owner which documents are safe?
    No. A document with no confirmed hit was not found by this run at this threshold with this reference, which is a statement about the run's coverage rather than about the model. Extraction runs produce positive evidence only; there is no negative result to hand over. If somebody needs a bounded statement about a specific string, that requires a different instrument designed to measure a known target, not a broad extraction campaign.

saying these in an interview costs you the question

  • Converts confirmed hits into an implied leak rate
  • Reads no confirmed hit as proof of absence
  • Headlines the candidate pool size as the finding
  • Claims the model stores a copy of the archive
  • Omits matched span length and the threshold used
  • Dismisses verified spans as too few to matter

context