skip to content

A canary inserted once did not come back out under a fixed extraction budget — what does that result prove?

level: seniorimportance: must knowfreq 50%

answer

  1. one shape, one rate, one effort
  2. which column was one?
  3. what actually leaks is repeated
  4. a hard case answered, not the common case
  5. did the copies even reach training?

basics

~10 s

A clean canary result proves only that a string of that shape, inserted that many times, resisted that extraction attempt against that model version. It is not evidence that the model does not memorize.

solid answer

~40 s

A clean canary result bounds four things, and each is narrow: the shape that was planted, the insertion rate that was used, the extraction effort that was spent, and the exact model and decoding path tested. The most important is the rate. Verbatim recall rises sharply with how often a string appears in training, and what actually escapes a real corpus is usually the repeated material - a signature block, a support macro pasted into thousands of transcripts. A canary inserted once measures the hardest case, so a negative tells you little about the repeated ones. The honest write-up states shape, rate, attempt budget and model version, and says what was not covered. "Our canary was not extractable, so the model does not memorize" is the sentence to strike.

code

text · 10 lines
text
memorization check - support assistant, fine-tune run 2026-Q1
  canary shape        : fixed label + 12 random alphanumerics (~62 bits)
  insertions          : 1   (corpus: 4.1M transcript turns)
  present in consumed data? : yes, 1 occurrence confirmed
  access used         : generation + per-token scores, no weights
  extraction attempts : 50,000 candidate completions
  result              : NOT recovered; rank indistinguishable from
                        unseen strings of the same shape
  ...
  conclusion as filed : "the model does not memorize training data"

go deeper

for a junior

Learn the asymmetry first: a canary that comes back proves something specific, and a canary that does not only rules out what was tried. Absence of a result is not absence of memorization.

for a middle

Explain why insertion rate dominates. Recall climbs steeply with repetition, so a single-insertion canary tests the hardest case while the strings that really leak are the repeated ones.

for a senior

Demonstrate that you check the canary survived the corpus pipeline, record shape, rate, attempt budget and model version, and write the conclusion in those same terms rather than as a general claim.

for a principal

Own the wording that leaves the team. A bounded finding restated as a clean bill of health is how an honest measurement turns into a false assurance nobody can retract later.

## Reading a negative result correctly The result of a canary run is an experiment, and like any experiment its scope is exactly its conditions. A canary that came back is strong evidence about the model. A canary that did not is a bounded statement, and the difference between what it bounds and what people hear is where the damage happens. Four conditions frame every negative result: **The shape.** The format and the entropy of the planted string. Verbatim recall is mediated by tokenization and by the frame around a string; a bare random blob and a random value embedded in a recognisable field are different objects. A result covers the shape that was tested. **The insertion rate.** How many times the canary appeared in the corpus. This is the column that dominates everything else. **The extraction effort.** How many attempts, through which interface, using which technique. A negative under fifty thousand attempts is not a negative under fifty million, and it is certainly not a negative under a technique better than the one that was run. **The artefact.** Which model version, which decoding path, which serving configuration. A model fine-tuned again next month is a different model, and recall may sit differently in it. ## Why the rate is the column that matters most A single occurrence is the hardest thing for a model to memorize. Recall of a string climbs steeply with repetition, because each additional occurrence pushes the weights further toward reproducing it. This is not a marginal effect — it is the dominant driver of which strings come back out. Now look at what a real corpus contains. The strings that leak in practice are rarely the once-mentioned oddity. They are the repeated ones: a template that carries an internal endpoint, a support macro pasted into thousands of chats, a signature block appended to every message from one team, a key that was quoted in a hundred escalations. Those appear many times, and a canary inserted once is a poor proxy for any of them. So a run with a single insertion measures the corner of the space where memorization is weakest, and reports back that memorization was weak there. That is a true statement and a nearly useless one for the question people actually want answered. The fix is not complicated: plant a ladder. Several canaries at several rates — one, ten, a hundred, a thousand — and read where recall starts to appear. That result has a shape you can reason with: it tells you roughly how many repetitions of a string of that form it takes before it becomes recoverable, which is a claim you can hold against what the corpus actually contains. ## The other silent failure Before reading any negative, confirm the canary reached training. Corpus pipelines deduplicate, filter, truncate and reject, and a strange high-entropy string is exactly the kind of thing a cleaning stage removes. A run where the intended copies were stripped produces a beautiful clean result that measures the pipeline, not the model. Counting the occurrences in the data the run actually consumed — not in the file that was handed to it — closes this. ## What is legitimately covered It is worth being precise about the true statements a negative supports, because there are some, and a reviewer who says "this proves nothing" is overcorrecting: - Strings of the tested shape, at the tested rate, resisted the extraction that was run. - The technique used did not recover the string within its attempt budget. - That holds for the tested model version and decoding path. What it does not support: any claim about differently-shaped secrets, any claim about repeated strings, any claim about a stronger technique, any claim about the pretraining corpus you never controlled, and any claim of the form "customer data cannot be extracted." ## Fine-tuning is not the whole model One more scope limit gets skipped constantly. A canary planted in a fine-tuning corpus measures memorization of the fine-tuning corpus. If the base model was pretrained on data you did not control, nothing about your canary speaks to what is recoverable from that. A negative on your own inserted string and a leak from the base model's training data are entirely compatible. ## How to write it up State the shape, the entropy, the insertion counts, the extraction interface and attempt budget, the model version, and the confirmation that the copies were present in the consumed data. Then state the conclusion in the same terms: not "no memorization", but "strings of this shape at these rates were not recovered under this effort against this version." That sentence survives contact with a sceptical reviewer. The shorter one does not.

  • How would you redesign the run so the result answers the question people are actually asking?
    Plant a ladder rather than a single canary: several strings of the same shape at several insertion rates — one, ten, a hundred, a thousand — and read where recall first appears. That yields a threshold in repetitions, which you can compare against how often real templates and boilerplate actually occur in the corpus. Add a second shape if the secrets you care about are structured differently.
  • The report says the canary was inserted once. What do you check before believing the negative at all?
    That the copy survived into the data the training run consumed. Deduplication, filtering, truncation and length rules routinely strip an odd high-entropy string, and then the run measures the pipeline instead of the model. Count occurrences in the consumed dataset, not in the file that was handed to the pipeline.
  • Does a clean canary result say anything about the base model the assistant was fine-tuned from?
    Nothing at all. A canary planted in the fine-tuning corpus measures recall of the fine-tuning corpus. Whatever the base model absorbed during pretraining on data you never controlled is untouched by the experiment, and a clean fine-tune result sits perfectly comfortably alongside verbatim recall from pretraining data.
  • Is a positive result read with the same caution as a negative one?
    No, and that asymmetry is the point. A recovered canary demonstrates that a string of that shape at that rate is extractable through query access alone — a direct, actionable finding. A negative only rules out what was tested. Evidence of presence is strong; evidence of absence is bounded by the effort spent looking.

A metal detector that swept for coins and found none has told you about coins, not about the field.

saying these in an interview costs you the question

  • Reads a clean canary as proof of no memorization
  • Never states the insertion rate with the result
  • Ignores that the extraction budget was finite
  • Assumes one shape covers all secret formats
  • Forgets to check the canary reached training
  • Extends a fine-tune result to the base model
  • Treats a negative as strongly as a positive

context