skip to content

What is a canary planted in a training corpus a known number of times, and why does its owner then try to extract it?

level: juniorimportance: should knowfreq 44%

answer

  1. a measurement, not an argument
  2. the owner goes first, on purpose
  3. a string nobody could have guessed
  4. inserted a counted number of times
  5. then queried back out, weights unused

basics

~20 s

A canary is a known random string the owner inserts into the training corpus a counted number of times, then tries to pull back out of the trained model using only the access an outsider has.

solid answer

~50 s

The owner of a model that was fine-tuned on sensitive text wants to know whether that text can be recovered by someone who can only send the model inputs and read its outputs. Rather than argue about it, they plant evidence: a random, high-entropy string of a chosen shape, inserted into the corpus a recorded number of times before training. After training, they sit in the extractor's seat deliberately — generation and per-token scores only, no weights — and see whether the model reproduces the canary, or scores it far more confidently than same-shaped strings it never saw. If it comes back, memorization of that string at that insertion count is demonstrated, and the only secret leaked is a fake one the owner minted. If it does not, they have one negative result about one shape at one rate, not a clean bill of health.

go deeper

for a junior

Be ready to say plainly what is planted, where, how many times, and how it is read back. Know that the string must be random enough that the model could not have produced it by chance.

for a middle

Explain why the owner deliberately restricts themselves to query access, and how a comparative reading works: the model preferring the exact inserted string over same-shaped strings it never saw.

for a senior

Show that you record shape, insertion count and extraction effort before the run, and that you check the copies survived the corpus pipeline. Be able to state what a clean result does not cover.

for a principal

Own the framing for people outside the team: this is evidence about one experiment, not an assurance about a corpus, and saying so plainly is what keeps the artefact trustworthy.

## The problem the canary is built for A model fine-tuned on real text — historical customer-support transcripts, say — has fitted that text into its weights. The question a reviewer eventually has to answer is whether an outside party, holding nothing but the ability to send the model inputs and read what comes back, can recover any of the original strings word for word. This is an empirical question about a specific model, and it is very hard to settle by reasoning. "We deduplicated" and "the model does not store text" are both arguments, not evidence, and neither is checkable by the person who has to sign off. A canary replaces the argument with an experiment. Before training, the owner writes a known string into the corpus a known number of times. After training, the owner runs the extraction an outsider could run and records the result. ## What the string has to be The canary must be something the model could not produce for any reason other than having been trained on it. That means high entropy: drawn at random from a space so large that guessing it is out of the question. A memorable phrase, a plausible-looking address, or anything the model's ordinary language competence could generate is useless as a canary — recovering it would prove nothing, because the model would have produced it anyway. It is also, by construction, worthless. The owner minted it. If the experiment succeeds and the string escapes, no real person is harmed. That is one of the design's quiet virtues: it is the only way to measure the leakage of a secret without risking a real one. ## What the owner records before training Two things, and both are load-bearing for reading the result later: - **The shape** — the format and the number of random bits. Memorization is not uniform across formats; a long random token and a short numeric code behave differently. - **The insertion rate** — how many times the string appears in the corpus. One occurrence, ten, a hundred. This is the single most important column, because how strongly a model reproduces a string rises sharply with how often it saw it. A canary the corpus pipeline silently strips — by deduplication, by a filter, by a length rule — measures nothing, so the owner also confirms the inserted copies actually survived into the training data. ## The seat the owner sits in The measurement is only worth something if it uses the access the outside party would have. Typically that is generation from the deployed decoding path, plus whatever per-token scores the interface exposes, and no weights. If the owner instead reads the answer out of the weights or out of the training log, they have measured nothing about what a querying outsider can do. The result is read one of two ways. The strong version is direct: the model emits the canary. The more common and more sensitive version is comparative: the model assigns the exact inserted string a far better score than same-shaped strings drawn from the same random space that were never in the corpus. A model that has never seen any of them should have no reason to prefer one. ## What the number is, and is not A positive result is strong evidence: it demonstrates that a string of that shape, at that insertion count, is recoverable through query access alone. That is a finding you can act on — deduplicate harder, filter that class of string out of the corpus, change what goes into the fine-tune at all. A negative result is much weaker than it feels. It bounds what was tested: that shape, that insertion rate, that extraction effort, that model version, that decoding path. It says nothing about a differently-shaped secret, nothing about a string that appeared three hundred times instead of once, and nothing about a better extraction technique than the one that was run. The failure mode this whole exercise most often produces in practice is a confident sentence — "our canary was not extractable, so the model does not memorize" — that the experiment does not support. ## What it is not Three different things in security carry the word *canary*, and mixing them up is a common stumble. A decoy document or account that alerts when someone touches it is a deception artefact — it reports intrusion. A synthetic event injected to make a detection rule fire is a heartbeat — it reports pipeline health. Neither involves training. The canary here is written into a corpus, learned by a model, and measured back out of the model's behaviour; the artefact under test is the weights. It is also not an evaluation example. An eval item is held out precisely so the model does not see it, and it measures accuracy. A canary is put in on purpose, and it measures recall of the training data itself.

  • Why does the owner run the extraction with query access only, when they hold the weights?
    Because the claim being tested is about what an outsider can do. Reading the string out of the weights or the training log answers a different question. The measurement is only comparable to a real extraction if it uses the same vantage a real extractor has: generation from the deployed decoding path and whatever scores the interface returns.
  • Why is a fake secret used instead of checking whether a real customer string comes back?
    Two reasons. A real string may already appear in public text, so recovering it proves nothing about training exposure. And running an extraction against a genuine secret means that if it works, you have just demonstrated a live disclosure. A minted, high-entropy canary is unguessable and worthless, so a successful extraction costs nothing but tells you the same thing.
  • What would make a canary run silently meaningless?
    The inserted copies never reaching training. Corpus pipelines deduplicate, filter and truncate, and a strange random string is exactly what such a stage removes. Before reading any result, confirm the intended number of occurrences was present in the data the run actually consumed; otherwise a clean result just records that the canary was deleted.

Like dropping one numbered marker into a river to see whether anything from upstream reaches the sea. Finding it proves passage; not finding it only rules out that marker on that day.

saying these in an interview costs you the question

  • Calls it a decoy file that alerts when opened
  • Treats the canary as a held-out evaluation item
  • Uses a memorable phrase the model could generate anyway
  • Reads the canary out of the weights, not by querying
  • Reports a negative result as proof of no memorization
  • Never records how many times it was inserted

context