Why must a canary planted in a training corpus be a high-entropy random string, and how is its recovery scored?
answer
- the model must have no other route
- there has to be a control group
- compare against strings it never saw
- a rank, not a raw score
- the space size makes it mean something
basics
~20 sThe planted string must be one the model could not produce except by having been trained on it. Recovery is scored comparatively: how far the model ranks the exact inserted string above same-shaped strings it never saw.
solid answer
~40 sA canary is a controlled stimulus, so it has to be unguessable. If the planted string is a natural phrase or a plausible identifier, the model can emit it without ever having seen it, and the measurement has no control group. Drawing it at random from a large space fixes that: the only route to preferring that exact string is training exposure. Scoring then does not need the model to spit the canary out unprompted, which is a coarse test. The sensitive reading is comparative - take alternatives with the identical format, drawn from the same random space, none of which were in the corpus, and ask how far the model ranks the real canary above them. A rank-based version of that gap, usually called exposure in this literature, is what gets reported.
go deeper
Remember the core requirement: the planted string has to be random enough that the model could not have produced it without training on it. A guessable canary measures nothing.
Explain the comparative reading: same-shaped strings drawn from the same random space, none in the corpus, and the canary's rank among them. Say why a rank travels better than a raw score.
Show you would check the alternative set for contamination and match the canary's format to the secrets you actually care about, then state which format the result covers.
Be able to say what the instrument is good for. High entropy buys a trustworthy positive; it buys nothing for a negative, and confusing the two is how a measurement becomes a false assurance.
## Entropy is what makes the experiment a control The canary experiment asks one question: can this exact string be recovered from the model by someone who only queries it? For the answer to mean anything, there must be exactly one explanation for a positive result — that the model was trained on the string. Any other route to producing it destroys the experiment. That is what entropy buys. If the planted string is drawn uniformly at random from a very large space, then a model that never saw it has no basis for preferring it over any other member of that space. A well-formed random identifier of a few dozen bits is already far outside anything a model could arrive at by chance or by fluency. The opposite choice fails silently, which is why it is worth naming. A canary like a common-looking name, a round number, or a sentence in ordinary prose can be produced by a model that never saw it, because producing plausible language is exactly what the model does. Recovering such a string is not evidence of memorization, and — worse — it is evidence that reads as alarming. A team can spend a week responding to a leak that never happened. ## Two ways to read the result, and one of them is much better The crude reading is generation: query the model in the ways an outsider would and see whether the canary appears in what comes back. This is a real test and a positive is meaningful, but it is coarse. A string can be strongly fitted into the weights and still fail to surface, because whether it surfaces depends on how the model is prompted, how the deployed decoding path samples, and how many attempts are made. The sensitive reading is comparative and uses the scores the interface returns. Construct a set of alternatives with exactly the canary's format, drawn from the same random space, none of which are in the corpus. Ask the model how confidently it assigns each one. A model that saw none of them should treat them as interchangeable — they are all equally arbitrary. If it assigns the one true canary a far better score than the others, that gap has one cause. Reporting that as a rank rather than a raw score is what makes it comparable across runs. Saying "the model rated the canary above all but a vanishing fraction of same-shaped alternatives" is a statement whose meaning does not shift when the model, the tokenizer or the score scale changes. The rank-based form of this measure is usually called exposure in the literature. The size of the random space is what makes that rank interpretable at all. A rank near the top of a space of a hundred candidates is barely a signal; the same rank in a space of billions is not something an unexposed model produces. ## Format matters as much as randomness Two strings can carry the same number of random bits and behave completely differently. Memorization is mediated by how the string is broken into tokens and by how distinctive its surrounding context is. A random string embedded in a recognisable frame — a fixed prefix, a stable field label — behaves differently from a bare random blob, because the frame gives an extractor a place to start and gives the model a stable cue. That is a real limitation on what any single canary measures, and it is the reason the shape has to be recorded alongside the result. A run that measured one format has measured one format. If the secrets you actually care about have a different structure — free-text notes, numbers, structured records — a canary shaped unlike them is a poor proxy for them. ## What the choice of alternatives can quietly break The comparative reading depends on the alternatives being genuinely unseen and genuinely same-shaped. Two things spoil it. If the alternative set is generated in a way that makes some members more natural-looking than others, the model's preference among them reflects fluency rather than exposure. And if any alternative happens to occur in the corpus — which is easy when the space is small or the format is common — the comparison is against a contaminated baseline and the measured gap shrinks for the wrong reason. ## Where the entropy argument stops helping High entropy makes a positive result trustworthy. It does nothing for the interpretation of a negative one. A model can fail to prefer a fifty-bit random string inserted once while reliably reproducing a short, structured, frequently repeated string that carries far less entropy — because repetition, not entropy, is the dominant driver of verbatim recall. The property that makes a canary a clean instrument is not the property that makes real data leak, and keeping those two straight is most of the skill in reading these reports.
- Why report a rank against alternatives rather than the model's raw confidence in the canary?Raw scores are not comparable across models, tokenizers or score scales, and a model can be broadly confident or broadly diffuse for reasons unrelated to memorization. A rank against same-shaped strings the model never saw removes that baseline: it asks only whether this one string stands out from its own random space, which is the thing exposure would cause.
- What goes wrong if some of your alternative strings happen to appear in the corpus?The baseline is contaminated. The comparison assumes the alternatives were unseen, so if a few were in the training data the model rates them highly too, the measured gap shrinks, and the run under-reports memorization. With a small random space or a common format that happens easily, which is another reason the space has to be large.
- Does a higher-entropy canary memorize less than a low-entropy one?Not in a way you can rely on. Entropy is a property of your instrument, not a defence. Repetition is the dominant driver of verbatim recall, so a short, low-entropy string that appears many times is typically far more recoverable than a long random one inserted once. Entropy makes a positive result interpretable; it does not predict what leaks.
saying these in an interview costs you the question
- Picks a memorable phrase as the canary
- Only checks whether the model emits it unprompted
- Compares confidence against unrelated text, not same-shaped strings
- Reports a raw score with no candidate space
- Assumes high entropy means the string is safe from recall
- Ignores that the canary's format shapes the result