skip to content

What is a canary memorization result still worth after the model has been fine-tuned again next quarter?

level: principalimportance: nice to knowfreq 27%

answer

  1. the artefact you measured no longer exists
  2. three expiries, one is not yours
  3. a routine, not a number
  4. hold shape and rate constant
  5. sign the scoped sentence, refuse short

basics

~20 s

Very little on its own: the result described one artefact, one corpus and one extraction technique. After a new fine-tune it describes a model that no longer exists, so the deliverable is a standing re-measurement.

solid answer

~40 s

Treat the canary run as perishable evidence and decide up front what makes it stale. Three things expire it: a new fine-tune or corpus change, because the model measured is gone; a change to the decoding path or to what the interface returns, because the extractor's vantage moved; and a better extraction technique than the one that produced the number, because the negative was always relative to the effort spent. The deliverable is therefore a routine, not a number - canaries planted in every training corpus by default, a fixed ladder of insertion rates so runs are comparable, re-measurement bound to release, and a named owner. The judgment is what you sign between runs. A scoped sentence ages honestly; "our model does not leak training data" does not.

go deeper

for a junior

Take away one thing: the result belongs to the model version that was tested. A model retrained next quarter has not inherited last quarter's clean run.

for a middle

Be able to list what invalidates a result — new training, a changed interface or decoding path, and a stronger extraction technique — and explain why the last one needs no change on your side.

for a senior

Show how you would keep results comparable: fixed shapes and a fixed ladder of insertion rates, re-measured per release, with the copies confirmed present in the consumed data each time.

for a principal

Own the wording and the spend. Decide what the organisation signs between runs, refuse the unbounded sentence, and know when further extraction budget is worth less than fixing the corpus.

## The result has an expiry date, and somebody has to name it A canary run produces a measurement of one artefact under one set of conditions. The moment any of those conditions moves, the number describes something that is no longer in production. A lead's job here is not to run the experiment better; it is to decide what the organisation is allowed to say on the strength of it, for how long, and what it funds to keep the claim alive. Three kinds of change expire a result, and they expire it for different reasons: **The artefact changed.** A new fine-tune, more data, more steps, a different base model. Memorization is settled during training, so a retrained model is simply a different subject. The prior number transfers no better than a load test of last quarter's binary. **The vantage changed.** The extractor's access is part of the threat model. Exposing per-token scores where only text was returned, adding a longer context, changing sampling defaults, or shipping weights to a partner all widen what an extractor can do — and a result gathered under the old access does not cover the new one. **The effort changed.** A negative is always relative to the extraction that was run. Technique in this area improves, and a string that resisted one method can fall to a better one against the very same weights. This is the expiry people forget, because nothing in your system changed at all. ## What the standing arrangement looks like The useful output of thinking this through is not a report but a routine. - **Canaries in every training corpus by default.** They cost effectively nothing to insert and they are worthless if you only think of them after the corpus is frozen. Retrofitting is impossible — you cannot plant evidence in a run that has already happened. - **A fixed ladder of insertion rates.** Holding shapes and rates constant across runs is what turns isolated numbers into a trend. A trend is the thing that actually informs a decision: recall appearing at a lower repetition count than last quarter is a signal; a single absolute number is not. - **Re-measurement tied to release, not to the calendar.** Every model that ships gets measured; a quarterly cadence measures whatever happened to be deployed on the day. - **A named owner and a re-read on technique.** Somebody has to notice that the extraction method used last time is no longer the best available, and re-run against the unchanged model. ## The claim you sign, and the one you refuse This is where the judgment sits, because the pressure is always toward a shorter sentence. A stakeholder wants "our model does not leak customer data." That sentence is unbounded in time, in shape, in rate and in technique, and every one of those bounds is real. Signing it means being wrong later, publicly, with a document that says you said it. The defensible form carries its scope: strings of these shapes, at these insertion rates, were not recovered from this model version through this interface under this extraction effort, as of this date. It is longer and less satisfying, and it is the only version that survives the next fine-tune. There is a further honesty worth stating out loud: a measurement is not a guarantee. A canary run is empirical evidence about what one attempt found, and it belongs in a different category from a formal bound on what any attempt could find. When someone needs an assurance that does not expire when technique improves, a measurement is the wrong instrument and no amount of re-running fixes that. ## What you fund, and what you decline Re-measurement is cheap relative to almost anything else in this area — inserting strings costs nothing, and the extraction run is bounded compute you choose the size of. That makes it easy to defend, and it also makes it easy to over-buy: an ever-larger attempt budget produces a negative that is marginally less narrow and still narrow. Past a point, the money is better spent on the corpus side, where deduplicating and filtering the classes of string you found recoverable changes the outcome rather than the measurement of it. The decision to escalate belongs to the ladder, not to the calendar. If recall shows up at rates that plausibly occur in your real data — a template repeated a few hundred times, say — that is a corpus problem to fix before the next run, and the measurement has done its job. If it only shows up at rates far above anything your corpus contains, the sensible call is to keep the routine cheap and spend the attention elsewhere. ## The thing not to promise Never let the canary become the answer to "is that customer's record gone." It measures recovery of a string you planted, under one attempt. It is evidence about a class of behaviour, not an accounting of any individual's data, and a lead who lets those two be conflated in a document has created a liability that the experiment was never able to cover.

  • Which expiry is easiest to miss, and why?
    The improvement in extraction technique, because nothing in your own system changed. The model, the corpus and the interface are all the same, yet the negative result was always relative to the effort that was spent looking. Somebody has to own re-running against an unchanged model when a better method becomes available, or the number silently decays into a claim nobody rechecked.
  • A stakeholder wants one sentence for a customer questionnaire. What do you give them?
    A scoped one: strings of the tested shapes, at the tested insertion rates, were not recovered from the named model version through the deployed interface under a stated extraction effort, as of a date. Refuse the unbounded version. The short sentence is the one that gets quoted back after the next fine-tune, and it will not be true then.
  • When is more extraction budget the wrong thing to fund?
    Once the ladder has already shown where recall appears. Beyond that, a bigger attempt budget buys a marginally less narrow negative and changes nothing about the model. The money moves to the corpus side — deduplicating and filtering the classes of string the ladder showed to be recoverable — because that changes the outcome rather than the measurement of it.

saying these in an interview costs you the question

  • Treats one canary run as a standing assurance
  • Runs canaries on a calendar, not per release
  • Changes shapes and rates between runs, losing the trend
  • Ignores that better extraction techniques expire a negative
  • Signs an unbounded no-leakage statement
  • Offers a canary result as proof a record is gone
  • Plans the canary after the corpus is already frozen

context