skip to content

After fine-tuning on a repo, you delete the file and rotate the key: what can still be extracted?

level: seniorimportance: should knowfreq 50%

answer

  1. three controls, only one touches the model
  2. the weights were frozen before you deleted anything
  3. filters only cover what you can list
  4. rotation changes worth, not possibility

basics

~20 s

The span itself, unchanged. Deleting the source alters the store, not the trained checkpoint, and rotation does not remove the string either — it makes recovering it worthless. Only a training-time action changes what the parameters hold.

solid answer

~50 s

Three controls get confused here and only one of them touches the model. Deleting the file from the repository changes what the *next* training run sees and nothing about the checkpoint you already shipped: the parameters were fixed when the run ended, and a user with ordinary generation access can still pull the span. An output filter helps only for strings you can enumerate and write down, which covers the credential you found and none of the tail behind it, and an exact-match rule is defeated by any formatting variant the model is equally happy to produce. Rotation is the control that actually reduces harm, and it works by changing the *value* of a recovery rather than its possibility — the string still comes back, it just no longer authorizes anything. Anything that genuinely empties the set is a training-time action, and it costs a run.

go deeper

for a junior

Remember that a trained checkpoint has no live link to the files it came from: deleting the source afterwards leaves the model exactly as it was.

for a middle

Explain what each control buys — store deletion, output filter, rotation, retrain — and why only the last changes what the parameters actually hold.

for a senior

Demonstrate the incident order: invalidate fast because it removes real harm, then scope which other classes went in, and never report a removal you cannot back.

for a principal

Own the decision on whether a retraining run is funded and the exact wording given to a customer or regulator, because 'no longer valid' and 'removed' are different claims with different consequences.

## Why this question separates people The instinct after finding a secret in a model's output is to fix the data: delete the file, rewrite the history, close the ticket. That instinct is correct for a store and wrong for a checkpoint, and the difference is that a checkpoint has no live relationship with the data it came from. Training read the corpus once, produced parameters, and stopped. Nothing you do to the corpus afterwards reaches back into those parameters. The recoverable set is what it was. ## The four controls, and what each actually buys **Removing the data from the store.** Buys everything for future training runs and nothing for the current model. This is worth doing immediately and worth being clear about: it is a fix for the *next* checkpoint, not for the one in production or, worse, the one a customer has already downloaded. **Filtering the output.** Buys coverage of exactly the strings you can list. That is genuinely useful for a known credential you have found, and it is structurally incapable of covering what you have not found — which is the entire tail and the whole reason the problem is hard. An exact-match rule also sits on a narrow ledge: the model can express the same content with different whitespace, different quoting, different chunking, and the filter has to have anticipated each. Treat it as harm reduction on known items, never as a boundary. **Rotating or invalidating.** Buys the most harm reduction per hour of anything on this list, and it is the correct first action in an incident. But be precise about the mechanism: the string is still in the weights and still comes back word for word. What changed is that it now refers to nothing. This distinction is exactly what gets lost when the incident is written up, and "we rotated it" quietly becomes "it was removed". **Retraining without the data, or training with a privacy mechanism from the start.** The only actions that change what the parameters contain. Both cost a training run and a delay, and the first also requires that you can identify what to exclude, which is a data-governance capability rather than a machine-learning one. ## Two things that look like removal and are not **Refusal or alignment training.** Teaching the model to decline requests that look like extraction raises the cost of eliciting a span. It does not change what is stored, and a model that refuses one framing will often complete the same continuation under a framing that looks entirely innocuous, because the underlying prediction was never altered. It is a rate control on the adversary, not a property of the weights. **Approximate unlearning.** Methods that claim to remove the influence of specific examples after the fact exist, and they are hard to verify: you can show that a particular probe no longer recovers a particular string, which bounds that probe. Reporting that as verified removal is the kind of claim that reads badly later, when someone recovers a variant. ## The classes rotation cannot save you from Rotation works because a credential's harm depends on it still being valid. That property is rarer than it looks. Customer names and account identifiers, personal data, unredacted incident narrative, internal path and topology structure, third-party licensed source — none of these can be invalidated. They stay recoverable and they stay meaningful. For those classes the honest options collapse to two: retrain without them, or do not ship the checkpoint. ## The order to work in Invalidate first, because it is fast and it removes real harm. Then scope: what else of that class was in the corpus, at what duplication, and which other spans of the same kind are therefore likely candidates. Then decide whether a retrain is funded, which is not a security decision alone. And throughout, keep two sentences apart in every artefact you write: *the credential no longer works* and *the string is no longer in the model*. Today only the first is true. ## What an interviewer is scoring Whether you know a checkpoint is a frozen artefact rather than a view over a database; whether you can rank controls by what they change rather than by how satisfying they feel; and whether you will state a residual honestly when the pressure in the room is to declare the incident closed.

  • What do you say when someone asks whether the secret has been removed from the model?
    That the credential no longer authorizes anything, and that the string itself remains recoverable from the checkpoint until a run without it replaces it. Those are two different statements and only the first is true today. Reporting removal on the strength of a rotation or a blocklist is the sentence that later reads as a misrepresentation.
  • Why doesn't refusal training close this?
    Because it raises the cost of elicitation without changing what is stored. A model taught to decline one framing will often complete the same span under another that looks innocuous, since the underlying continuation is unchanged. It is a rate control on the adversary, not a property of the weights, and it degrades as soon as someone phrases the request differently.
  • Which leaked classes cannot be neutralised the way a credential can?
    Anything whose harm does not depend on it still being valid: customer names and identifiers, personal data, unredacted incident text, internal topology detail, third-party licensed code. You cannot rotate a person's name. For those the real options are retraining without them or not shipping the checkpoint.

saying these in an interview costs you the question

  • Says deleting the training file removes the string from the model
  • Treats an exact-match output blocklist as complete coverage
  • Reports a rotated credential as removed from the checkpoint
  • Assumes refusal training empties the recoverable set
  • Presents approximate unlearning as verified removal

context