skip to content

Should a checkpoint fine-tuned on an internal monorepo ship to customers, and on what evidence?

level: principalimportance: nice to knowfreq 34%

answer

  1. you cannot prove absence
  2. probing bounds only what you probed
  3. decide on the corpus, not the model
  4. some classes cannot be rotated away
  5. somebody has to sign for the residue

basics

~20 s

Decide on the corpus, not on your probing: your own extraction attempts bound only what you probed. The real inputs are which classes of sensitive string went in, which can be invalidated afterwards, and what a retrain costs.

solid answer

~50 s

You cannot enumerate the recoverable set, so evidence from your own extraction attempts is a floor and never a clearance — finding nothing bounds the prefixes you tried at the effort you spent. The decision has to rest on the corpus. Ask which classes went in — credentials, customer identifiers, personal data, unredacted incident text, third-party licensed source — and split them by whether they can be neutralised after the fact. A credential can be invalidated; a customer's name cannot. That leaves three options somebody has to price: ship with everything invalidatable invalidated and the residue accepted in writing; re-run the fine-tune on a filtered corpus, which costs a run, a delay, and real utility because the filter must be over-inclusive; or keep the sensitive material out of weights entirely and behind an access-controlled layer. Then write the customer-facing claim to what you can defend.

go deeper

for a junior

Know that you cannot search a model for a string, so 'we looked and found nothing' is a far weaker statement than it sounds in a release meeting.

for a middle

Explain why a negative extraction result bounds only what was probed, and why the contents of the corpus are the input the decision should actually rest on.

for a senior

Lay out the options concretely — invalidate, refilter and retrain, or keep the material out of weights — with what each costs and what residual each leaves behind.

for a principal

Own the residual and the wording. Nobody can prove absence, so the call is which classes you accept, who funds the retrain and the quality it costs, and what claim you will defend a year from now.

## Why this is a judgment call and not a test result Every other release gate answers a question with a measurement. This one cannot, because there is no operation that searches a model for a string. You can probe, and probing is worth doing, but a negative result is evidence about your prefixes and your effort, not about the parameters. Customers will collectively spend far more effort than your team did, with better prefixes, because they use the product and know its conventions. So somebody has to sign for a residual that nobody can measure. That is what makes it a lead's decision rather than an engineer's. ## Move the decision to the input you can actually audit The corpus is enumerable; the recoverable set is not. So the review is a data review: - Which classes of sensitive content were in the fine-tuning data at all — credentials and tokens, customer names and identifiers, personal data, unredacted incident and support text, internal topology and path structure, third-party licensed source under terms you do not want to redistribute. - Roughly how duplicated each class was, since duplication is the strongest predictor of exact recovery. - Which classes can be neutralised after training and which cannot. That last split is the load-bearing one and it is usually shorter than people expect. Invalidation works only where the harm depends on the string still being live. Credentials, signing material and session tokens qualify. Names, identifiers, personal data, licensed text and internal structure do not: they remain recoverable and they remain meaningful. ## The three options, and who owns each **Ship with the invalidatable invalidated.** Fastest. Its honesty depends entirely on the residual being written down and accepted by someone with the authority to accept it, not buried in a risk register nobody reads. Security scopes the residual; the business accepts it. **Refilter and re-run the fine-tune.** Changes what the parameters hold, which is the only thing that does. The cost is not just the compute: to be confident, the filter has to be over-inclusive, and an over-inclusive filter removes real code the assistant would have learned from. Completion quality drops in exactly the areas where the filter was broadest. That trade is a product decision, and the product owner has to be in the room, because security cannot unilaterally spend someone else's quality budget. **Keep the material out of weights.** Serve a base model and put the proprietary content behind an access-controlled layer that the model reads at request time rather than absorbs at training time. This is a different architecture with different costs and different failure modes, but it has one structural advantage: content that was never trained into parameters cannot be recovered from them, and access to it stays revocable per user. ## Shipping weights is a one-way door Exposing the model behind your own endpoint and handing out the checkpoint are not the same decision, even though the model is identical. Behind your endpoint you keep metering, logging, the ability to notice a harvest, and the option to withdraw the model tomorrow. Once the file is out, its holder has unlimited, unlogged, unrated access and every later mitigation you build applies only to your own copy. If the corpus review leaves you uncertain, that uncertainty argues for the endpoint and against distribution — the two decisions deserve different thresholds. ## Write the claim you can defend The sentence people want is "the model does not contain your data". You cannot support it and you should not sign it. What you can support is specific and narrower: which classes were excluded from the corpus and how that exclusion was verified; which classes were present and have since been invalidated; what was accepted as residual. That statement survives someone recovering a span later; the broad one does not, and the gap between them is where regulatory and contractual trouble lives. ## What separates a principal answer here Not the technical content — the willingness to name the residual out loud, to price the utility loss the safe option costs and say who absorbs it, to distinguish the endpoint decision from the distribution decision, and to refuse a claim that a clean probe does not support even when the release date is the loudest thing in the room.

  • The team probed for a week and recovered nothing — is that clearance to ship?
    No. A negative result bounds the prefixes and classes tried at the effort spent, and customers will collectively spend far more, with better prefixes because they use the product daily. Absence of recovery in your run is weak evidence about the tail, and it is not the kind of evidence a distribution decision should rest on.
  • What does an over-inclusive corpus filter cost, and who pays it?
    Utility, and the product owner pays. Filtering hard enough to be confident removes real material the assistant would have learned from, so quality drops in exactly the areas where the filter was broadest. That is a business trade rather than a security one, which is why both owners have to make it together rather than security making it alone.
  • Why is shipping the weights a different decision from serving the same model yourself?
    Because a distributed checkpoint cannot be recalled, and its holder gets unlimited, unlogged, unmetered access with no way for you to notice a harvest. Behind your own endpoint you keep metering, logging and the option to withdraw. The model is identical; the exposure is not, so the two decisions deserve different thresholds.
  • What can you honestly put in a customer-facing statement?
    Which classes were excluded from the training corpus and how that was verified, which were present and have since been invalidated, and what residual was accepted. Not 'the model does not contain your data' — that claim cannot be supported and it is the one that fails badly if a single span is later recovered.

saying these in an interview costs you the question

  • Treats a clean internal extraction run as proof the model is safe to ship
  • Claims the model contains no customer data because probing found none
  • Assumes an output filter can make the ship decision for you
  • Ignores classes of content that cannot be invalidated after the fact
  • Prices the retrain without pricing the utility the filter destroys

context