skip to content

Deleting From Weights

Deleting the row does not touch weights already trained on it, and 'gone' is a claim that needs an instrument. Interviewers probe why deletion and the privacy attack are one subject read two ways.

on this pageshow

explore

questions

5

A customer's row is deleted from the training store — what can an adversary who can only query the deployed model still learn?

level: juniorimportance: must knowfreq 62%

answer

  1. storage and parameters are two different places
  2. the fit already happened
  3. a delete changes the next run only
  4. no weights needed to test it
  5. advantage over the base rate

basics

~20 s

Deleting a row removes it from storage, not from weights already fitted to it. Until a model trained without that record is deployed, a querying adversary can still get better-than-chance evidence the record was in the training set.

solid answer

~50 s

Deletion and training touch different objects. The store is where records live; the weights are the output of a run that already read that record and moved parameters because of it. Removing the row changes what the *next* run sees and nothing about the model currently serving traffic. So an adversary who holds no weights and can only send inputs and read outputs still faces a model whose fit is measurably stronger on examples it trained on — the signal a membership test measures as advantage over the base rate. Nothing about the delete reduces that advantage. What reduces it is replacing the deployed model with one fitted on a corpus that never contained the record, which happens on the retrain schedule, not on the delete. Older checkpoints and any model fine-tuned from them carry the same influence until they are retired too.

go deeper

for a junior

Be ready to state the split in one breath: the store lost the row, the weights did not. Know that a model can leak that a record was used without storing the record.

for a middle

Explain the mechanism, not just the slogan: the fit is tighter on examples seen during training, and a query-only test converts that gap into above-chance evidence about a single candidate record.

for a senior

Show you know which artefacts are in scope — live weights, prior checkpoints, fine-tuned derivatives — and that the claim only becomes true when a model fitted without the record actually ships.

for a principal

Own the framing you give non-engineers: deletion here is a schedule and a claim about which parameters serve traffic, not an operation with an instant receipt. Say so before somebody promises otherwise in writing.

### Two different objects with two different histories A trained model is not a container that holds rows. Training is a process: it reads a dataset and produces parameters. Every example in that dataset had some influence on where the parameters ended up — large for an unusual example the model had to bend to fit, small for one of ten thousand near-identical ones, but never exactly zero. When a customer withdraws consent and the row is deleted from the training store, exactly one of those two objects changes. The corpus no longer contains the record and the next training run will not see it. The deployed weights are unchanged, because nothing about a delete re-runs the fit. A deployed model does not learn that a record has been withdrawn; it does not learn anything after training ends. So the correct first answer to 'is it gone' is a question back: gone from *what*? From the store, yes, immediately and auditably. From the deployed parameters, not yet, and not until something replaces them. ### What the adversary here actually has The adversary in this picture is deliberately weak. They do not hold the weights, they cannot see gradients, and they may only send inputs to the served endpoint and read what comes back. That is the vantage of anyone with an account. The limit matters, because it is exactly the vantage a company implicitly claims is harmless when it says the model does not store the data. What that adversary can do is compare how the model behaves on a candidate record against how it behaves on records it plausibly never saw. Models fit what they trained on more tightly than what they did not — that is imperfect generalization, and it is present in any model that overfits at all, which is every useful model to some degree. A membership test turns that gap into a decision about one candidate record, and it is scored as advantage over the base rate: if half the candidates you test really were in the training set, a coin gets 50 percent, and only the excess over 50 is a disclosure. Because that advantage comes from the fit, and the fit did not change, the advantage did not change when the row was deleted. ### Why 'the model does not store training data' is not an answer This is the reflex answer and it is a category error. Nobody claimed the model contains a copy of the record. The claim being made against you is weaker and easier to support: that the model behaves measurably differently because the record was there. Influence is not storage, and a membership test does not need storage — it needs a difference in behaviour between seen and unseen examples. Two related confusions are worth separating. Verbatim recall — a large model reproducing a training string exactly — is a stronger and rarer phenomenon. Membership evidence is the weak, common one, and it is the one that decides whether a deletion claim holds up. ### A concrete shape Take a bank's phone channel with a speaker-verification model, where each enrolled customer's audio contributed to fitting the matcher. A customer withdraws consent. Operations deletes the enrolment from the store and closes the ticket. The matcher in production was fitted with that customer's audio in the training set and keeps serving. Someone holding a plausible sample of that customer's voice and an account on the channel can still ask the model questions and get an above-chance read on whether that voice was part of what it was trained on. The withdrawal did not change the model; it changed the queue for the next one. ### What actually moves the needle Only a new model. A run over a corpus that never contained the record produces parameters that are a function of a dataset the record was never in, so there is nothing about that record for a membership test to find — with two honest caveats: near-duplicates and strongly correlated records (the same customer enrolled twice, a family member with similar audio) leave a residue that is not the deleted record but behaves a little like it; and any checkpoint or derivative model that *was* trained with it is still carrying it wherever it is still deployed or distributed. That is why deletion in this area is never a database operation with a receipt. It is a claim about which parameters are currently serving, and it becomes true on a schedule. ### The answer to give Say clearly that deletion from the store bounds future runs only; that the live model retains the record's influence until a model fitted without it is deployed; that the instrument that would demonstrate residual influence is a membership test scored against the base rate; and that any older artefact still in circulation is still in scope.

  • The team says the model never stores raw inputs, so nothing can leak. Why is that not a defence?
    Nobody is claiming the record is stored. The claim is that the model behaves differently because the record was in the fit, and a membership test measures that difference from outputs alone. Storage is irrelevant to it. A model can hold no copy of anything and still give an adversary above-chance evidence about who was in the training set.
  • Besides the live model, what else still carries the record after the delete?
    Every checkpoint produced by a run that included it, any model fine-tuned or distilled from such a checkpoint, and any copy of those still distributed or held on a device. A deletion claim that covers only the currently served weights is false for those artefacts. They have to be retired or replaced as part of the same commitment.
  • Does it change anything if the deleted record was one of ten million?
    It usually shrinks the signal but does not zero it. An ordinary, densely surrounded record has little individual influence and is hard to detect. An unusual one — a rare accent, an outlier feature vector, a duplicated record — keeps a detectable footprint even in a very large corpus, and those are exactly the people a deletion request is most likely to come from.

Removing a student's exam paper from the filing cabinet does not change the grade curve that was already computed from it.

saying these in an interview costs you the question

  • Thinks a store deletion propagates into the trained weights
  • Says a model does not keep training data, so nothing leaks
  • Assumes the retrain has already happened
  • Treats a database delete receipt as proof of removal
  • Confuses verbatim recall with membership evidence
  • Forgets older checkpoints and fine-tuned derivatives

context

open as a page

Why is a full retrain without a record the only removal an adversary with query access cannot contest?

level: middleimportance: should knowfreq 45%

basics

~20 s

Retraining over data that never contained the record leaves nothing for a membership test to find: the parameters come from a corpus it was never in. Approximate unlearning only corrects existing weights and claims closeness to that ideal.

open as a page

Your red-team membership test, 200 queries per record, finds nothing on the retrained model — what does that establish?

level: seniorimportance: should knowfreq 38%

basics

~20 s

It establishes that one attack, at one strength, on the records you sampled, did not beat the base rate. It bounds that attack, not the model, and says nothing about a stronger adversary, untested records, or checkpoints still in circulation.

open as a page

What can you commit to a regulator about a withdrawn voiceprint, given a querying adversary and a monthly retrain?

level: principalimportance: should knowfreq 33%

basics

~20 s

Commit to three separable claims: the record is deleted from stores and excluded from future runs, effective now; the served model is fitted without it at the next monthly retrain; residual influence is bounded by evidence you state.

open as a page

You publish a checkpoint after each honoured deletion — what does that give an adversary holding both versions and a candidate record?

level: seniorimportance: nice to knowfreq 24%

basics

~20 s

Two releases differ by the one change between them: that record. An adversary holding both versions and a candidate can read from that difference whether the candidate was the record removed, which is a bit about a person's withdrawal.

open as a page