skip to content

Why is a full retrain without a record the only removal an adversary with query access cannot contest?

level: middleimportance: should knowfreq 45%

answer

  1. one changes the fit, one corrects it
  2. a claim about the process, not a measurement
  3. closeness, with assumptions attached
  4. the bound degrades as requests pile up
  5. one training run per honoured request

basics

~20 s

Retraining over data that never contained the record leaves nothing for a membership test to find: the parameters come from a corpus it was never in. Approximate unlearning only corrects existing weights and claims closeness to that ideal.

solid answer

~50 s

Exact retraining changes the causal history of the weights: the parameters now come from a run in which the record played no part, so residual influence is not small, it is absent. That is the only removal whose claim does not depend on an assumption. Approximate unlearning starts from weights that *did* see the record and applies a correction intended to undo its contribution. Its guarantee is a closeness claim — that the parameters are indistinguishable, to a stated degree, from what a retrain would have produced — and it holds under conditions: bounded per-example influence, a recorded training trajectory, and requests not chosen to defeat it. Those bounds also compose across requests. Two honest caveats on the exact side: near-duplicate and correlated records leave a residue, and older checkpoints still in circulation are unaffected. The price of exact removal is a training run per honoured request unless you batch.

go deeper

for a junior

Know the two options by name and by strength: retrain the model without the record, or correct the existing weights. The first is certain, the second is a bound, and the difference is the whole answer.

for a middle

Explain why the retrain claim is about the process — the parameters come from a corpus the record was never in — while the approximate claim is a closeness statement that needs assumptions to hold.

for a senior

Bring the caveats unprompted: correlated and duplicated records, checkpoints already distributed, and the way approximate bounds compose across a queue of requests rather than resetting each time.

for a principal

Own the cost line. One training run per honoured request against a batching interval is the tradeoff you are funding, and the interval you pick becomes the delay you are committing to in a policy somebody outside engineering will read.

### The claim each method can make The question 'is the record's influence gone' has two very different answers depending on how you removed it. **Exact retraining** re-runs the fit over a corpus that never contained the record. The resulting parameters are a function of a dataset in which the record does not appear, so there is no residual influence to bound: the quantity is zero by construction. That is a statement about the process, not a measurement, which is why it is the only version of the claim that does not rest on assumptions somebody could attack. **Approximate unlearning** never re-runs the fit. It starts from the weights that already absorbed the record and applies a correction — an update meant to push the parameters to roughly where they would have been. Its guarantee is therefore a *closeness* claim: that the corrected parameters are hard to distinguish from the ones a retrain would have produced, to a stated degree, with a stated failure probability. The shape is familiar to anyone who has read a privacy parameter: a bound with numbers attached, not an absence. ### What the closeness claim is standing on The assumptions are where an interviewer will push, and they are the reason the guarantee is much weaker than the word 'unlearning' suggests. - **Bounded influence.** The correction has to know roughly how much the record moved the parameters. That is well behaved for simple, convex models and much less so for a deep network, where the same record can matter enormously at one point in the trajectory and negligibly at another. - **Knowledge of the training history.** Many methods need the training trajectory, the batch composition, or intermediate state. If your pipeline did not record it, the method's premise is missing, and a claim whose premise is missing is not a weaker claim, it is no claim. - **Non-adversarial requests.** The bounds typically assume the set of deletion requests is not chosen to be maximally awkward. An adversary who can influence which records are inserted and then request their deletion is outside the model the bound was proved in. - **Composition.** Every honoured request consumes some of the guarantee. Sequential deletions degrade it, and a method that is defensible for a handful of requests per quarter may assert almost nothing after a thousand. None of this makes approximate methods useless. It makes them a bound you must state out loud, alongside its assumptions, rather than an answer you can give as a yes. ### The adversary who decides whether the claim holds The test of either method is not a proof read by a lawyer, it is an adversary with query access to the served model, holding a candidate record. They ask whether the model still fits that candidate the way it fits records it trained on. Against a genuine retrain, they have nothing to find. Against an approximate removal, whether they find something is exactly what the closeness bound is about — and a bound is a promise about the worst case *within its assumptions*, which is why an empirical result that the attack failed is not the same as the bound being sound. ### The two caveats on the exact side Exact retraining is the strong option, not a perfect one. First, **correlated and duplicated records**. If the withdrawn enrolment has a near-twin still in the corpus — the same customer enrolled twice under two accounts, a household member with similar audio, a row derived from the same underlying event — the retrained model still fits things that look like the deleted record. That is not the record's influence; it is a correlated record's influence, and it will still produce a membership-looking signal for a candidate that resembles both. Deletion policy has to decide whether it is removing a record or a person. Second, **artefacts already in circulation**. A retrain fixes the weights it produces. Every earlier checkpoint, every model fine-tuned or distilled from one, and every copy shipped to a device was produced by a run that included the record and is unaffected by any amount of retraining afterwards. A removal commitment that does not include retiring those is incomplete. ### The limit that decides which one you use Cost. Exact removal is one full training run per honoured request if you honour them one at a time, which for a large model is the dominant number in the whole conversation. In practice organisations batch: accumulate requests, retrain on an interval, and the interval becomes the delay between the request and the moment the claim is true. Some architectures reduce the bill by partitioning the corpus so only affected shards are refit, which is exact but constrains how the model is trained in the first place. Approximate methods exist because the bill is real, and they are a legitimate choice — as long as what gets written down is the bound they actually assert and not the word 'deleted'.

  • What single fact about the deleted record most weakens an exact-retrain claim?
    That a near-duplicate or strongly correlated record is still in the corpus. The retrained model never saw the deleted row, but it still fits something that resembles it, so a candidate matching both still looks familiar to the model. The removal was of a record; the concern was usually about a person, and those differ whenever the same subject contributed more than once.
  • How do teams get exact removal without paying a full training run per request?
    By batching requests into a scheduled retrain, so many removals share one run and the interval becomes the honoured delay, or by structuring training so the corpus is partitioned and only affected parts are refit. The second is still exact, but it constrains the architecture and training procedure up front, which is a decision made long before the first request arrives.
  • Why does a sequence of approximate removals weaken the guarantee more than one does?
    The bounds compose. Each correction is asserted relative to the state it started from, and errors accumulate as the corrected weights drift from any real retrained model. A guarantee that is defensible for a handful of requests can assert almost nothing after a large number, so the method needs a periodic exact retrain to reset the accounting.

saying these in an interview costs you the question

  • Treats approximate unlearning as equivalent to deletion
  • Cannot name any assumption the closeness bound needs
  • Claims retraining also fixes older distributed checkpoints
  • Ignores near-duplicate records when claiming exact removal
  • Thinks a failed attack proves the bound is sound
  • Never mentions the per-request training cost

context