You publish a checkpoint after each honoured deletion — what does that give an adversary holding both versions and a candidate record?
answer
- neither version alone, the difference
- one release, one record
- the bit says they asked to leave
- batch so many change at once
- retire the older artefact
basics
~20 sTwo releases differ by the one change between them: that record. An adversary holding both versions and a candidate can read from that difference whether the candidate was the record removed, which is a bit about a person's withdrawal.
solid answer
~50 sHonouring deletion by shipping a model per request turns the remediation itself into a channel. Each pair of consecutive releases differs by one removal, so an adversary who holds both and a plausible candidate record can compare how the two versions behave on that candidate: a systematic loosening of fit on it across the pair points at it as the record that left. Neither version alone gives that; the *difference* does, which is the point people miss. And the bit is not innocuous — it says a specific person was enrolled and has withdrawn, which is often more sensitive than membership alone. The structural fix is to stop making releases record-aligned: batch removals so a release covers many changes at once, do not publish which request a build honoured, and retire the superseded version. The cost is delay between the request and the claim.
go deeper
Understand the core idea: comparing two versions of a model can reveal what changed between them, so releasing one model per deletion request can point at the person who asked.
Explain the adversary's vantage precisely — both versions plus a candidate record, no privileged access — and what the comparison against unchanged control records buys them.
Show the operational consequences: batch size per release, not publishing the request-to-build alignment, and retiring superseded checkpoints so half the pair stops existing.
Own the counter-intuitive tradeoff in the policy conversation: the most responsive deletion process is the leakiest one, and the retrain interval you negotiate is simultaneously a cost, an SLA and a privacy parameter.
### The channel you created by being responsive There is a genuinely counter-intuitive result in this area: the more promptly and precisely you honour deletion requests by publishing models, the more you leak about the people who filed them. The reason is arithmetic. If two consecutive published checkpoints differ by exactly one honoured removal, then whatever differs between them is caused by that removal (plus whatever training noise the run introduced). An adversary who holds both versions is no longer reasoning about one model's behaviour in isolation, which is the setting every membership defence is designed for. They are reasoning about a difference whose cause is a single record. ### What the adversary needs, and what they get Their vantage is specific and worth stating precisely: both checkpoints, or at least query access to both while the older one is still reachable, plus a candidate record they suspect was the one removed. They do not need the training set, gradients, or any privileged access. What they can do is compare the two versions' behaviour on their candidate against their behaviour on records they know did not change. If the newer version fits the candidate noticeably less tightly than the older one did, while nothing comparable happened to the control records, the candidate is a plausible answer to 'what was removed'. Over a sequence of releases, a candidate can be located to a particular release, which narrows the claim further. The payoff is one bit, and it is a worse bit than plain membership. Plain membership says a person's data was in the training set. This says a specific person was in the training set *and asked to leave*. In many contexts — a health service, a legal service, a channel someone withdrew from after an incident — the act of withdrawal is itself the sensitive fact. ### Why the owner's instrument is the adversary's instrument This is the same test the owner runs to verify the removal worked, pointed the other way. Verification asks 'does the new model still show a signal on this record'. The attack asks 'which record's signal disappeared between these two models'. One instrument, two directions, and it is why a deletion programme cannot treat verification and privacy attack as separate subjects. If your verification method works, an outsider holding both artefacts has a working attack. ### The properties that reduce it These are structural, not clever: **Do not make a release correspond to one request.** If removals are accumulated and applied in a scheduled retrain, each release differs by many records at once, and no single candidate is implicated by the difference. The number of removals per release is the thing that controls the leak, and it is a scheduling decision. **Do not publish the alignment.** Release notes, ticket references, or a timestamp that lines up with a request re-create the alignment even when the batch was large. If somebody can tell which build honoured which request, batching bought nothing. **Retire the old artefact.** The attack needs both versions. A superseded checkpoint that stays reachable — still serving a region, still on devices, still downloadable — keeps supplying half of the pair long after the new one shipped. **Per-record bounds help here too.** Training that bounds how much any single record can move the parameters directly bounds how much a single removal can move them, which is the quantity this attack reads. That is the same machinery that limits membership signal in a single model, doing double duty across a pair. ### The tradeoff you are actually making Every one of those mitigations costs promptness. Batching means the interval between the request and the moment the removal is real, and someone outside engineering has to agree that interval is acceptable — and be told what it is, rather than being promised an immediate deletion that the pipeline never performed. The uncomfortable framing to bring to that conversation: honouring requests individually and immediately is the version that leaks most, and a policy written by someone who has not heard this will ask for exactly that.
- Why is the leaked bit worse than ordinary membership evidence?Ordinary membership says a person's data was used. This says the person was in the set and requested removal, so it discloses an action they took, not just a fact about the corpus. In contexts where withdrawal itself signals something — a health or legal service, a channel someone left after an incident — that action is the sensitive part, and it was created by the remediation.
- Does batching removals actually change the adversary's problem, or just annoy them?It changes the problem. With one removal per release, a detected change implicates one candidate. With many, a change is consistent with any of them, and the adversary can at best conclude that a candidate is among the batch. The uncertainty scales with batch size, which makes the retrain interval a privacy parameter as well as a cost parameter.
- What does keeping the superseded checkpoint reachable cost you here?It supplies the other half of the pair. The attack needs both versions; if the old build is still serving somewhere, still installed on devices, or still downloadable, an adversary can construct the comparison at any time. Retiring superseded artefacts is part of the removal, not cleanup afterwards, and it is usually the step that gets deferred.
Publishing a class roster before and after one student leaves reveals which student left, even though neither roster is labelled.
saying these in an interview costs you the question
- Analyses each release in isolation and misses the difference
- Thinks prompt per-request releases are strictly safer
- Treats the leaked bit as equivalent to plain membership
- Leaves the superseded checkpoint reachable
- Batches removals but publishes which request each build honoured
- Assumes the adversary needs the training set or gradients