A speaker withdraws consent after three published models trained on their recordings - how do you honour the erasure?
answer
- immutability is not a legal exemption
- consent lives outside the corpus
- re-cut shards, keep a tombstone
- rebuild equals set minus suppressions
- retention is erasure on a clock
basics
~20 sYou cannot delete the bytes and keep every frozen snapshot byte-rebuildable. Remove the recordings, mint new snapshots without them, and keep the old ids as tombstoned manifests plus a counted suppression list - so a rebuild is exact except for a named, recorded exclusion.
solid answer
~50 sImmutability is an engineering property, not an exemption from the right to erasure, so the source bytes go and reproducibility downgrades honestly. The workable architecture keeps consent **outside** the corpus: rows carry a subject id, a separate mutable registry holds consent state, and every read of any snapshot applies the **current suppression list**, so past and future rebuilds honour accumulated erasures automatically. The old snapshot id is retained as a tombstone - the original manifest, the suppressed subject ids, the reason class, an ISO 8601 timestamp and the approver - so an auditor can still see exactly what trained a model and exactly how the rebuildable set now differs. Whether the three **models** must be rebuilt is a separate policy call, not a property a hash can answer: pick immediate rebuild or erase-now-retrain-on-next-cycle per model, and write the position down. Retention expiry runs the same path on a clock, so the retention window and the how-long-must-this-stay-rebuildable window are reconciled at design time.
go deeper
Recall the tension in one line: a frozen training set and a deletion request pull in opposite directions, and the deletion wins.
Explain the mechanics: shards are re-cut rather than edited, a new id is minted, and the old manifest is kept as a tombstone rather than destroyed.
Show the operational path: consent held outside the corpus, a suppression list applied on every read, and the copies - exports, evaluation sets, backups - covered too.
Own the policy: which guarantee you give up, what erasure latency you commit to, when a model is rebuilt versus refreshed, and what you tell an auditor about non-rebuildable snapshots.
## The impossible triangle Three things are wanted at once and only two are achievable together: 1. The withdrawn recordings are **actually gone**, everywhere they were copied. 2. Every published snapshot id remains **byte-exactly rebuildable**. 3. Published models are **left as they are**. The right to erasure (GDPR Article 17, and its analogues elsewhere) makes the first non-negotiable, so the design question is which of the other two you give up and how you record the loss. The answer that reads as senior engineering and as a defensible position is: keep (1) and (3) by default, degrade (2) deliberately, and make the degradation itself a recorded, counted, auditable fact rather than a silent one. ## Consent as an indirection, not as data The structural move is to stop storing consent state inside the immutable corpus. Every utterance carries a **subject id**; a separate, mutable registry holds each subject's current consent state and the date it changed. Then: - The corpus stays immutable and content-addressed, as versioning requires. - A withdrawal is a single write to a mutable registry, which is the only place it can be handled in near real time. - Every reader - training jobs, evaluation, exports, ad-hoc analysis - applies the **current suppression list** as it resolves a snapshot, so no path can accidentally read withdrawn data, including a rebuild of a two-year-old id. This is what makes the erasure hold going forward. It does not delete the bytes, which is a separate, physical step. ## Deleting the bytes without rewriting history Shards are content-addressed, so a shard cannot be edited: the affected shards are **re-cut without the withdrawn utterances**, producing new hashes and a new snapshot id. The old manifest is kept as a **tombstone** rather than deleted, because destroying it destroys the ability to explain what the old model was trained on. The tombstone records the original manifest, the suppressed subject ids or row ids, the reason class, the timestamp and the approver, and it counts the removals. Rebuilding an old id then produces **the recorded set minus a counted exclusion list**, and the rebuild reports that difference instead of pretending to be identical. One alternative worth naming because architectures genuinely differ here: some platforms store each subject's raw audio under a per-subject key and destroy the key on withdrawal, leaving the shard bytes in place but unreadable. It is cheaper than re-cutting terabytes, and whether it counts as erasure at all is a question for counsel in your jurisdiction, not an engineering preference. ## What happens to the three models Nothing about a content hash tells you whether a model 'contains' a withdrawn speaker's data - that is a policy position, and organisations take different ones. What a lead owes is a **decision rule written in advance**, applied per model: | situation | usual position | |---|---| | model retrains on a short cycle and the contribution is diffuse | erase the data now, let the next scheduled refresh drop it out | | model is long-lived, rarely refreshed, or the withdrawal is contractually load-bearing | rebuild the model explicitly, on a stated deadline | | model is archived and not serving | erase the data; archive the model with the tombstone attached | Commit to a **latency** for each path, because 'eventually' is not an answer an erasure request accepts. ## Do not forget the copies Erasure applies to everything derived, not just the row in the corpus: exported analysis extracts, evaluation sets, cached transcripts, quality-review samples, backups and any downstream dataset built from the corpus. Each needs either the same suppression path or a stated expiry that bounds how long a copy can survive the request. A deletion that misses the copies is not a deletion. ## Retention expiry is the same machinery on a clock Retention is foreseeable erasure: on its schedule it removes shards that published models depend on. Because it is foreseeable, it is planned - the retention clock and the 'how long must this model stay rebuildable' clock are reconciled when both are set, not discovered in conflict during an audit. When they cannot be reconciled, the honest outcome is a documented statement of which models are explainable but no longer rebuildable, and from what date.
- What must a tombstoned snapshot record to stay useful to an auditor?The original manifest, the suppressed subject or row identifiers, the reason class, a timestamp in ISO 8601, who approved it and how many rows went. That is enough to state precisely how the rebuildable set differs from the set that trained the model, without restoring any erased content to do so.
- Does retention expiry need the same machinery as an erasure request?It uses the same deletion path, but it is foreseeable, so it is planned rather than reacted to. Set the retention window against how long each model must remain rebuildable; where they conflict, record which models become explainable-but-not-rebuildable and from when, rather than discovering it during an audit.
- Why not simply refuse to include any contributor whose consent might later be withdrawn?That is not a real option for a contributed corpus - consent can be withdrawn by anyone at any time - so the architecture has to assume withdrawal is routine. Designing the suppression path and the tombstone up front costs far less than re-cutting terabytes under a deadline the first time it happens.
saying these in an interview costs you the question
- Says immutability means the data cannot be deleted, so the request is refused
- Deletes rows in place and lets published ids resolve to fewer rows silently
- Treats deleting the corpus row as enough while exports and derived sets keep copies
- Assumes every erasure request forces an immediate rebuild of every model
- Plans retention expiry independently of how long models stay rebuildable
- Believes a content hash can be edited to drop one row from a snapshot