How would you fulfil a data subject's erasure request end to end across a data platform, from intake to proof of completion?
answer
- verify, then resolve every identifier
- locate with tags and lineage
- per-store action, including exceptions
- processors and downstream copies
- evidence without keeping the data
basics
~20 sVerify the request, resolve the person to every identifier, locate their data through classification and lineage, delete or anonymise per store while honouring retention duties, notify processors, purge old versions, suppress re-ingestion, and log evidence.
solid answer
~50 sI would run it as a tracked workflow with a deadline. **Intake**: record the request and verify identity, following the rules the privacy team sets. **Resolve identity**: map the person to every key the platform uses — customer ID, emails, device and account IDs — through an identity graph. **Locate**: use classification tags and lineage to list every table, extract and feature set that holds those keys. **Act per store**: delete or anonymise, or record why a legal retention duty keeps some records. **Propagate** to processors and downstream tools that received copies. **Purge** old snapshots and file versions, and add the keys, as keyed hashes, to a **suppression list** for ingestion. **Verify** by re-querying, then write **evidence** — what, where, when — without keeping the personal data itself. Batch requests so file rewrites stay affordable within the deadline.
go deeper
Know that an erasure request must reach every place the person's data is stored, not only the main table.
Explain identity resolution, discovery through tags and lineage, and the purge of old versions.
Design the end-to-end workflow with exceptions, processors, suppression, verification and evidence under a deadline.
Decide the architecture and ownership, such as a central erasure service and per-store connectors, and the metrics that show requests complete on time.
## Why a workflow, not a query An erasure request names a **person**, but a data platform stores **identifiers** in dozens of places. What the law requires — which rights apply, which exemptions, which deadline — is decided by the applicable regime and the privacy team; the platform's job is to **execute reliably and prove it**. That needs a tracked workflow, not an engineer running ad hoc deletes. ## The steps 1. **Intake and verification.** Log the request with a ticket and a due date derived from the applicable regime (under the GDPR, one month from receipt, extendable in some cases). Identity is verified by the process the privacy team defines. 2. **Identity resolution.** Map the person to **every identifier**: internal customer IDs, email addresses, phone numbers, device IDs, account IDs in acquired systems. An identity graph or mapping table makes this repeatable. 3. **Discovery.** Query the catalog for datasets tagged as holding those identifier types, and use **lineage** to follow derived tables and extracts. Free-text and log stores need content search. 4. **Decide per location.** Delete, anonymise, or **retain with a recorded reason** where another legal duty applies (invoices kept for tax, for example). This decision comes from the privacy team's rules, not from the engineer running the job. 5. **Execute.** Batched deletes per table; downstream derived tables recomputed or patched. 6. **Propagate.** Send deletion instructions to processors and downstream tools that received copies (marketing, support, analytics vendors), and track their confirmations. 7. **Physical purge.** Expire snapshots and remove unreferenced files so time travel cannot return the data; handle backups by purge or by re-applying deletions on restore. 8. **Suppress.** Add the identifiers, stored as **keyed hashes**, to a suppression list checked at ingestion and in backfills. 9. **Verify and record.** Re-query every location; store evidence: request ID, locations processed, actions, timestamps, exceptions — **not** the deleted data. ## Design choices that make it scale | Choice | Why it helps | |---|---| | Classification tags on identifier columns | discovery becomes a catalog query | | Column-level lineage | derived copies are found automatically | | Partitioning or clustering by subject key where feasible | deletes touch fewer files | | Batch windows well inside the deadline | file rewrites are amortised | | A central erasure service with per-store connectors | each store implements one interface | | Evidence log separate from the data | proof survives without keeping personal data | ## Common failures - Deleting from the main customer table only. - Missing identifiers from an acquired system, so part of the person remains. - No propagation to vendors, who keep processing copies. - Evidence that stores the deleted records "for audit", recreating the problem. ## Why interviewers ask it It combines identity, discovery, execution and proof under a deadline. A senior answer keeps **legal decisions with the privacy team**, builds **discovery on tags and lineage**, remembers **processors, purge and suppression**, and keeps **evidence without data**.
- Why must the evidence log avoid storing the erased data?The log exists to prove the data is gone. If it keeps the deleted records or clear identifiers, it becomes another store of that person's data and would itself need erasing. Store request IDs, locations, actions and timestamps, and identifiers only as keyed hashes if needed.
- Part of a customer's data must be kept for tax records. How does the workflow handle it?The privacy team's rule marks those records as retained under a legal duty. The workflow deletes everything else, restricts the retained records to the purpose that requires them, records the reason and expiry, and deletes them when that duty ends.
saying these in an interview costs you the question
- Running ad hoc deletes without a tracked workflow and deadline
- Resolving the person only by the identifier given in the request
- Forgetting processors and downstream tools that received copies
- Keeping the deleted records in the audit log as proof
- Letting engineers decide legal exemptions case by case