A privacy regulation requires you to erase an individual's personal data on request, but your system soft-deletes everything. How do you actually satisfy such an erasure request, and what does a purge process have to cover?
answer
- flag ≠ erasure — data still readable
- hard delete / anonymize / crypto-shred
- batched, ordered, idempotent, audited
- follow copies: replicas, index, warehouse, logs
- backups = retention window; drop partition beats mass DELETE
basics
~20 sA flag is not erasure — the data is still readable. Satisfy the request by hard-deleting or irreversibly anonymizing the personal fields, keeping only what law requires, and run it as a batched, idempotent job that also reaches replicas, search indexes, exports and analytics copies.
solid answer
~60 sFirst separate two events that share a button: **"user deleted their account"** (an application state, soft delete is fine) and **"erasure request"** (the personal data must stop existing). For erasure you have three real techniques: - **Hard delete** the rows containing personal data. - **Anonymize in place** — irreversibly overwrite name, email, address, IP with tombstones — when you must retain the record itself for legal reasons (invoices, tax records, fraud history). - **Crypto-shredding** — store personal fields encrypted per subject and destroy that subject's key, which is useful when the data is spread across immutable stores such as event logs and backups. The purge job must be **batched** (small transactions, so you do not hold locks or bloat the undo/WAL), **ordered** with respect to foreign keys, **idempotent and resumable**, and **audited** — you record that the erasure happened, without re-storing what you erased. Then follow the data: replicas, logical replication targets, materialized views, caches, search indexes, the warehouse, log files, exports and backups (usually handled by a documented backup retention window rather than surgical edits). Legal holds and statutory retention override erasure for specific records; that conflict is a policy decision, not an engineering one.
go deeper
Know the headline: a flag hides data but does not erase it, so erasure needs a real delete or irreversible anonymization.
Describe the two-step pattern — soft delete starts a retention clock, a purge job performs the erasure — and mention batching and foreign-key ordering.
Cover anonymize-vs-delete-vs-crypto-shred, an idempotent resumable job, auditing the erasure, and following every downstream copy including caches, search and the warehouse.
Own the policy surface: a personal-data inventory, retention and legal-hold rules, backup windows, key-management for crypto-shredding, and a retention architecture (partitioning) that makes purge cheap by design.
## Soft delete is not erasure Setting `deleted_at` hides a row from your application. The name, email address and payment details are still in the table, still in every index, still in every replica and backup, and still visible to anyone with database access or a leaked dump. Under privacy regimes that grant a right to erasure, that is not deletion. Treating the flag as compliance is the single most common mistake in this area. So the schema has to support two distinct operations: 1. **Account closure / user-initiated delete** — a reversible application state. Soft delete with a grace period is the right tool, because users change their minds and because you often need the record for the winding-down of the relationship. 2. **Erasure** — the data must genuinely stop existing. This is irreversible by definition. A useful pattern is to make closure the trigger and erasure a scheduled follow-up: `deleted_at` starts a retention clock, and a purge job erases rows whose clock has expired unless a legal hold applies. ## The three erasure techniques **Hard delete.** The cleanest: the rows go away. Works when the record has no independent legal reason to exist (profile, preferences, device tokens, session history). Requires deleting in foreign-key-safe order, or leaning on declared `ON DELETE CASCADE`. **Anonymization / pseudonymization.** Often you cannot delete the row: an invoice must survive for tax law, a moderation decision must survive for safety, an aggregate must stay consistent. The answer is to strip the row of the fields that identify a person — overwrite `email`, `full_name`, `phone`, `ip_address` with fixed tombstones or a one-way surrogate — while retaining the non-identifying skeleton. Two cautions: it must be *irreversible* (a hash of an email is reversible by brute force over a known address space, so a hash alone is not anonymization), and you must consider re-identification via the remaining columns (a rare postcode plus a birth date can identify someone even without a name). **Crypto-shredding.** Encrypt each subject's personal fields with a per-subject key held in a key store; erasure destroys the key, making every copy of the ciphertext — including ones inside append-only event logs, immutable archives and old backups — permanently unreadable. This is the practical answer for architectures where physically rewriting history is impossible. It costs key-management complexity and makes those fields unsearchable. ## What a purge job must do **Batch it.** `DELETE FROM events WHERE user_id = 42` against ten million rows takes long locks, generates a huge amount of write-ahead log and replication lag, and can blow up the undo/rollback segment. Delete in bounded chunks (a few thousand rows) with a commit between batches, ideally driven by an indexed key range. **Order it.** Children before parents, unless real cascade actions are declared. In a soft-delete schema the cascade was never exercised, so the constraint graph may be incomplete — verify it rather than assuming. **Make it idempotent and resumable.** It will be interrupted. Re-running must be safe and must pick up where it stopped, so drive it from a queue of erasure requests with a per-request status rather than from a single long statement. **Audit the erasure, not the erased data.** Regulators expect proof that you complied. Record subject id, request time, completion time, and which stores were covered — never a copy of the erased values. **Follow the data everywhere.** The relational table is the beginning: physical and logical replicas, read replicas used by reporting, materialized views, caches, search indexes, the analytics warehouse and its downstream marts, object storage (uploaded documents, avatars), message queues and event streams, application logs and metrics with user identifiers, third-party processors you sent data to, and CSV exports someone emailed. An erasure design that stops at the primary database is not an erasure design. This is why teams inventory personal-data locations up front; you cannot erase from a store you did not know existed. **Backups are the standard exception.** Surgically editing backups is impractical and dangerous. The accepted approach is a documented, bounded backup retention window plus a rule that any restore re-applies outstanding erasure requests. Say this explicitly in an interview — it shows you have done it rather than theorized it. ## Storage reality after a purge In multi-version engines a `DELETE` marks row versions dead rather than freeing space; background cleanup (vacuum, purge threads) reclaims it, and index space often does not shrink until a rebuild. Deleting hundreds of millions of rows can leave the table as large as before and slower than before. Where the data is naturally time-ordered, **dropping a partition** is drastically cheaper than deleting rows — a metadata operation that frees the space instantly, with no bloat. Designing retention around partitioning is the difference between a purge that runs in seconds and one that runs for days. ## The conflicts you should name Erasure fights with statutory retention (financial records), with legal holds (litigation), with fraud prevention (you may need a hashed identifier to keep a banned actor out), and with derived analytics that were already aggregated. The engineering job is to surface these conflicts and implement whatever the policy decides — typically: erase identifying columns, retain the minimum legally-required skeleton, and keep an irreversible one-way token where continuity is genuinely required.
- When can't you simply hard-delete the row, and what do you do instead?When the record itself must survive for a legal or safety reason — invoices for tax law, transaction records for anti-money-laundering, a ban decision for platform safety. You then anonymize: irreversibly overwrite the identifying columns with tombstones while keeping the non-identifying skeleton. Check that the remaining columns cannot re-identify the person on their own, and make sure the replacement is genuinely one-way rather than a reversible hash of a small value space.
- How do you handle personal data sitting in old backups?You generally do not edit backups. The accepted practice is a documented, bounded backup retention window so the copies age out, combined with a documented procedure that any restore replays outstanding erasure requests before the system serves traffic. Where the data is inside truly immutable stores, crypto-shredding — destroying the per-subject key — makes the retained ciphertext unreadable without rewriting anything.
- Why is a mass DELETE sometimes worse than expected, and what is the alternative?In MVCC engines a delete creates dead row versions that background cleanup must reclaim; the table and its indexes often stay the same size, and the operation generates large amounts of write-ahead log plus replication lag. If the data is time-ordered, partition by time and drop whole partitions instead — a metadata operation that frees space immediately, with no bloat and no long-running transaction.
saying these in an interview costs you the question
- Claiming a deleted_at flag satisfies a right-to-erasure request
- Hashing an email address and calling it anonymized
- Running one enormous DELETE in a single transaction on a live system
- Stopping at the primary database and forgetting replicas, search indexes, the warehouse and logs
- Promising surgical deletion from backups instead of relying on a retention window or crypto-shredding