skip to content

Compliance, Retention & Erasure

How a platform deletes one person from immutable lakes and warehouses, enforces retention windows and proves it. Right to be forgotten in a lakehouse has no lazy answer.

on this pageshow

explore

questions

6

Why is deleting one customer's data from a data lake much harder than deleting their row from an application database?

level: middleimportance: must knowfreq 60%

answer

  1. immutable files, not rows
  2. copies at every layer
  3. old snapshots still hold it
  4. backups, extracts, derived tables
  5. replays bring it back

basics

~20 s

Lake data sits in immutable files copied through raw, cleaned and curated layers, kept in old snapshots, backups and extracts. Deleting a person means rewriting files everywhere they appear, expiring old versions, and stopping replays from reloading them.

solid answer

~40 s

In an application database the person is a few rows, updated in place, and a `DELETE` removes them. In a lake, the same person is spread across **immutable files**: a delete means **rewriting** each file that holds their rows (or writing delete markers that later compaction applies). They also appear in **every layer** — raw landing, cleaned, curated marts — and in **derived tables**, **extracts** and **feature sets**. Table formats keep **old snapshots** for time travel, so the rows stay readable until those snapshots are expired and their files physically removed; **backups** hold them longer still. Finally, if raw data is ever **replayed or backfilled**, the person comes back. So erasure in a lake is a workflow: find every location, rewrite, expire, purge, and keep a suppression list so reprocessing does not reintroduce them.

go deeper

for a junior

Know that lakes keep data in files and copies across layers, so deleting one person touches many places.

for a middle

Explain rewriting immutable files, snapshot expiry and why replays can reintroduce deleted data.

for a senior

Design the erasure workflow with location via tags and lineage, batching, purge, suppression and evidence.

for a principal

Decide the platform rules, such as partitioning by subject, raw-zone retention and backup handling, that keep erasure affordable at scale.

## The contrast | Aspect | Application database | Data lake or lakehouse | |---|---|---| | Storage unit | rows and pages updated in place | large immutable files in object storage | | Delete operation | `DELETE` removes rows | rewrite affected files, or write delete markers purged later | | Copies | usually one primary plus replicas | raw, cleaned and curated layers, marts, extracts, feature sets | | Old versions | overwritten, then reclaimed | kept as snapshots for time travel until expired | | Backups | database backups | object versions, snapshots, backup copies, archive tiers | | Reloading | not usual | replays and backfills from raw are routine | ## Where the person hides 1. **Raw landing zones**, often kept "forever" so pipelines can be rebuilt. 2. **Every downstream layer**: cleaned tables, dimensional models, aggregates small enough to single someone out. 3. **Derived artefacts**: machine-learning feature tables, extracts sent to other teams, cached dashboard results. 4. **Old table versions**: a lakehouse table format keeps prior snapshots so readers can query the past; a deleted row is still in files referenced by those snapshots. 5. **Object-store versioning and backups**: deleted files may survive as previous object versions. 6. **Logs and free text**: event payloads and support notes with the person's details embedded. ## What deletion actually takes - **Locate**: use classification tags and lineage to find every table holding the person's identifiers, not only the obvious customer table. - **Rewrite**: run deletes per table; the table format rewrites files (copy-on-write) or records delete files (merge-on-read) that later compaction applies. - **Expire and purge**: expire snapshots older than the deletion and remove unreferenced files, otherwise time travel still returns the rows. - **Handle backups**: either purge them, or record the deletion so it is re-applied if a backup is ever restored. - **Suppress reintroduction**: keep a list of erased identifiers (stored as keyed hashes) and filter it in ingestion and backfills. - **Record evidence**: which locations were processed, when, and with what result. ## Batching Rewriting large files for each request is expensive, so platforms usually **batch** erasure requests and run them on a schedule within the deadline the applicable regime sets, often partitioning tables so that one person's rows are concentrated rather than scattered across every file. ## Why interviewers ask it It is the classic data-platform privacy question. A good answer contrasts **update-in-place** with **immutable files**, lists the **hidden copies** (layers, snapshots, backups, extracts), and knows that deleting is not finished until old versions are **physically purged** and **replays cannot restore** the data.

  • Why does a delete in a lakehouse table not immediately remove the data from storage?
    The table format writes new files or delete markers and commits a new snapshot, but earlier snapshots still reference the old files. Until those snapshots are expired and unreferenced files removed, time travel or a direct file read can still return the rows.
  • How do you stop a backfill from reloading erased customers?
    Keep a suppression list of erased identifiers, stored as keyed hashes rather than in clear, and apply it as a filter at ingestion and in any replay or backfill job, so raw data that still contains them is dropped on the way in.
  • Why batch erasure requests instead of processing each one immediately?
    Each request forces file rewrites across many tables; batching amortises that cost so one rewrite handles many people. The batch interval must stay well inside the deadline the applicable regime sets.

saying these in an interview costs you the question

  • Believing a DELETE statement on the main table completes the erasure
  • Forgetting snapshots kept for time travel
  • Ignoring raw landing zones and extracts sent to other teams
  • Letting backfills from raw data reintroduce erased people
  • Storing the suppression list of erased people in clear text
open as a page

How do you define and enforce retention periods across a warehouse and a lake so that expired data is actually deleted on time?

level: middleimportance: should knowfreq 45%

basics

~20 s

Give every dataset a retention class tied to a named date column, partition by that date, run automated jobs that drop expired partitions and then physically purge old versions, and log evidence. Legal holds and exceptions are explicit, recorded overrides.

open as a page

How would you fulfil a data subject's erasure request end to end across a data platform, from intake to proof of completion?

level: seniorimportance: should knowfreq 45%

basics

~20 s

Verify the request, resolve the person to every identifier, locate their data through classification and lineage, delete or anonymise per store while honouring retention duties, notify processors, purge old versions, suppress re-ingestion, and log evidence.

open as a page

Engineers want to keep raw event data with identifiers forever so pipelines can always be rebuilt; how do you decide what the platform actually keeps?

level: principalimportance: should knowfreq 30%

basics

~20 s

Weigh the real value of full rebuilds against the risk, erasure cost and legal limits of keeping identifiable raw data. Usually keep identifiable raw data for a bounded rebuild window, then keep pseudonymised or aggregated forms, with legal minimums as explicit exceptions.

open as a page

When does crypto-shredding, encrypting each person's data with their own key and destroying the key on erasure, beat physically deleting rows in a lakehouse, and what does it cost?

level: seniorimportance: nice to knowfreq 30%

basics

~20 s

Crypto-shredding encrypts each person's identifying fields with a per-person key; erasure destroys the key, making every copy unreadable, including backups and old snapshots. It wins where rewriting copies is impractical, but costs key management at scale and query-time decryption.

open as a page