skip to content

Why is deleting one customer's data from a data lake much harder than deleting their row from an application database?

level: middleimportance: must knowfreq 60%

answer

  1. immutable files, not rows
  2. copies at every layer
  3. old snapshots still hold it
  4. backups, extracts, derived tables
  5. replays bring it back

basics

~20 s

Lake data sits in immutable files copied through raw, cleaned and curated layers, kept in old snapshots, backups and extracts. Deleting a person means rewriting files everywhere they appear, expiring old versions, and stopping replays from reloading them.

solid answer

~40 s

In an application database the person is a few rows, updated in place, and a `DELETE` removes them. In a lake, the same person is spread across **immutable files**: a delete means **rewriting** each file that holds their rows (or writing delete markers that later compaction applies). They also appear in **every layer** — raw landing, cleaned, curated marts — and in **derived tables**, **extracts** and **feature sets**. Table formats keep **old snapshots** for time travel, so the rows stay readable until those snapshots are expired and their files physically removed; **backups** hold them longer still. Finally, if raw data is ever **replayed or backfilled**, the person comes back. So erasure in a lake is a workflow: find every location, rewrite, expire, purge, and keep a suppression list so reprocessing does not reintroduce them.

go deeper

for a junior

Know that lakes keep data in files and copies across layers, so deleting one person touches many places.

for a middle

Explain rewriting immutable files, snapshot expiry and why replays can reintroduce deleted data.

for a senior

Design the erasure workflow with location via tags and lineage, batching, purge, suppression and evidence.

for a principal

Decide the platform rules, such as partitioning by subject, raw-zone retention and backup handling, that keep erasure affordable at scale.

## The contrast | Aspect | Application database | Data lake or lakehouse | |---|---|---| | Storage unit | rows and pages updated in place | large immutable files in object storage | | Delete operation | `DELETE` removes rows | rewrite affected files, or write delete markers purged later | | Copies | usually one primary plus replicas | raw, cleaned and curated layers, marts, extracts, feature sets | | Old versions | overwritten, then reclaimed | kept as snapshots for time travel until expired | | Backups | database backups | object versions, snapshots, backup copies, archive tiers | | Reloading | not usual | replays and backfills from raw are routine | ## Where the person hides 1. **Raw landing zones**, often kept "forever" so pipelines can be rebuilt. 2. **Every downstream layer**: cleaned tables, dimensional models, aggregates small enough to single someone out. 3. **Derived artefacts**: machine-learning feature tables, extracts sent to other teams, cached dashboard results. 4. **Old table versions**: a lakehouse table format keeps prior snapshots so readers can query the past; a deleted row is still in files referenced by those snapshots. 5. **Object-store versioning and backups**: deleted files may survive as previous object versions. 6. **Logs and free text**: event payloads and support notes with the person's details embedded. ## What deletion actually takes - **Locate**: use classification tags and lineage to find every table holding the person's identifiers, not only the obvious customer table. - **Rewrite**: run deletes per table; the table format rewrites files (copy-on-write) or records delete files (merge-on-read) that later compaction applies. - **Expire and purge**: expire snapshots older than the deletion and remove unreferenced files, otherwise time travel still returns the rows. - **Handle backups**: either purge them, or record the deletion so it is re-applied if a backup is ever restored. - **Suppress reintroduction**: keep a list of erased identifiers (stored as keyed hashes) and filter it in ingestion and backfills. - **Record evidence**: which locations were processed, when, and with what result. ## Batching Rewriting large files for each request is expensive, so platforms usually **batch** erasure requests and run them on a schedule within the deadline the applicable regime sets, often partitioning tables so that one person's rows are concentrated rather than scattered across every file. ## Why interviewers ask it It is the classic data-platform privacy question. A good answer contrasts **update-in-place** with **immutable files**, lists the **hidden copies** (layers, snapshots, backups, extracts), and knows that deleting is not finished until old versions are **physically purged** and **replays cannot restore** the data.

  • Why does a delete in a lakehouse table not immediately remove the data from storage?
    The table format writes new files or delete markers and commits a new snapshot, but earlier snapshots still reference the old files. Until those snapshots are expired and unreferenced files removed, time travel or a direct file read can still return the rows.
  • How do you stop a backfill from reloading erased customers?
    Keep a suppression list of erased identifiers, stored as keyed hashes rather than in clear, and apply it as a filter at ingestion and in any replay or backfill job, so raw data that still contains them is dropped on the way in.
  • Why batch erasure requests instead of processing each one immediately?
    Each request forces file rewrites across many tables; batching amortises that cost so one rewrite handles many people. The batch interval must stay well inside the deadline the applicable regime sets.

saying these in an interview costs you the question

  • Believing a DELETE statement on the main table completes the erasure
  • Forgetting snapshots kept for time travel
  • Ignoring raw landing zones and extracts sent to other teams
  • Letting backfills from raw data reintroduce erased people
  • Storing the suppression list of erased people in clear text