skip to content

A GDPR erasure request requires deleting all of a user's data from a compacted Kafka topic. How do tombstones and retention/compaction make this work, and what are the caveats?

level: seniorimportance: should knowfreq 45%

answer

  1. Tombstone = key + null value
  2. compact policy required, not delete
  3. delete.retention.ms keeps tombstone ~24h
  4. Eventually consistent; closed segments only
  5. Crypto-shredding when physical delete is hard

basics

~20 s

On a compacted topic, produce a tombstone — a record with the user's key and a null value. Log compaction eventually removes all prior records for that key and then drops the tombstone after delete.retention.ms. Deletion is asynchronous, not immediate, and only works per key.

solid answer

~50 s

For a key-compacted topic (cleanup.policy=compact), GDPR erasure is done with a tombstone: a record carrying the subject's key and a null value. Log compaction guarantees the latest value per key is retained; a null value signals deletion, so after the next compaction pass all earlier records for that key are physically removed, and the tombstone itself is purged after delete.retention.ms. Caveats: this is eventually consistent — compaction runs on closed (non-active) segments per dirty-ratio thresholds (min.cleanable.dirty.ratio, min.compaction.lag.ms), so deletion is not immediate and you can't bound it tightly. It only works if data is keyed by the subject; data spread across many keys or topics needs a deletion strategy per stream. Downstream consumers, changelog topics, and tiered/archived storage must honor tombstones too. For time-retained topics (cleanup.policy=delete), retention.ms ages data out but can't target one user. Crypto-shredding (destroying that user's encryption key) is the common complement when physical deletion can't be guaranteed.

go deeper

for a junior

Know that a tombstone is a record with a null value used to delete a key on a compacted topic.

for a middle

Explain compact vs delete policy, delete.retention.ms, and that deletion is eventual not instant.

for a senior

Reason about compaction thresholds, segment rolling, per-key limitation, and downstream propagation gaps.

for a principal

Architect org-wide GDPR erasure: keying strategy, max.compaction.lag bounds, crypto-shredding for unreachable copies, and propagation to all derived stores.

## The problem GDPR's right to erasure means you must be able to delete a person's data. Kafka logs are **append-only and immutable** — you cannot edit or delete an individual record in place. Two mechanisms make erasure possible: **retention** and **compaction with tombstones**. ## Retention vs compaction Every topic has a `cleanup.policy`: - **`delete`** (default): records are removed once they exceed **`retention.ms`** (age) or **`retention.bytes`** (size). This ages out *all* old data on a schedule but **cannot target a single user**. - **`compact`**: Kafka keeps **at least the latest value for each message key**; older values for the same key are eligible for removal. Used for changelog/state topics where you want the current state per key. ## Tombstones In a **compacted** topic, a **tombstone** is a record with a **non-null key and a `null` value**. It means 'this key is deleted.' During the next compaction pass: 1. All earlier records for that key are removed. 2. The tombstone is retained for a grace window — **`delete.retention.ms`** (default 24h) — so consumers that are behind still observe the deletion. 3. After that window, the tombstone itself is purged. So to erase a user from a compacted topic keyed by user id: **produce a tombstone with their key**. ## Why deletion is not immediate Compaction is a **background process** with knobs: - It only acts on **non-active (closed) segments** — the segment currently being written is never compacted, so very recent data lingers until the segment rolls (`segment.ms` / `segment.bytes`). - The **log cleaner** triggers based on **`min.cleanable.dirty.ratio`** (how much of the log is 'dirty'/uncompacted) and **`min.compaction.lag.ms`** / **`max.compaction.lag.ms`** (minimum/maximum age before a record is compacted). Result: erasure is **eventually consistent**. You generally **cannot promise deletion within seconds**; you can tune `max.compaction.lag.ms` to bound the worst case. ## Caveats and gaps - **Keyed-by-subject required:** tombstones erase **per key**. If a user's data is scattered under many keys, embedded in events keyed by something else, or in multiple topics, a single tombstone won't reach it — you need a per-topic deletion design. - **Downstream propagation:** consumers, **Streams changelog/state stores**, materialized views, search indexes, data lakes, and **tiered storage** must also process the tombstone or honor the deletion — otherwise copies survive. - **Compaction must be on:** a `delete`-policy topic ignores tombstones for targeted erasure; tombstones only have erasure semantics under `compact`. - **Crypto-shredding complement:** when guaranteed physical deletion is hard (backups, tiered/object storage, replicas), encrypt each subject's data with a per-subject key and, on erasure, **destroy the key** — the ciphertext becomes permanently unreadable, satisfying erasure without chasing every byte. ## Summary Tombstone + compaction = targeted, eventual, per-key deletion. Retention = bulk, time-based aging. Crypto-shredding = make-unreadable when delete-everywhere isn't feasible.

  • Why can't you guarantee a user's record is deleted within a fixed short time after sending a tombstone?
    Compaction runs in the background only on closed (non-active) segments and is gated by thresholds like min.cleanable.dirty.ratio and min/max.compaction.lag.ms. The active segment isn't compacted until it rolls. You can bound the worst case with max.compaction.lag.ms but not make it instant.
  • What is crypto-shredding and when do you use it for GDPR?
    Encrypt each data subject's records with a per-subject key; on an erasure request, delete that key so the ciphertext can never be decrypted again. You use it when physically deleting every copy (backups, tiered storage, replicas) is impractical — the data is rendered permanently unreadable instead.

saying these in an interview costs you the question

  • Saying a tombstone deletes the data immediately/synchronously
  • Thinking tombstones work on cleanup.policy=delete topics
  • Believing retention.ms can target a single user
  • Forgetting downstream copies (changelogs, indexes, tiered storage) must also honor the deletion
  • A tombstone has a value (it must be null)

context