What is a tombstone in a compacted Kafka topic, and how does delete.retention.ms govern its lifecycle?
answer
- tombstone = null value = delete marker
- delete.retention.ms default 24h
- grace window so lagging consumers see the delete
- only meaningful under compact
- active segment tombstone not yet processed
basics
~20 sA tombstone is a record with a null value that marks a key for deletion. After compaction, the tombstone removes prior values for that key, and the tombstone itself is retained for at least delete.retention.ms before being purged, so lagging consumers can still observe the delete.
solid answer
~50 sIn a compacted topic, a record with key=K and value=null is a tombstone: it signals 'key K is deleted'. During compaction the cleaner removes all earlier records for K, and after a grace period removes the tombstone too. That grace period is delete.retention.ms (default 86400000 ms = 24h). The window matters because consumers (and downstream stores like KTables) must see the tombstone to delete their local copy of K; if it vanished too quickly, a slow consumer that already read the old value would never learn of the deletion. The clock for tombstone removal is based on when the segment containing the tombstone becomes eligible after a cleaner pass — specifically, a tombstone is retained until at least delete.retention.ms after it's been moved to the cleaned portion of the log. Tombstones only have this special meaning under cleanup.policy containing compact.
go deeper
Recognize that a null-value record is a tombstone meaning 'delete this key'.
Explain delete.retention.ms, its 24h default, and why the grace window protects lagging consumers.
Reason about consumer-lag vs window sizing and the divergence risk for materialized state.
Design hard-delete/GDPR compliance flows over compacted changelogs and set org-wide tombstone retention policy.
## What a tombstone is In Kafka a **tombstone** is simply a record whose **value is null** (the key is still present). In a **compacted** topic this null value is interpreted as a **delete marker** for that key. After compaction processes it, prior values for the key are gone and the tombstone marks the key as removed. ## Why tombstones can't be deleted immediately Compaction is the only mechanism to delete a *specific key* from a compacted topic (which otherwise has no time-based expiry). But there's a race: imagine a consumer that has read the old value of key K but hasn't yet reached the tombstone. If the cleaner deleted the tombstone right away, that consumer's local view (e.g., a Kafka Streams state store, a cache, a materialized table) would keep K forever — it would never see the delete. ## delete.retention.ms To prevent that, the tombstone is kept around for a grace window controlled by **`delete.retention.ms`** (default **86400000 ms = 24 hours**). The semantics: after the tombstone has been retained through a compaction pass, it remains visible for at least delete.retention.ms so any consumer caught up within that window will observe the delete before the marker disappears. After the window, a subsequent cleaner pass can physically remove the tombstone. ## Practical tuning - If consumers can lag for days, **raise** delete.retention.ms so deletes aren't missed. - If you want deletes to physically vanish quickly (e.g., GDPR-style hard delete from the changelog), **lower** it — but only after confirming all consumers are well within the window. - delete.retention.ms has no effect unless cleanup.policy includes **compact** (and a tombstone in a pure delete topic is just a normal null-value record with no special semantics). ## Edge cases - A tombstone in the **active segment** is not yet processed; only after the segment rolls and becomes part of the dirty log can the cleaner act on it. - Producing a new non-null value for K after a tombstone simply resurrects the key — compaction will then keep that newer value. - Tombstones still consume offset space until removed; very high churn of deletes can keep the dirty ratio elevated.
- What is the default value of delete.retention.ms and what unit?86400000 milliseconds, i.e. 24 hours.
- What happens to a downstream KTable if delete.retention.ms is too small and a consumer lags beyond it?The consumer may never see the tombstone, so its materialized state retains the deleted key indefinitely — a ghost record. The local store and the compacted source diverge.
saying these in an interview costs you the question
- Saying a tombstone has an empty (zero-length) value — it must be null; an empty byte array is a normal record.
- Claiming tombstones are deleted instantly during compaction.
- Believing delete.retention.ms affects normal (non-null) records — it only governs tombstone purge timing.
- Saying tombstones work in a delete-only topic — null values have no delete semantics there.