How does the cleaner handle tombstones, and what role does delete.retention.ms play?
answer
- tombstone = key present, value null
- delete.retention.ms default 24h
- kept so consumers observe the delete
- removed by a later pass after the window
- too low -> lagging consumer misses the delete
basics
~20 sA tombstone is a record with a real key but a null value, signaling 'this key is deleted'. The cleaner keeps tombstones around for delete.retention.ms (default 24h) so consumers can observe the deletion, then removes them in a later pass.
solid answer
~50 sIn a compacted topic, deleting a key is done by producing a record with that key and a null value — a tombstone. During compaction the cleaner first removes older records that share the tombstone's key (the value is gone), but it deliberately retains the tombstone itself for delete.retention.ms (default 86400000 ms = 24h). This window guarantees a consumer that is briefly offline or replaying still sees the null and learns the key was deleted, rather than the deletion vanishing silently. Only after a tombstone has been in the clean log past delete.retention.ms does a subsequent compaction pass actually drop it, finally reclaiming the key. The timer is measured from when the tombstone becomes part of the cleanable tail (after a clean), not from produce time — so very short delete.retention.ms combined with fast cleaning can race a slow consumer and lose the delete signal.
go deeper
Know a tombstone is a null-value record meaning 'delete this key' and it sticks around for a while.
Explain delete.retention.ms (default 24h) and why the tombstone is kept before being dropped.
Detail the two-phase removal, the consumer-observability contract, and the race from too-low retention.
Set delete.retention.ms against worst-case consumer lag/downtime SLAs and reason about deletion correctness across the consumer fleet.
## What a tombstone is In a **compacted** topic, you cannot delete a key by sending a delete command — you express deletion in the data itself. A **tombstone** is a record with a normal **key** but a **null value**. Semantically it means 'the latest state of this key is: gone'. Downstream state stores (e.g., Kafka Streams, a database materialized from the topic) interpret a null value as 'remove this key'. ## The two-phase removal Compaction handles tombstones in two stages: 1. **Collapse the key's history:** like any key, the tombstone supersedes older records with the same key, so those older values are removed in the normal compaction pass. 2. **Retain the tombstone temporarily:** unlike a normal record, the cleaner does **not** immediately drop the tombstone, because if it did, a consumer that hadn't yet read the deletion would never learn the key was removed and would keep stale state forever. ## delete.retention.ms `delete.retention.ms` (topic-level, default **86,400,000 ms = 24 hours**) is the minimum time a tombstone is preserved in the **clean** portion of the log after compaction. The contract: any consumer that reads at least once every `delete.retention.ms` is guaranteed to observe every tombstone and therefore every deletion. The clock effectively starts when the tombstone enters the cleanable tail (i.e., it has been carried through a compaction). After it has lingered past `delete.retention.ms`, the **next** compaction pass is allowed to physically remove it, finally reclaiming the key entirely from the log. ## The race condition If you set `delete.retention.ms` too low, a slow or lagging consumer can miss the tombstone: the cleaner removes it before the consumer reads that far, so the consumer's view of the key is stuck at its old value (a 'zombie' record). This is why the default is a full day — it tolerates consumer downtime, rebalances, and catch-up. ## Edge cases & related configs - A tombstone has a key but null value; a record with a **null key** cannot be compacted at all and is a config/usage error on compacted topics. - `min.compaction.lag.ms` can additionally delay when a tombstone (or any record) first becomes compactable. - With `cleanup.policy=compact,delete`, normal retention.ms can also age out tombstones independently. ## Practical guidance Set `delete.retention.ms` to comfortably exceed your worst-case consumer downtime / lag, so deletions are never silently lost. Lower it only when you specifically need keys to disappear faster and you control consumer freshness.
- Why not delete a tombstone immediately once older records for its key are gone?Because a consumer that hasn't yet read the deletion would never see the null value and would retain stale state forever. delete.retention.ms keeps the tombstone long enough for all reasonably-fresh consumers to observe it.
- What happens on a compacted topic if you produce a record with a null key?It cannot be compacted (compaction is key-based) and is effectively a misuse; the broker rejects or warns depending on version. Tombstones must have a non-null key and a null value.
saying these in an interview costs you the question
- Saying a tombstone has a null key — it has a real key and a null value.
- Claiming tombstones are removed immediately during compaction — they're retained for delete.retention.ms.
- Confusing delete.retention.ms (tombstone lifetime) with retention.ms (segment age-out for the delete policy).
- Thinking the retention timer starts at produce time rather than when the tombstone enters the cleaned tail.