skip to content

What does Kafka's cleanup.policy control, and how do delete and compact differ? How do retention.ms and retention.bytes fit in?

level: juniorimportance: must knowfreq 70%

answer

  1. delete = age/size; compact = latest-per-key
  2. retention.bytes is PER PARTITION
  3. segment granularity → active segment never deleted
  4. tombstone kept for delete.retention.ms
  5. compact,delete combines both

basics

~10 s

cleanup.policy sets how old data is removed. delete drops whole log segments once they exceed retention.ms (age) or retention.bytes (size). compact keeps the latest value per key forever. retention.ms/bytes only apply to delete.

solid answer

~40 s

cleanup.policy is a per-topic setting choosing how Kafka reclaims log space. With delete (the default), Kafka removes entire log segments once the data in them exceeds the time limit retention.ms or the size limit retention.bytes (whichever triggers first). With compact, Kafka instead runs the log cleaner to retain only the most recent record for each key, so a topic becomes a changelog/snapshot of latest values. You can combine them as compact,delete, which compacts but still applies retention limits. Retention applies at segment granularity, so the active segment is never deleted and data can outlive retention.ms until the segment rolls (segment.ms/segment.bytes). retention.bytes is per partition, not per topic. Tombstones (null-value records) in compacted topics are kept for delete.retention.ms before final removal so consumers can observe deletes.

go deeper

for a junior

Know the two policies: delete removes old data by time/size; compact keeps the latest value per key.

for a middle

Explain segment-level eviction, retention.bytes being per-partition, and compact,delete combination.

for a senior

Reason about why data outlives retention.ms, tombstone handling, and choosing policy for changelog vs event topics.

for a principal

Set org-wide retention/compaction standards, capacity-plan disk from partitions×retention.bytes, and design compacted state topics for recovery.

## What is a Kafka topic log? A Kafka topic is split into partitions; each partition is an append-only **log** persisted to disk as a series of **segment** files. New records always append to the *active* (newest) segment. A segment 'rolls' (closes, and a new one opens) when it reaches `segment.bytes` (default 1 GiB) or `segment.ms` (default 7 days). Only **closed** segments are eligible for cleanup — this is why data can survive past its nominal retention until the segment rolls. ## cleanup.policy The per-topic config `cleanup.policy` decides how Kafka reclaims space: - **`delete`** (default): time/size-based eviction. Kafka deletes whole closed segments once their data violates a limit. - `retention.ms` — maximum age of data (default 604800000 = 7 days). A segment is deletable once its newest record is older than this. - `retention.bytes` — maximum size **per partition** (default -1 = unlimited). When a partition exceeds it, oldest segments are dropped. Note: per partition, NOT per topic — total topic disk is roughly `retention.bytes × partitions`. - Whichever limit is hit first triggers deletion. - **`compact`**: key-based retention. A background **log cleaner** thread rewrites the log keeping only the **latest record per key**. Older values for the same key are removed. The result is a 'latest-state snapshot' — ideal for changelogs, KTable state, `__consumer_offsets`. retention.ms/bytes do NOT bound a purely compacted topic; data for a key lives until a newer record for that key arrives. - **Tombstones**: a record with a null value marks a key for deletion. After compaction the tombstone itself is retained for `delete.retention.ms` (default 24h) so lagging consumers can see the delete, then it is removed. - `min.cleanable.dirty.ratio`, `min.compaction.lag.ms`, `max.compaction.lag.ms` tune when/how aggressively the cleaner runs. - **`compact,delete`** (both): compaction PLUS time/size retention — keep latest-per-key but also evict by age/size. Useful when a compacted changelog must not grow unbounded. ## Edge cases - The active segment is never cleaned, so a low-traffic topic can hold data far past `retention.ms`. - `retention.bytes` being per-partition surprises people sizing clusters. - Setting `retention.ms` does NOT immediately delete — cleanup runs on a check interval (`log.retention.check.interval.ms`). - Lowering retention.ms can purge data on the next check; treat config changes as potentially destructive.

  • Why might data persist longer than retention.ms?
    Cleanup operates at segment granularity and the active segment is never deleted. Data sits in the active segment until it rolls (segment.ms/segment.bytes), so low-throughput topics can exceed retention.ms.
  • Is retention.bytes per topic or per partition?
    Per partition. Total topic footprint is approximately retention.bytes multiplied by the partition count.

saying these in an interview costs you the question

  • Saying retention.bytes is per topic.
  • Claiming retention.ms deletes individual records immediately (it's segment-level and runs on a check interval).
  • Thinking compact bounds size by time — pure compaction has no time/size cap unless you add delete.
  • Confusing tombstones with normal records — they have null value and are governed by delete.retention.ms.

context