skip to content

What does cleanup.policy=compact do to a Kafka topic, and how is it different from the default delete policy?

level: juniorimportance: must knowfreq 70%

answer

  1. latest value per key vs age-out whole segments
  2. compact = changelog / snapshot
  3. delete = retention.ms / retention.bytes
  4. per-partition, dirty tail only
  5. __consumer_offsets, KTable changelog

basics

~10 s

Compaction keeps at least the latest value for each message key, deleting older values for that key. The default delete policy instead drops whole segments once they exceed retention.ms or retention.bytes, regardless of key.

solid answer

~40 s

cleanup.policy=delete (the default) removes data by age/size: once a log segment is older than retention.ms or the partition exceeds retention.bytes, the whole segment is deleted. cleanup.policy=compact keeps the topic indefinitely but guarantees that for every key, the most recent value survives; older records with the same key are garbage-collected by a background log cleaner. This makes a compacted topic behave like a changelog or key/value snapshot — replaying it reconstructs the latest state per key. Compaction operates per partition and only on the 'dirty' (uncompacted) tail, never on the active segment. It's the mechanism behind Kafka internal topics like __consumer_offsets and Kafka Streams KTable changelogs. You can also combine both: cleanup.policy=compact,delete applies retention on top of compaction.

go deeper

for a junior

Know the one-liner: compact keeps the newest value per key; delete ages out old data.

for a middle

Explain per-partition operation, dirty vs active segment, and a concrete use case like __consumer_offsets.

for a senior

Discuss convergence semantics, why offsets are preserved, and combining compact+delete.

for a principal

Frame compaction as the substrate for event-sourcing/changelog patterns and reason about state-recovery guarantees for stream processors.

## What a Kafka topic stores A Kafka **topic** is split into **partitions**; each partition is an append-only **log** of records. Every record has a **key**, a **value**, and an **offset** (its position in the log). Physically each partition is a series of **segment** files. ## The default: cleanup.policy=delete By default Kafka retains data by time and size. `retention.ms` (default 7 days) and `retention.bytes` decide when an *entire segment* is eligible for deletion. Deletion is coarse: it removes the oldest segments wholesale, with no regard for what keys they contain. Old data simply ages out. ## Compaction: cleanup.policy=compact With compaction the contract changes. Kafka promises that **for each key, the record with the highest offset is retained**. Older records for the same key may be removed by a background thread called the **log cleaner**. The result is that the log eventually converges to (at most) one value per key — a snapshot of current state — while still being an ordered log you can replay. Key points: - Compaction is **per partition**, so keys must be partitioned consistently for it to make sense. - The cleaner only compacts the **dirty** portion (records appended since the last clean); the **active segment** (the one currently being written) is never compacted. - Records with a **null value** are **tombstones**: they signal deletion of a key and are themselves removed after `delete.retention.ms`. - Compaction does **not** guarantee instant removal of old values — it is asynchronous and governed by thresholds like `min.cleanable.dirty.ratio`. ## Why it exists Compacted topics are ideal for **changelogs**: __consumer_offsets, Kafka Streams state-store changelogs, and Debezium/CDC snapshots all rely on 'latest value per key'. A new consumer can bootstrap full state by reading the compacted log from the beginning. ## Combining policies `cleanup.policy=compact,delete` runs both: compaction keeps the latest per key, *and* retention.ms/bytes can still drop very old segments, bounding total size.

  • Name a real Kafka topic that uses compaction and why.
    __consumer_offsets stores the latest committed offset per (group, topic, partition) key; compaction keeps only the newest commit so the topic doesn't grow unbounded while still allowing full state recovery on broker restart.
  • Does compaction guarantee there is exactly one record per key after it runs?
    No. It guarantees the latest value survives, but duplicates can linger in the dirty/active portion or until the cleaner runs; the active segment is never compacted. It converges over time, not instantly.

saying these in an interview costs you the question

  • Saying compaction deletes data by time like retention — it removes superseded values by key, not by age.
  • Claiming compaction guarantees exactly one record per key immediately.
  • Thinking compaction reorders or rewrites offsets (offsets are preserved; some just disappear).

context