skip to content

What is the Kafka log cleaner, and how do you enable it for a topic?

level: juniorimportance: must knowfreq 60%

answer

  1. cleanup.policy=compact
  2. log.cleaner.enable default true since 0.9.4
  3. keep latest value per key
  4. background CleanerThread, never blocks clients
  5. operates on the tail, not active segment

basics

~20 s

The log cleaner is a background process that runs compaction: it scans a topic's log and keeps only the latest record per key, deleting older duplicates. You enable it by setting the topic config cleanup.policy=compact.

solid answer

~40 s

The log cleaner is the broker subsystem that performs log compaction. Instead of deleting whole segments by age/size (the delete policy), compaction retains at least the most recent value for every distinct message key and removes superseded older records. You turn it on per topic with cleanup.policy=compact (or compact,delete to combine both). The broker also needs log.cleaner.enable=true, which has defaulted to true since Kafka 0.9.4, so in practice setting cleanup.policy on the topic is enough. Dedicated cleaner threads (log.cleaner.threads, default 1) do the work asynchronously in the background; they never block producers or consumers. Compaction operates only on the 'tail' of the log — segments below the active segment — so the latest writes are always preserved verbatim.

go deeper

for a junior

Know cleanup.policy=compact keeps the latest value per key, done by a background cleaner.

for a middle

Distinguish delete vs compact vs compact,delete; know cleaner threads are async and operate on the tail.

for a senior

Explain the active-segment boundary, null-key handling, and combining policies for bounded changelogs.

for a principal

Reason about when to model state as a compacted topic vs an external store, and capacity-plan cleaner threads against write rate.

## What problem compaction solves Kafka stores each partition as an append-only **log** split into **segment** files. The default retention strategy, `cleanup.policy=delete`, removes *entire old segments* once they exceed `retention.ms` (time) or `retention.bytes` (size). That is fine for event streams but wrong for **changelog / state** topics where you want to keep the *latest value per key forever* (e.g., a user-profile table keyed by user id). **Log compaction** is the alternative: for every distinct **key**, Kafka guarantees it retains *at least* the most recent record; older records with the same key become eligible for removal. The result is a log that still contains the full *current state* of every key while shedding the history. ## Who does the work: the log cleaner The **log cleaner** is the broker component that performs compaction. It is a pool of background threads: - `log.cleaner.enable` (broker, default **true** since 0.9.4) — master switch. - `log.cleaner.threads` (broker, default **1**) — number of `CleanerThread` workers. These threads run continuously and asynchronously. They do **not** block produce or fetch requests; compaction is invisible to clients except that duplicate keys eventually disappear. ## How you turn it on for a topic Compaction is selected **per topic** via: ``` cleanup.policy=compact ``` Values are `delete` (default), `compact`, or both `compact,delete` (compact, *and* also drop segments past retention). Example: ``` kafka-topics.sh --bootstrap-server localhost:9092 \ --create --topic user-profiles \ --config cleanup.policy=compact ``` ## Key boundaries - Compaction works on the **tail** (everything before the **active** segment). The active segment — the one currently being appended — is never compacted, so the newest writes are always present. - Compaction is **key-based**, so every record must have a non-null key; null-key records cannot be compacted meaningfully. - It removes *duplicates*, not all old data — the surviving record per key stays, regardless of age (unless `compact,delete` plus retention also applies).

  • What does cleanup.policy=compact,delete do?
    It applies compaction (latest value per key) AND still enforces retention.ms/retention.bytes, so old keys can also age out by time/size — useful for changelogs you don't want to grow unbounded.
  • Does compaction guarantee zero duplicates for a key?
    No. It guarantees the latest value survives, but duplicates can remain in the uncompacted head and in the active segment, so consumers must be idempotent / take the last value seen per key.

saying these in an interview costs you the question

  • Saying compaction deletes records by age like the delete policy — it is key-based, not time-based.
  • Claiming you must set log.cleaner.enable=true manually — it has defaulted to true since 0.9.4.
  • Saying compaction blocks producers/consumers — cleaner threads run in the background.
  • Thinking compaction guarantees exactly one record per key everywhere in the log.

context