skip to content

Log Cleaner and Compaction Internals

The background cleaner threads that rewrite compacted segments: dirty ratio, the offset map, and tombstone removal. Interviewers ask it when they want to know why compaction has not happened yet on a topic.

part ofApache Kafkaoverview, primer and where to startread it →
on this pageshow

questions

5

What is the Kafka log cleaner, and how do you enable it for a topic?

level: juniorimportance: must knowfreq 60%

answer

  1. cleanup.policy=compact
  2. log.cleaner.enable default true since 0.9.4
  3. keep latest value per key
  4. background CleanerThread, never blocks clients
  5. operates on the tail, not active segment

basics

~20 s

The log cleaner is a background process that runs compaction: it scans a topic's log and keeps only the latest record per key, deleting older duplicates. You enable it by setting the topic config cleanup.policy=compact.

solid answer

~40 s

The log cleaner is the broker subsystem that performs log compaction. Instead of deleting whole segments by age/size (the delete policy), compaction retains at least the most recent value for every distinct message key and removes superseded older records. You turn it on per topic with cleanup.policy=compact (or compact,delete to combine both). The broker also needs log.cleaner.enable=true, which has defaulted to true since Kafka 0.9.4, so in practice setting cleanup.policy on the topic is enough. Dedicated cleaner threads (log.cleaner.threads, default 1) do the work asynchronously in the background; they never block producers or consumers. Compaction operates only on the 'tail' of the log — segments below the active segment — so the latest writes are always preserved verbatim.

go deeper

for a junior

Know cleanup.policy=compact keeps the latest value per key, done by a background cleaner.

for a middle

Distinguish delete vs compact vs compact,delete; know cleaner threads are async and operate on the tail.

for a senior

Explain the active-segment boundary, null-key handling, and combining policies for bounded changelogs.

for a principal

Reason about when to model state as a compacted topic vs an external store, and capacity-plan cleaner threads against write rate.

## What problem compaction solves Kafka stores each partition as an append-only **log** split into **segment** files. The default retention strategy, `cleanup.policy=delete`, removes *entire old segments* once they exceed `retention.ms` (time) or `retention.bytes` (size). That is fine for event streams but wrong for **changelog / state** topics where you want to keep the *latest value per key forever* (e.g., a user-profile table keyed by user id). **Log compaction** is the alternative: for every distinct **key**, Kafka guarantees it retains *at least* the most recent record; older records with the same key become eligible for removal. The result is a log that still contains the full *current state* of every key while shedding the history. ## Who does the work: the log cleaner The **log cleaner** is the broker component that performs compaction. It is a pool of background threads: - `log.cleaner.enable` (broker, default **true** since 0.9.4) — master switch. - `log.cleaner.threads` (broker, default **1**) — number of `CleanerThread` workers. These threads run continuously and asynchronously. They do **not** block produce or fetch requests; compaction is invisible to clients except that duplicate keys eventually disappear. ## How you turn it on for a topic Compaction is selected **per topic** via: ``` cleanup.policy=compact ``` Values are `delete` (default), `compact`, or both `compact,delete` (compact, *and* also drop segments past retention). Example: ``` kafka-topics.sh --bootstrap-server localhost:9092 \ --create --topic user-profiles \ --config cleanup.policy=compact ``` ## Key boundaries - Compaction works on the **tail** (everything before the **active** segment). The active segment — the one currently being appended — is never compacted, so the newest writes are always present. - Compaction is **key-based**, so every record must have a non-null key; null-key records cannot be compacted meaningfully. - It removes *duplicates*, not all old data — the surviving record per key stays, regardless of age (unless `compact,delete` plus retention also applies).

  • What does cleanup.policy=compact,delete do?
    It applies compaction (latest value per key) AND still enforces retention.ms/retention.bytes, so old keys can also age out by time/size — useful for changelogs you don't want to grow unbounded.
  • Does compaction guarantee zero duplicates for a key?
    No. It guarantees the latest value survives, but duplicates can remain in the uncompacted head and in the active segment, so consumers must be idempotent / take the last value seen per key.

saying these in an interview costs you the question

  • Saying compaction deletes records by age like the delete policy — it is key-based, not time-based.
  • Claiming you must set log.cleaner.enable=true manually — it has defaulted to true since 0.9.4.
  • Saying compaction blocks producers/consumers — cleaner threads run in the background.
  • Thinking compaction guarantees exactly one record per key everywhere in the log.

context

open as a page

How does the cleaner handle tombstones, and what role does delete.retention.ms play?

level: middleimportance: must knowfreq 50%

basics

~20 s

A tombstone is a record with a real key but a null value, signaling 'this key is deleted'. The cleaner keeps tombstones around for delete.retention.ms (default 24h) so consumers can observe the deletion, then removes them in a later pass.

open as a page

How does the cleaner decide which log to compact next, and what is min.cleanable.dirty.ratio?

level: seniorimportance: must knowfreq 55%

basics

~20 s

Each cleaner thread picks the log with the highest 'dirty ratio' — the fraction of the cleanable log that hasn't been compacted yet. A log only becomes eligible once that ratio reaches min.cleanable.dirty.ratio (default 0.5).

open as a page

Explain the head vs tail of a compacted log and why the active segment is never compacted.

level: seniorimportance: should knowfreq 35%

basics

~20 s

The tail is the older, already-compacted part of the log (at most one record per key). The head is newer records appended since the last clean, which may still contain duplicates. The active segment — the one being written to — is always part of the head and is never compacted.

open as a page

Mechanically, how does a single compaction pass deduplicate keys, and what is the offset map?

level: seniorimportance: should knowfreq 40%

basics

~20 s

The cleaner makes two passes over the dirty head. First it builds an in-memory offset map: key -> highest offset seen. Then it recopies the log, keeping each record only if its offset equals the map's value for that key, and merges results into new compacted segments.

open as a page