skip to content

How does the cleaner decide which log to compact next, and what is min.cleanable.dirty.ratio?

level: seniorimportance: must knowfreq 55%

answer

  1. dirtyBytes / (dirtyBytes + cleanBytes)
  2. default 0.5 eligibility threshold
  3. greedy: highest dirty ratio first
  4. clean tail vs dirty head + checkpoint offset
  5. max.compaction.lag.ms forces time-based cleaning

basics

~20 s

Each cleaner thread picks the log with the highest 'dirty ratio' — the fraction of the cleanable log that hasn't been compacted yet. A log only becomes eligible once that ratio reaches min.cleanable.dirty.ratio (default 0.5).

solid answer

~40 s

The cleaner divides each partition's log into a clean 'tail' (already compacted) and a dirty 'head' (appended since the last clean). The dirty ratio = dirtyBytes / (dirtyBytes + cleanBytes) over the cleanable region. On each cycle a cleaner thread selects the *eligible* log with the highest dirty ratio — a greedy 'most dirty first' policy that maximizes bytes reclaimed per pass. A log is eligible only when its ratio is at least min.cleanable.dirty.ratio (default 0.5): below that, compacting wastes I/O for little gain. Lowering it (e.g., 0.1) makes compaction more aggressive and timely at the cost of more disk I/O and CPU; raising it lets duplicates accumulate longer but reduces overhead. Kafka also exposes min.cleanable.dirty.ratio per topic, plus time-based forcing via max.compaction.lag.ms and a minimum delay via min.compaction.lag.ms.

go deeper

for a junior

Know there's a 'dirty ratio' default 0.5 that decides when a topic gets compacted.

for a middle

Define dirty ratio as dirty/(dirty+clean) and know the cleaner picks the most-dirty log.

for a senior

Explain greedy selection, eligibility threshold, and the time-based forcing configs and their trade-offs.

for a principal

Capacity-plan cleaner I/O across many topics and choose ratio + lag settings against compliance and storage SLAs.

## Clean tail vs dirty head The cleaner tracks, per partition, a **cleaner checkpoint offset**: everything below it has already been compacted (the **clean tail**); everything from there up to the start of the active segment is the **dirty head** — records appended since the last clean that may contain duplicate keys. ## The dirty ratio The cleaner ranks logs by **dirty ratio**: ``` dirtyRatio = dirtyBytes / (dirtyBytes + cleanBytes) ``` where `dirtyBytes` is the size of the uncompacted head and `cleanBytes` is the size of the already-compacted tail (both restricted to the cleanable region, i.e., excluding the active segment and anything held back by lag configs). A ratio near 1.0 means almost the whole cleanable log is unprocessed duplicates; near 0 means it was recently cleaned. ## Selection: greedy most-dirty-first On each iteration a free `CleanerThread` scans all compacted logs assigned to it and **picks the one with the highest dirty ratio** among those that are *eligible*. This greedy strategy reclaims the most space per pass and prevents any one log from starving — a continuously written log climbs to the top of the ranking. ## min.cleanable.dirty.ratio `min.cleanable.dirty.ratio` (topic-level, default **0.5**) is the **eligibility threshold**: a log is not even considered until its dirty ratio reaches this value. Rationale: compacting a log that is only 10% duplicates rewrites 90% of it for little reclaimed space, so the default waits until at least half is dirty. - **Lower (e.g., 0.1):** compaction runs sooner and more often → tombstones/duplicates removed faster, smaller log, but **more CPU + disk I/O** because logs are rewritten frequently. - **Higher (e.g., 0.9):** less overhead, but duplicates and tombstones linger much longer, inflating the log and delaying state convergence. ## Time-based forcing Ratio alone can leave a low-write log uncleaned indefinitely. Two configs add a time dimension: - `max.compaction.lag.ms` — the maximum time a message can remain *uncompacted* in the dirty head; once exceeded, the log is forced eligible **regardless of ratio**. Critical for GDPR-style 'delete within N days' guarantees. - `min.compaction.lag.ms` — the minimum time a message must stay in the head before it can be compacted, guaranteeing consumers a window to read recent records. ## Practical takeaways - The greedy ratio policy means hot topics get cleaned first automatically. - Tune `min.cleanable.dirty.ratio` down only if you need fresher compaction; budget the extra I/O. - Use `max.compaction.lag.ms` when compliance requires bounded time-to-removal, since ratio-based selection gives no time guarantee.

  • A compacted topic has low write volume and its data is never getting compacted promptly. What config fixes a 'must remove within N days' requirement?
    max.compaction.lag.ms — it forces a log eligible once a record has sat uncompacted for that long, independent of the dirty ratio. The ratio policy alone gives no time bound.
  • What is the trade-off of setting min.cleanable.dirty.ratio very low?
    Compaction becomes aggressive and timely (less duplicate accumulation, smaller logs) but consumes much more CPU and disk I/O because logs are rewritten far more frequently.

saying these in an interview costs you the question

  • Saying the cleaner just goes round-robin or oldest-first — it is greedy by dirty ratio.
  • Computing dirty ratio over the whole log including the active segment — the active segment and lag-held data are excluded.
  • Claiming min.cleanable.dirty.ratio guarantees a time bound — it is purely a size ratio; use max.compaction.lag.ms for time.
  • Thinking lowering the ratio is free — it costs CPU and I/O from frequent rewrites.

context