How does the log cleaner decide which partitions to compact, and what role does min.cleanable.dirty.ratio play?
answer
- dirtyRatio = dirty / (clean + dirty)
- min.cleanable.dirty.ratio default 0.5
- cleanest = pick highest ratio first
- min.compaction.lag.ms = floor, max = deadline
- active segment always excluded
basics
~20 sThe log cleaner splits each partition into a 'clean' head (already compacted) and a 'dirty' tail (new records). It computes the dirty ratio = dirty bytes / total bytes and only cleans a partition when that ratio exceeds min.cleanable.dirty.ratio (default 0.5). The dirtiest eligible log is cleaned first.
solid answer
~50 sCompaction runs on a pool of cleaner threads (log.cleaner.threads). Each partition log is conceptually divided into a clean section (compacted in a prior pass) and a dirty section (records appended since). The cleaner computes dirty ratio = dirtyBytes / (cleanBytes + dirtyBytes). A partition becomes eligible only when this ratio >= min.cleanable.dirty.ratio (default 0.5), preventing wasteful constant recompaction. Among eligible logs the cleaner picks the one with the highest dirty ratio to maximize reclaimed space per pass. min.compaction.lag.ms forces records to wait before becoming eligible; max.compaction.lag.ms forces a partition to be cleaned within a deadline even if the ratio threshold isn't met. Lowering min.cleanable.dirty.ratio (e.g. to 0.1) makes compaction more aggressive — fresher state, more CPU/IO; raising it batches more work. The active segment is always excluded, so there's an irreducible amount of uncompacted data.
go deeper
Know that compaction is a background job that doesn't run on every write.
Explain the dirty/clean split and the 0.5 default ratio as a trigger.
Tune ratio plus the lag configs and reason about CPU/IO vs freshness trade-offs.
Set cluster-wide cleaner sizing/policy, diagnose stuck cleaners, and bound time-to-compact for compliance.
## The cleaner and the dirty/clean split The **log cleaner** is a background subsystem with a configurable thread pool (**`log.cleaner.threads`**, default 1). For each compacted partition it maintains a **cleaner checkpoint** marking the boundary between: - the **clean** portion: already compacted in an earlier pass (one value per key there), and - the **dirty** portion: records appended since the last clean. ## The dirty ratio and min.cleanable.dirty.ratio The cleaner computes: ``` dirtyRatio = dirtyBytes / (cleanBytes + dirtyBytes) ``` A log is **eligible** for cleaning only when `dirtyRatio >= min.cleanable.dirty.ratio` (default **0.5**). This avoids re-scanning a mostly-clean log over and over for little gain. Among all eligible logs, the cleaner greedily selects the one with the **highest** dirty ratio — that's where it reclaims the most space per unit of work. - **Lower** the ratio (e.g. 0.1) → compaction triggers sooner → state is fresher and old values vanish faster, at the cost of more CPU and disk IO and more frequent rewrites. - **Raise** it → fewer, larger passes → less overhead but staler tails and more disk used by superseded values. ## Time-based controls - **`min.compaction.lag.ms`** (default 0): a record cannot be compacted until at least this long after it was appended. Guarantees a minimum time window in which every version of a key is visible — useful for consumers that need to see intermediate updates. - **`max.compaction.lag.ms`** (default Long.MAX, effectively off): an upper bound — a dirty record must be compacted within this time even if the dirty ratio never crosses the threshold. Important for compliance (bounded time-to-delete for tombstones) and for low-traffic topics that would otherwise never reach the ratio. ## The active segment caveat The **active segment** (currently being appended) is never compacted, and segments holding records younger than min.compaction.lag.ms are excluded. So a partition always carries some uncompacted bytes; you cannot drive the log to perfectly one-record-per-key in real time. ## How a pass works (high level) 1. Build an offset map of key → highest offset over the dirty section. 2. Recopy retained records (those at their highest offset, plus tombstones still within delete.retention.ms) into new segments, skipping superseded records. 3. Swap in the new segments and advance the cleaner checkpoint. ## Operational signals Metrics like `max-clean-time-secs`, `cleaner-recopy-percent`, and the per-log `uncleanable` state surface cleaner health. A stuck cleaner (e.g. on a corrupt record) can let the dirty log grow unbounded — watch `time-since-last-run-ms`.
- A low-traffic compacted topic never seems to compact old values. What config fixes that?Set max.compaction.lag.ms to a finite value so the cleaner is forced to compact within a deadline even though the dirty ratio never crosses min.cleanable.dirty.ratio. Optionally lower the ratio too.
- Why does Kafka pick the log with the highest dirty ratio rather than round-robin?To maximize space reclaimed per cleaner pass — the dirtiest log has the most superseded data to remove, giving the best return on the CPU/IO cost of a pass.
saying these in an interview costs you the question
- Saying compaction runs continuously on every append — it's threshold-driven and batched.
- Confusing min.compaction.lag.ms (eligibility floor) with delete.retention.ms (tombstone purge window).
- Claiming the dirty ratio includes the active segment in a way that triggers immediate cleaning.
- Thinking min.cleanable.dirty.ratio is a percentage 0-100 (it's a fraction 0.0-1.0).