What problem does min.compaction.lag.ms solve, and how does it interact with consumers that need to see every update to a key?
answer
- minimum age before compactable
- protects intermediate update visibility (CDC/audit)
- floor; max.compaction.lag.ms is ceiling
- NOT the same as delete.retention.ms
- default 0 = eligible ASAP
basics
~20 smin.compaction.lag.ms sets a minimum age a record must reach before the cleaner may compact it away. It guarantees that every version of a key stays in the log for at least that long, so consumers reading within that window can observe intermediate updates, not just the latest value.
solid answer
~40 sBy default (min.compaction.lag.ms=0) a record becomes compactable as soon as a newer value for its key exists and the segment is in the dirty section — so a fast cleaner can erase intermediate states almost immediately. Some consumers need to see every transition of a key (e.g. an audit/CDC stream or a processor that reacts to each change, not just the final state). Setting min.compaction.lag.ms to, say, 10 minutes guarantees that no record is compacted until it's at least 10 minutes old, giving consumers a guaranteed window to read all versions. It pairs with max.compaction.lag.ms (the upper deadline) to bracket compaction timing. Note it gates eligibility by record age; it does not by itself keep the log uncompacted forever, and it is distinct from delete.retention.ms (which is specifically about tombstone visibility).
go deeper
Know it's a delay before records can be compacted.
Explain the intermediate-update visibility window it provides and the default of 0.
Size it against consumer lag and distinguish it cleanly from delete.retention.ms.
Architect CDC/audit pipelines over compacted topics with explicit visibility-window SLAs.
## The intermediate-state problem A compacted topic eventually keeps only the latest value per key. But some downstream needs **every** version: a change-data-capture pipeline that emits each update, an auditor that records all transitions, or a stream processor whose logic depends on intermediate values. If the cleaner compacts too eagerly, a consumer that's even slightly behind would skip straight from an early value to the latest, losing the in-between records. ## What min.compaction.lag.ms does **`min.compaction.lag.ms`** (default **0**) defines the **minimum age** a message must reach before it is eligible to be compacted away. A record younger than this lag is **not** a candidate for removal even if a newer value for its key already exists. So if you set it to 600000 (10 minutes), every version of every key is guaranteed to remain in the log for at least 10 minutes after it was written, giving consumers a bounded window to observe all updates. ## Interaction with the dirty ratio and the deadline - The lag is an **eligibility floor**: a record must be *both* older than min.compaction.lag.ms *and* part of a log whose dirty ratio crossed min.cleanable.dirty.ratio (or hit max.compaction.lag.ms) to actually be compacted. - **`max.compaction.lag.ms`** is the complementary **ceiling**, forcing compaction by a deadline. Together they bracket how stale or fresh the compacted view can be. ## How it differs from delete.retention.ms These are frequently confused: - **min.compaction.lag.ms** delays compaction of *any* record by age, protecting intermediate update visibility. - **delete.retention.ms** specifically keeps *tombstones* visible after compaction so lagging consumers see deletes. One is about not-yet-compacted versions; the other is about how long delete markers persist post-compaction. ## Edge cases - Setting min.compaction.lag.ms high effectively widens the uncompacted tail and increases disk usage, since superseded values stick around longer. - It interacts with segment rolling: records in the active segment are uncompactable regardless, so very low-traffic topics may exceed the lag naturally before a segment even rolls. - It does not affect time-based delete retention; on a compact,delete topic, retention.ms still ages out old segments independently.
- A CDC consumer occasionally misses intermediate updates to a key. Which compaction config would you raise and why?Raise min.compaction.lag.ms so each version of a key stays in the log long enough for the consumer to read it before the cleaner removes superseded records. Size it above the consumer's worst-case lag.
- How is min.compaction.lag.ms different from delete.retention.ms?min.compaction.lag.ms delays compaction of any record by its age (protecting intermediate, non-null versions). delete.retention.ms keeps tombstones visible after compaction so deletes aren't missed. Different records, different purposes.
saying these in an interview costs you the question
- Confusing it with delete.retention.ms.
- Saying it prevents compaction entirely (it only delays eligibility by age).
- Claiming default is nonzero — the default is 0.
- Thinking it guarantees consumers see every update regardless of lag — only within the configured window.