skip to content

Explain the head vs tail of a compacted log and why the active segment is never compacted.

level: seniorimportance: should knowfreq 35%

answer

  1. checkpoint offset splits clean tail / dirty head
  2. tail: <=1 record per key; head: duplicates
  3. active segment = currently appended, never compacted
  4. rolls at segment.bytes / segment.ms
  5. recent data verbatim -> duplicates visible at the tip

basics

~20 s

The tail is the older, already-compacted part of the log (at most one record per key). The head is newer records appended since the last clean, which may still contain duplicates. The active segment — the one being written to — is always part of the head and is never compacted.

solid answer

~50 s

Kafka conceptually splits a compacted partition into a clean tail and a dirty head, divided by the cleaner checkpoint offset. The tail has already been deduplicated, so each key appears at most once; the head is everything written since the last clean and may hold multiple records per key. The cleaner only ever processes up to the start of the active segment — the segment currently receiving appends. The active segment is excluded because it is being written concurrently, its size/contents change constantly, and Kafka guarantees the most recent records are always readable verbatim. This also means a key written only in the active segment is never compacted yet, so the very latest data is preserved exactly and read latency stays low. The head therefore always contains at least the active segment plus any rolled-but-not-yet-cleaned segments; segment.bytes and segment.ms control how quickly the active segment rolls and thus how soon new data becomes cleanable.

go deeper

for a junior

Know the newest segment being written isn't compacted, so fresh data may have duplicates.

for a middle

Define head (dirty, new) vs tail (clean, old) and that the active segment is always excluded.

for a senior

Explain the checkpoint offset, why active-segment exclusion is required, and how roll cadence affects compaction latency.

for a principal

Reason about segment sizing trade-offs for compaction freshness vs file overhead across high-throughput compacted topics.

## The clean/dirty split A compacted partition's log is logically divided by the **cleaner checkpoint offset**: - **Tail (clean):** offsets *below* the checkpoint. The cleaner has already run here, so each key appears **at most once** — this is the deduplicated 'state'. - **Head (dirty):** offsets at/above the checkpoint up to the active segment. These are records appended since the last clean; they may contain **multiple records per key** (duplicates not yet collapsed). After each compaction pass the checkpoint advances, converting freshly-cleaned head into new tail. ## The active segment A partition log is a sequence of **segment** files. Exactly one is the **active segment**: the newest file, the one currently being **appended** to by producers. When it reaches `segment.bytes` (default 1 GiB) or `segment.ms` elapses, it is **rolled** (closed) and a brand-new active segment is created. ## Why the active segment is never compacted 1. **Concurrent writes:** producers are appending to it right now; rewriting a file under active append is unsafe and would require heavy locking. 2. **Stable, verbatim latest data:** Kafka guarantees the newest records are immediately and exactly readable. If the active segment could be compacted, recent records might be reorganized or removed mid-stream. 3. **Offset/index stability:** the active segment owns the log-end offset and its index is being extended; compaction rewrites segments and rebuilds indexes, which only makes sense for closed segments. So the cleanable region is strictly **everything below the active segment**. A key that has only ever been written into the active segment will not be compacted until that segment rolls and a later cleaning pass covers it. ## Consequences - **Duplicates near the head are normal.** Even on a compacted topic, a consumer reading the tip will see un-collapsed duplicates and tombstones — consumers must take the last value per key. - **Roll cadence affects compaction latency.** Smaller `segment.bytes` / `segment.ms` roll the active segment sooner, making recent data cleanable faster (at the cost of more, smaller files). Large segments delay compaction of recent keys. - **The tail is the 'table'.** Tools/state stores treat the compacted tail as the materialized current state of each key. ## Relation to other configs - `min.cleanable.dirty.ratio` decides *when* the head is dirty enough to clean. - `min.compaction.lag.ms` can hold records in the head even after the segment rolls. - The active segment exclusion is structural and not directly configurable — you influence it indirectly via segment roll settings.

  • A team complains duplicates for a key keep appearing on their compacted topic even though compaction is on. What's the likely benign explanation?
    Those duplicates are in the head / active segment, which compaction never touches until the segment rolls and a pass runs. Compaction only guarantees dedup in the tail; consumers must keep the last value per key.
  • How can you make recently written keys become compactable sooner?
    Lower segment.bytes or segment.ms so the active segment rolls more frequently, moving recent records out of the active segment into the cleanable region sooner — trading off more, smaller segment files.

saying these in an interview costs you the question

  • Saying the whole log is always fully deduplicated — the head and active segment can hold duplicates.
  • Claiming the active segment is compacted in place — it is always excluded from cleaning.
  • Confusing the tail (old, clean) with the head (new, dirty) — getting the direction backwards.
  • Believing the active-segment exclusion is a tunable config rather than structural.

context