skip to content

What is a Kafka partition's commit log, and why is it split into segments on disk?

level: juniorimportance: must knowfreq 70%

answer

  1. append-only, immutable tail
  2. .log + .index + .timeindex
  3. base offset = filename (20 digits)
  4. active = writable, closed = read-only
  5. retention deletes whole closed segments

basics

~20 s

Each partition is an append-only log: records are only added to the end, never changed in place. Kafka splits that log into fixed-size files called segments so old data can be deleted or compacted one whole file at a time instead of editing one giant file.

solid answer

~40 s

A Kafka partition is physically an append-only commit log: producers only append records to the tail, and records are immutable once written. Instead of one ever-growing file, Kafka stores the partition as a sequence of segments on disk. Each segment is a set of files sharing a base name (the base offset), most importantly the .log file holding the records, plus .index (offset->byte position) and .timeindex (timestamp->offset) companions. Only the newest segment (the active segment) is being written; older segments are closed and read-only. Splitting into segments makes retention cheap: to delete or compact data, Kafka operates on whole closed segment files rather than rewriting a single massive file, and it can map/serve reads efficiently. Roll triggers (segment.bytes, segment.ms) decide when the active segment closes and a new one opens.

go deeper

for a junior

Know: partition = append-only log, split into segment files, only the newest is being written.

for a middle

Add the file trio (.log/.index/.timeindex), base-offset naming, active vs closed distinction.

for a senior

Explain why segmentation enables cheap retention/compaction and the roll triggers.

for a principal

Reason about segment sizing trade-offs (file handles vs retention granularity) and log.dirs/JBOD layout decisions.

## The commit log A Kafka **topic** is divided into **partitions**, and each partition is the real unit of storage and ordering. A partition is implemented as a **commit log**: an ordered, append-only sequence of records. "Append-only" means writes only ever happen at the end (the tail); existing records are never modified or moved. Each record gets a monotonically increasing integer called the **offset** (0, 1, 2, ...) that identifies its position in that partition. ## Why segments instead of one file If a partition were a single file, it would grow without bound, and deleting old data (retention) would require rewriting the whole file. Instead Kafka chops the log into **segments**. A segment is a group of files on disk that share a **base name** equal to the **base offset** — the offset of the first record in that segment, zero-padded to 20 digits. For base offset 0 you get files like: - `00000000000000000000.log` — the actual records (this IS the log data) - `00000000000000000000.index` — a sparse map from relative offset to byte position in the .log, so a consumer can seek without scanning - `00000000000000000000.timeindex` — a sparse map from timestamp to offset, used for time-based lookups and time retention When the next segment starts at, say, offset 6000, its files are named `00000000000000006000.log`, etc. ## Active vs closed segments At any moment exactly one segment per partition is the **active segment** — the only one currently open for writing. All earlier segments are **closed** (read-only). Retention (delete or compaction) only ever touches closed segments; the active segment is never deleted, which is why you can have data older than your retention window still on disk if no new writes have rolled the segment. ## When does a segment roll? The active segment is closed and a new one opened (a "roll") when any of these fire: - `segment.bytes` (default 1 GiB) — the segment reaches that size. - `segment.ms` (default 7 days) — that much time has passed since the segment was created, even if it isn't full. (`log.roll.ms`/`log.roll.hours` are the broker-level equivalents.) - A record arrives whose timestamp would overflow the index, or the index file fills. ## log.dir layout Brokers store data under directories listed in `log.dirs` (or single `log.dir`). Inside, there is one directory per partition named `<topic>-<partition>` (e.g. `orders-3`), and inside that live the segment files plus a `leader-epoch-checkpoint` and (since KRaft/newer) partition metadata. Multiple `log.dirs` let a broker spread partitions across disks (JBOD). ## Edge cases - The active segment can be older than `retention.ms` and still present, because retention skips the active segment. - A small/idle partition may keep one segment for a long time until `segment.ms` rolls it. - Reducing `segment.bytes` creates more, smaller files — finer retention granularity but more open file handles.

  • Why can't retention delete the active segment?
    The active segment is the one currently being written, so Kafka never deletes or compacts it. Data must first roll into a closed segment (via segment.bytes/segment.ms) before retention can act, which is why data older than retention.ms can linger on a low-traffic partition.
  • What do the three files of a segment hold?
    .log holds the actual record batches; .index is a sparse offset->byte-position map for seeking; .timeindex is a sparse timestamp->offset map for time lookups and time-based retention.

saying these in an interview costs you the question

  • Saying records can be updated or deleted in place — the log is append-only and immutable.
  • Claiming the whole partition is one file — it is a sequence of segment files.
  • Confusing offset (logical record position) with byte position in the file.
  • Saying every segment is writable — only the active segment is.

context