What is a Kafka partition's commit log, and why is it split into segments on disk?
answer
- append-only, immutable tail
- .log + .index + .timeindex
- base offset = filename (20 digits)
- active = writable, closed = read-only
- retention deletes whole closed segments
basics
~20 sEach partition is an append-only log: records are only added to the end, never changed in place. Kafka splits that log into fixed-size files called segments so old data can be deleted or compacted one whole file at a time instead of editing one giant file.
solid answer
~40 sA Kafka partition is physically an append-only commit log: producers only append records to the tail, and records are immutable once written. Instead of one ever-growing file, Kafka stores the partition as a sequence of segments on disk. Each segment is a set of files sharing a base name (the base offset), most importantly the .log file holding the records, plus .index (offset->byte position) and .timeindex (timestamp->offset) companions. Only the newest segment (the active segment) is being written; older segments are closed and read-only. Splitting into segments makes retention cheap: to delete or compact data, Kafka operates on whole closed segment files rather than rewriting a single massive file, and it can map/serve reads efficiently. Roll triggers (segment.bytes, segment.ms) decide when the active segment closes and a new one opens.
go deeper
Know: partition = append-only log, split into segment files, only the newest is being written.
Add the file trio (.log/.index/.timeindex), base-offset naming, active vs closed distinction.
Explain why segmentation enables cheap retention/compaction and the roll triggers.
Reason about segment sizing trade-offs (file handles vs retention granularity) and log.dirs/JBOD layout decisions.
## The commit log A Kafka **topic** is divided into **partitions**, and each partition is the real unit of storage and ordering. A partition is implemented as a **commit log**: an ordered, append-only sequence of records. "Append-only" means writes only ever happen at the end (the tail); existing records are never modified or moved. Each record gets a monotonically increasing integer called the **offset** (0, 1, 2, ...) that identifies its position in that partition. ## Why segments instead of one file If a partition were a single file, it would grow without bound, and deleting old data (retention) would require rewriting the whole file. Instead Kafka chops the log into **segments**. A segment is a group of files on disk that share a **base name** equal to the **base offset** — the offset of the first record in that segment, zero-padded to 20 digits. For base offset 0 you get files like: - `00000000000000000000.log` — the actual records (this IS the log data) - `00000000000000000000.index` — a sparse map from relative offset to byte position in the .log, so a consumer can seek without scanning - `00000000000000000000.timeindex` — a sparse map from timestamp to offset, used for time-based lookups and time retention When the next segment starts at, say, offset 6000, its files are named `00000000000000006000.log`, etc. ## Active vs closed segments At any moment exactly one segment per partition is the **active segment** — the only one currently open for writing. All earlier segments are **closed** (read-only). Retention (delete or compaction) only ever touches closed segments; the active segment is never deleted, which is why you can have data older than your retention window still on disk if no new writes have rolled the segment. ## When does a segment roll? The active segment is closed and a new one opened (a "roll") when any of these fire: - `segment.bytes` (default 1 GiB) — the segment reaches that size. - `segment.ms` (default 7 days) — that much time has passed since the segment was created, even if it isn't full. (`log.roll.ms`/`log.roll.hours` are the broker-level equivalents.) - A record arrives whose timestamp would overflow the index, or the index file fills. ## log.dir layout Brokers store data under directories listed in `log.dirs` (or single `log.dir`). Inside, there is one directory per partition named `<topic>-<partition>` (e.g. `orders-3`), and inside that live the segment files plus a `leader-epoch-checkpoint` and (since KRaft/newer) partition metadata. Multiple `log.dirs` let a broker spread partitions across disks (JBOD). ## Edge cases - The active segment can be older than `retention.ms` and still present, because retention skips the active segment. - A small/idle partition may keep one segment for a long time until `segment.ms` rolls it. - Reducing `segment.bytes` creates more, smaller files — finer retention granularity but more open file handles.
- Why can't retention delete the active segment?The active segment is the one currently being written, so Kafka never deletes or compacts it. Data must first roll into a closed segment (via segment.bytes/segment.ms) before retention can act, which is why data older than retention.ms can linger on a low-traffic partition.
- What do the three files of a segment hold?.log holds the actual record batches; .index is a sparse offset->byte-position map for seeking; .timeindex is a sparse timestamp->offset map for time lookups and time-based retention.
saying these in an interview costs you the question
- Saying records can be updated or deleted in place — the log is append-only and immutable.
- Claiming the whole partition is one file — it is a sequence of segment files.
- Confusing offset (logical record position) with byte position in the file.
- Saying every segment is writable — only the active segment is.