skip to content

Walk through the on-disk layout under a broker's log.dirs: what directories and files exist for a partition, and how do you inspect a segment?

level: seniorimportance: should knowfreq 35%

answer

  1. log.dirs -> <topic>-<partition> dirs
  2. JBOD: partition lives on one disk, no striping
  3. leader-epoch-checkpoint + partition.metadata
  4. checkpoints: recovery / replication(HWM) / cleaner
  5. kafka-dump-log.sh --print-data-log

basics

~10 s

Under each path in log.dirs there's one directory per partition named <topic>-<partition> (e.g. orders-3). Inside are the segment files (.log/.index/.timeindex), a leader-epoch-checkpoint, and partition metadata. You inspect a .log with kafka-dump-log.sh.

solid answer

~40 s

A broker stores data under the directories listed in `log.dirs` (default a single `/tmp/kafka-logs`). Each replica the broker hosts gets its own directory named `<topic>-<partition>`, e.g. `orders-3`. Inside that directory are the segment file sets — `<baseOffset>.log`, `.index`, `.timeindex` — plus `leader-epoch-checkpoint` (tracks leader epochs for truncation correctness), and `partition.metadata` (topic ID, KRaft). At the root of each log.dir there are bookkeeping files: `recovery-point-offset-checkpoint`, `replication-offset-checkpoint` (the HWM checkpoint), `cleaner-offset-checkpoint` (compaction progress), and `meta.properties` (cluster/broker id). Multiple `log.dirs` give JBOD: Kafka places each partition entirely in one chosen directory (no striping) and balances partition count across disks. To inspect a segment you use `kafka-dump-log.sh --files 00000000000000000000.log --print-data-log`, which decodes batches, offsets, timestamps, and record headers; add `.index`/`.timeindex` files to dump the index entries.

go deeper

for a junior

Know data lives under log.dirs in <topic>-<partition> folders containing segment files.

for a middle

Name the segment trio plus that there are checkpoint files and a way to dump segments.

for a senior

Explain JBOD placement, leader-epoch/partition.metadata, the checkpoint files, and kafka-dump-log.sh usage.

for a principal

Reason about disk-failure isolation, FD budgeting, recovery time vs checkpoints, and intra-broker reassignment artifacts.

## Top level: log.dirs A broker writes log data under the directories named by `log.dirs` (a comma-separated list; the singular `log.dir` is the fallback, default `/tmp/kafka-logs`). Each entry is typically a separate physical disk — this is Kafka's **JBOD** (Just a Bunch Of Disks) model. Kafka does NOT stripe a partition across disks; it assigns each partition's directory wholly to one chosen log.dir, balancing the count of partitions across disks. ## Per-log.dir bookkeeping files At the root of each log.dir you'll find: - `meta.properties` — cluster id, broker/node id, directory id. - `recovery-point-offset-checkpoint` — per-partition offset up to which data is known flushed; used to bound recovery work on restart. - `replication-offset-checkpoint` — the high-water-mark checkpoint per partition. - `cleaner-offset-checkpoint` — how far log compaction has cleaned each partition. - `log-start-offset-checkpoint` — the log start offset per partition (after retention/deletes). ## Per-partition directory Each replica hosted by the broker has a directory named `<topic>-<partition>`: ``` orders-3/ 00000000000000000000.log 00000000000000000000.index 00000000000000000000.timeindex 00000000000000006000.log 00000000000000006000.index 00000000000000006000.timeindex leader-epoch-checkpoint partition.metadata ``` - The **segment file trio** repeats per segment, named by base offset. - `leader-epoch-checkpoint` — maps leader epoch -> start offset; lets followers truncate correctly after leadership changes (replaces the old high-water-mark-only truncation, fixing data-loss edge cases). - `partition.metadata` — stores the topic id (UUID); important under KRaft where topics are identified by id, not just name. - Compacted topics may also show `.deleted`, `.cleaned`, `.swap` temp files during compaction, and a `.snapshot` file for the producer-state (idempotence/transactions). ## A topic being deleted A partition pending deletion is renamed to `<topic>-<partition>.<uuid>-delete` and removed asynchronously, so you may briefly see `-delete` suffixed directories. ## Inspecting segments Kafka ships `kafka-dump-log.sh` (class `kafka.tools.DumpLogSegments`): ``` kafka-dump-log.sh \ --files /var/lib/kafka/orders-3/00000000000000000000.log \ --print-data-log ``` This prints each record batch: baseOffset, lastOffset, producerId, baseSequence, compression, and per-record offset/timestamp/key/value/headers (with `--print-data-log`). Pointing `--files` at a `.index` or `.timeindex` dumps the sparse index entries (relative offset -> position, or timestamp -> offset). `--index-sanity-check`/`--verify-index-only` validate index integrity. There's also `--deep-iteration` to descend into compressed batches. ## Why this layout matters operationally - A failed disk in JBOD takes down only the partitions on that log.dir; with `log.dir.failure.timeout.ms`/online-dir-failure handling the broker can keep serving the others. - File-descriptor pressure scales with (segments per partition) x (partitions per broker) — sizing segments interacts with `ulimit -n`. - Recovery time after an unclean shutdown depends on how far behind the recovery-point checkpoints are, since Kafka must re-validate segments past the recovery point. ## Edge cases - After moving a partition between disks (intra-broker reassignment), you may see a temporary `*-future` directory holding the in-progress copy. - `/tmp/kafka-logs` as the default is a common production footgun — /tmp can be wiped on reboot.

  • Does a single partition span multiple log.dirs for striping?
    No. Kafka places each partition's directory entirely within one chosen log.dir (JBOD). Multiple log.dirs balance partition COUNT across disks; a single partition never stripes across them.
  • What is leader-epoch-checkpoint for?
    It maps each leader epoch to the offset where that epoch began. Followers use it to truncate their log to the correct point after a leadership change, fixing data-loss/divergence bugs that pure high-water-mark truncation had.

saying these in an interview costs you the question

  • Saying a partition is striped across all log.dirs — it lives wholly in one.
  • Forgetting log.epoch/leader-epoch-checkpoint and treating HWM as the only truncation anchor.
  • Reading .log files with cat/text tools instead of kafka-dump-log.sh — they're binary record batches.
  • Leaving log.dirs at the default /tmp/kafka-logs in production.

context