Walk through the on-disk layout under a broker's log.dirs: what directories and files exist for a partition, and how do you inspect a segment?
answer
- log.dirs -> <topic>-<partition> dirs
- JBOD: partition lives on one disk, no striping
- leader-epoch-checkpoint + partition.metadata
- checkpoints: recovery / replication(HWM) / cleaner
- kafka-dump-log.sh --print-data-log
basics
~10 sUnder each path in log.dirs there's one directory per partition named <topic>-<partition> (e.g. orders-3). Inside are the segment files (.log/.index/.timeindex), a leader-epoch-checkpoint, and partition metadata. You inspect a .log with kafka-dump-log.sh.
solid answer
~40 sA broker stores data under the directories listed in `log.dirs` (default a single `/tmp/kafka-logs`). Each replica the broker hosts gets its own directory named `<topic>-<partition>`, e.g. `orders-3`. Inside that directory are the segment file sets — `<baseOffset>.log`, `.index`, `.timeindex` — plus `leader-epoch-checkpoint` (tracks leader epochs for truncation correctness), and `partition.metadata` (topic ID, KRaft). At the root of each log.dir there are bookkeeping files: `recovery-point-offset-checkpoint`, `replication-offset-checkpoint` (the HWM checkpoint), `cleaner-offset-checkpoint` (compaction progress), and `meta.properties` (cluster/broker id). Multiple `log.dirs` give JBOD: Kafka places each partition entirely in one chosen directory (no striping) and balances partition count across disks. To inspect a segment you use `kafka-dump-log.sh --files 00000000000000000000.log --print-data-log`, which decodes batches, offsets, timestamps, and record headers; add `.index`/`.timeindex` files to dump the index entries.
go deeper
Know data lives under log.dirs in <topic>-<partition> folders containing segment files.
Name the segment trio plus that there are checkpoint files and a way to dump segments.
Explain JBOD placement, leader-epoch/partition.metadata, the checkpoint files, and kafka-dump-log.sh usage.
Reason about disk-failure isolation, FD budgeting, recovery time vs checkpoints, and intra-broker reassignment artifacts.
## Top level: log.dirs A broker writes log data under the directories named by `log.dirs` (a comma-separated list; the singular `log.dir` is the fallback, default `/tmp/kafka-logs`). Each entry is typically a separate physical disk — this is Kafka's **JBOD** (Just a Bunch Of Disks) model. Kafka does NOT stripe a partition across disks; it assigns each partition's directory wholly to one chosen log.dir, balancing the count of partitions across disks. ## Per-log.dir bookkeeping files At the root of each log.dir you'll find: - `meta.properties` — cluster id, broker/node id, directory id. - `recovery-point-offset-checkpoint` — per-partition offset up to which data is known flushed; used to bound recovery work on restart. - `replication-offset-checkpoint` — the high-water-mark checkpoint per partition. - `cleaner-offset-checkpoint` — how far log compaction has cleaned each partition. - `log-start-offset-checkpoint` — the log start offset per partition (after retention/deletes). ## Per-partition directory Each replica hosted by the broker has a directory named `<topic>-<partition>`: ``` orders-3/ 00000000000000000000.log 00000000000000000000.index 00000000000000000000.timeindex 00000000000000006000.log 00000000000000006000.index 00000000000000006000.timeindex leader-epoch-checkpoint partition.metadata ``` - The **segment file trio** repeats per segment, named by base offset. - `leader-epoch-checkpoint` — maps leader epoch -> start offset; lets followers truncate correctly after leadership changes (replaces the old high-water-mark-only truncation, fixing data-loss edge cases). - `partition.metadata` — stores the topic id (UUID); important under KRaft where topics are identified by id, not just name. - Compacted topics may also show `.deleted`, `.cleaned`, `.swap` temp files during compaction, and a `.snapshot` file for the producer-state (idempotence/transactions). ## A topic being deleted A partition pending deletion is renamed to `<topic>-<partition>.<uuid>-delete` and removed asynchronously, so you may briefly see `-delete` suffixed directories. ## Inspecting segments Kafka ships `kafka-dump-log.sh` (class `kafka.tools.DumpLogSegments`): ``` kafka-dump-log.sh \ --files /var/lib/kafka/orders-3/00000000000000000000.log \ --print-data-log ``` This prints each record batch: baseOffset, lastOffset, producerId, baseSequence, compression, and per-record offset/timestamp/key/value/headers (with `--print-data-log`). Pointing `--files` at a `.index` or `.timeindex` dumps the sparse index entries (relative offset -> position, or timestamp -> offset). `--index-sanity-check`/`--verify-index-only` validate index integrity. There's also `--deep-iteration` to descend into compressed batches. ## Why this layout matters operationally - A failed disk in JBOD takes down only the partitions on that log.dir; with `log.dir.failure.timeout.ms`/online-dir-failure handling the broker can keep serving the others. - File-descriptor pressure scales with (segments per partition) x (partitions per broker) — sizing segments interacts with `ulimit -n`. - Recovery time after an unclean shutdown depends on how far behind the recovery-point checkpoints are, since Kafka must re-validate segments past the recovery point. ## Edge cases - After moving a partition between disks (intra-broker reassignment), you may see a temporary `*-future` directory holding the in-progress copy. - `/tmp/kafka-logs` as the default is a common production footgun — /tmp can be wiped on reboot.
- Does a single partition span multiple log.dirs for striping?No. Kafka places each partition's directory entirely within one chosen log.dir (JBOD). Multiple log.dirs balance partition COUNT across disks; a single partition never stripes across them.
- What is leader-epoch-checkpoint for?It maps each leader epoch to the offset where that epoch began. Followers use it to truncate their log to the correct point after a leadership change, fixing data-loss/divergence bugs that pure high-water-mark truncation had.
saying these in an interview costs you the question
- Saying a partition is striped across all log.dirs — it lives wholly in one.
- Forgetting log.epoch/leader-epoch-checkpoint and treating HWM as the only truncation anchor.
- Reading .log files with cat/text tools instead of kafka-dump-log.sh — they're binary record batches.
- Leaving log.dirs at the default /tmp/kafka-logs in production.