skip to content

Where does KRaft store the cluster metadata on disk, and what is the on-disk layout of the __cluster_metadata log directory?

level: middleimportance: should knowfreq 45%

answer

  1. __cluster_metadata-0 = single partition
  2. Raft over the controller quorum
  3. quorum-state file = leader/epoch/voted-for
  4. .checkpoint = snapshots, enable truncation
  5. metadata.log.dir overrides log.dirs

basics

~10 s

KRaft stores metadata in an internal single-partition topic, __cluster_metadata-0, written as a normal Kafka log (segments, indexes) plus snapshot files. By default it lives under log.dirs; you can move it with metadata.log.dir.

solid answer

~40 s

KRaft keeps cluster metadata in the internal topic __cluster_metadata, which always has exactly one partition (__cluster_metadata-0) replicated via Raft across the controller quorum. On disk it's a standard Kafka log directory: numbered .log segments with their .index/.timeindex companions, a leader-epoch-checkpoint, plus Raft-specific files — quorum-state (current leader/epoch/voted-for) and periodic .checkpoint snapshot files that let the log be truncated. By default it sits inside the first entry of log.dirs; metadata.log.dir lets you place it on a dedicated (often faster) disk. Controllers hold the full authoritative log; brokers materialize a local replica of the same metadata to serve their own needs. Snapshots are critical: without them the metadata log would grow forever, and new/restarting nodes replay from the latest snapshot plus the tail rather than from offset zero.

go deeper

for a junior

Know metadata lives in an internal __cluster_metadata topic on disk under log.dirs.

for a middle

Describe the single-partition layout, segment/index files, quorum-state, snapshots, and metadata.log.dir.

for a senior

Explain Raft replication across the controller quorum, snapshot-driven truncation, and broker-as-observer materialization.

for a principal

Reason about isolating metadata I/O, single-Raft-log throughput limits, and recovery semantics from snapshot+tail.

## What metadata is Cluster metadata is the catalog of the cluster: topic and partition definitions, broker registrations, ACLs, configs, and controller/leader assignments. In ZooKeeper-based Kafka this lived in ZooKeeper znodes. In **KRaft** it lives in a Kafka topic that Kafka manages itself. ## The __cluster_metadata topic KRaft introduces an internal topic named **`__cluster_metadata`** with **exactly one partition**, so on disk you see the directory **`__cluster_metadata-0`**. This single partition is replicated not by ordinary ISR replication but by a **Raft** consensus protocol across the **controller quorum** (the nodes with `controller` in `process.roles`). One controller is the **active controller** (Raft leader); it appends metadata records, and followers replicate them. Every metadata change (create topic, register broker, change config) is a record appended to this log. ## On-disk layout `__cluster_metadata-0` is a normal Kafka **log directory**, so it contains: - **`*.log`** — segment files holding the actual metadata records (named by their base offset, e.g. `00000000000000000000.log`). - **`*.index`** and **`*.timeindex`** — offset and time indexes for each segment. - **`leader-epoch-checkpoint`** — maps leader epochs to start offsets. - **`partition.metadata`** — records the topic id. Plus KRaft/Raft-specific files in the directory: - **`quorum-state`** — the persisted Raft state: current leader id, current epoch, and who this node voted for. This survives restarts so a node doesn't forget its vote and break the consensus safety guarantees. - **`*.checkpoint`** snapshot files — point-in-time **snapshots** of the materialized metadata state at a given offset/epoch (named like `<offset>-<epoch>.checkpoint`). ## Where it lives By default the metadata partition is placed in the **first directory** of `log.dirs`. The config **`metadata.log.dir`** overrides this so you can put metadata on a dedicated disk — common in production to isolate metadata I/O from data I/O and to use faster storage for low-latency consensus. ## Why snapshots matter A Raft log is append-only; replaying every record since cluster creation would be slow and the log would grow without bound. KRaft periodically writes a **snapshot** capturing the full materialized state at offset N, then can delete log segments before N. A restarting controller or a broker catching up loads the **latest snapshot** and then replays only the tail of the log after it. This bounds recovery time and disk usage. ## Brokers vs controllers - **Controllers** hold the authoritative Raft replica and run the quorum. - **Brokers** are observers: they fetch the metadata log and build a local materialized copy (and their own `__cluster_metadata-0` replica on disk) so each broker has the current cluster state without querying a controller for every decision. ## Edge cases - Corruption or deletion of `quorum-state` can make a node lose its Raft identity; recovery may require re-bootstrapping from the quorum. - Because there's only one partition, metadata throughput is bounded by a single Raft log — fine because metadata events are comparatively low-rate.

  • Why does __cluster_metadata have only one partition?
    Metadata must be a single totally-ordered Raft log so all nodes agree on the exact sequence of changes; multiple partitions would break the single linearizable order that consensus and state materialization rely on.
  • What is the role of the .checkpoint snapshot files?
    They capture the full materialized metadata state at an offset so older log segments can be deleted and restarting nodes can load the snapshot plus tail instead of replaying from offset zero — bounding disk growth and recovery time.

saying these in an interview costs you the question

  • Saying __cluster_metadata has many partitions or is replicated by ISR like a normal topic.
  • Claiming metadata is stored back in ZooKeeper.
  • Thinking brokers don't keep any local metadata copy.
  • Forgetting snapshots and asserting the log replays from offset 0 on every restart.

context