skip to content

Describe the __cluster_metadata topic in KRaft: how it is structured, who reads and writes it, and why it is special.

level: middleimportance: must knowfreq 65%

answer

  1. single partition, partition 0
  2. Raft-replicated across controller voters
  3. active controller = only writer, brokers = observers
  4. offset + epoch ordering
  5. snapshots bound replay; kafka-metadata-shell to inspect

basics

~20 s

__cluster_metadata is the internal Kafka topic that holds all cluster metadata as an event log. It has a single partition replicated across the controller quorum via Raft. The active controller writes to it; controllers and brokers read it to stay in sync.

solid answer

~50 s

In KRaft, cluster metadata lives in a special internal topic, __cluster_metadata, with exactly one partition (partition 0). That single partition is Raft-replicated across the controller voters: the active controller is the partition leader and is the only writer; follower controllers replicate it for high availability. Brokers act as read-only observers — they fetch the log and apply records to build an in-memory view of topics, partitions, leaders, configs, and ACLs. Each record has an offset and an epoch, giving a totally ordered event log. To keep replay bounded, controllers periodically take metadata snapshots so a restarting node loads the latest snapshot plus the tail of the log instead of replaying from offset 0. Unlike normal topics it is not created by users, not directly consumable by clients, and is the literal source of truth for cluster state — you inspect it with kafka-metadata-shell or kafka-dump-log, not the regular consumer API.

go deeper

for a junior

Know __cluster_metadata is the internal log that stores all cluster metadata in KRaft.

for a middle

Explain single partition, Raft replication across voters, controller-writes/brokers-observe, and that it is the source of truth.

for a senior

Discuss offset/epoch ordering, commit-on-majority, snapshots for bounded replay, and the inspection tooling.

for a principal

Reason about failure modes (majority loss = metadata loss), snapshot tuning, and the design rationale of metadata-as-event-log vs ZooKeeper's tree.

## What it is `__cluster_metadata` is an **internal topic** that KRaft uses as the single, authoritative event log of all cluster metadata. Everything ZooKeeper used to hold — topic and partition definitions, partition leadership and ISR, broker registrations, dynamic configs, ACLs, client quotas, and producer-ID state — is now expressed as **records** appended to this log. ## Structure: one partition, Raft-replicated The topic has **exactly one partition** (`__cluster_metadata-0`). A single partition is deliberate: metadata changes must be **totally ordered** so every node applies them in the same sequence, and one Raft log gives exactly that. This partition is **not** replicated by the normal ISR mechanism. Instead it is replicated by the **KRaft Raft protocol** across the **controller quorum** (the set of nodes listed in `controller.quorum.voters`). - Each record carries an **offset** (its position in the log) and an **epoch** (which leadership term wrote it). Together these give a strict, gap-free ordering and let nodes detect stale leaders. - A record is **committed** once a **majority** of voters have persisted it (e.g. 2 of 3). Only committed records are applied. ## Who writes and who reads - **Active controller (the Raft leader)**: the *only* writer. All metadata mutations go through it and are appended here. - **Follower controllers (voters)**: replicate the log by fetching from the leader; ready to take over if the leader fails. - **Brokers**: **observers** (non-voting). They fetch the log to maintain an in-memory metadata cache (the "metadata image") and never write to it. This is how a broker learns it is now leader for a partition, that a topic was created, etc. ## Snapshots: bounding replay An append-only log grows forever, so replaying it from offset 0 on every restart would be unbounded. KRaft periodically writes a **metadata snapshot**: a compacted point-in-time image of all metadata at a given offset. On startup a node loads the **latest snapshot** and then replays only the **log tail** after that offset. This keeps startup fast and storage bounded (older log segments before the snapshot can be discarded). ## Why it is special vs a normal topic - Auto-created by the cluster, never by users; you cannot delete it. - Replicated by Raft, not by ISR; its leadership election is the controller election. - Not consumable through the ordinary KafkaConsumer client. You inspect it with tooling: **`kafka-metadata-shell.sh`** (browse the metadata as a filesystem-like tree) or **`kafka-dump-log.sh`** (decode the raw records on disk). - It is the **source of truth**: if it is corrupted or lost on a majority of voters, the cluster loses its metadata. ## Edge cases - It lives under the controllers' `metadata.log.dir` (or the configured log dir) as `__cluster_metadata-0`. - Because it is a single partition, throughput of metadata writes is bounded by one leader — fine, because metadata-change rate is tiny compared to data traffic. - Snapshot generation is controlled by configs like `metadata.log.max.record.bytes.between.snapshots` / `metadata.snapshot.max.*` thresholds.

  • Why does the metadata topic have only one partition?
    Metadata changes must be applied in the exact same order on every node, so a single totally-ordered Raft log is required. Multiple partitions would lose the global ordering guarantee. Metadata write throughput is tiny, so one partition is not a bottleneck.
  • How does a restarting controller avoid replaying the entire log from the beginning?
    KRaft periodically writes metadata snapshots — compacted images at a given offset. A restarting node loads the latest snapshot, then replays only the log records after that offset, keeping startup fast and storage bounded.

saying these in an interview costs you the question

  • Saying brokers write to __cluster_metadata (only the active controller writes; brokers are observers)
  • Claiming it has multiple partitions for scalability (it is single-partition by design for total ordering)
  • Thinking it is replicated by normal ISR rather than the Raft protocol
  • Believing you read it with a normal KafkaConsumer (use kafka-metadata-shell / kafka-dump-log)

context