skip to content

KRaft Storage Format and Cluster Bootstrap

Formatting storage with a cluster ID and bootstrapping the first quorum: kafka-storage.sh, meta.properties, and the metadata log directory. Asked because it is the first thing you touch when standing up a KRaft cluster and an easy place to get stuck.

part ofApache Kafkaoverview, primer and where to startread it →
on this pageshow

questions

5

Why must you run kafka-storage.sh format before starting a KRaft-mode Kafka broker or controller, and what does it create?

level: juniorimportance: must knowfreq 70%

answer

  1. No ZooKeeper → no auto cluster.id
  2. random-uuid then format -t -c
  3. meta.properties = cluster.id + node.id + version
  4. Missing file → node won't start
  5. Mismatch → InconsistentClusterIdException

basics

~20 s

In KRaft mode you must format each node's log directory first. kafka-storage.sh format writes a meta.properties file containing a shared cluster.id and the node.id, so the node knows which cluster it belongs to. Without it, the node refuses to start.

solid answer

~40 s

KRaft removes ZooKeeper, which previously assigned the cluster.id automatically on first connect. In KRaft, you generate a cluster.id yourself (kafka-storage.sh random-uuid) and then run kafka-storage.sh format -t <cluster-id> -c server.properties on every node. Format writes a meta.properties file into each configured log directory (log.dirs / metadata.log.dir) holding the cluster.id, the node.id, and a version. This is a one-time bootstrap step per node before first boot. If a node starts with an unformatted directory, it logs 'No `meta.properties` found' and exits with a non-zero code. The same cluster.id must be used across all nodes of one cluster; a mismatch causes the node to refuse to join (InconsistentClusterIdException), which prevents accidentally mixing nodes from different clusters in the same quorum.

go deeper

for a junior

Know that you must format storage first and that it writes meta.properties with the cluster.id; without it the node won't start.

for a middle

Explain the random-uuid + format -t -c two-step and the contents of meta.properties.

for a senior

Discuss InconsistentClusterIdException, --ignore-formatted, per-node vs shared identifiers, and the ZK→KRaft rationale.

for a principal

Frame meta.properties as the durable identity contract that prevents disk/node cross-pollination across clusters and how this replaces ZK's implicit identity assignment.

## Background: what changed with KRaft Apache Kafka historically depended on **ZooKeeper**, a separate coordination service, to store cluster metadata (which brokers exist, topic configs, partition leaders, etc.). When a broker first connected to a ZooKeeper ensemble, ZooKeeper handed it a **cluster.id** — a unique identifier for that Kafka cluster — and the broker persisted it locally. You never had to format anything by hand. **KRaft** (Kafka Raft) removes ZooKeeper. Metadata is now stored inside Kafka itself, in an internal topic called **`__cluster_metadata`**, managed by a built-in Raft consensus quorum of **controller** nodes. Because there is no ZooKeeper to mint the cluster.id and bootstrap the storage, you must do that explicitly before the first start. ## The two-step bootstrap 1. **Generate a cluster id** (do this once for the whole cluster): ``` KAFKA_CLUSTER_ID="$(bin/kafka-storage.sh random-uuid)" ``` This prints a base64-encoded 128-bit UUID, e.g. `MkU3OEVBNTcwNTJENDM2Qk`. 2. **Format each node's storage** using that id: ``` bin/kafka-storage.sh format -t $KAFKA_CLUSTER_ID -c config/kraft/server.properties ``` You run this on **every** node (broker, controller, or combined), pointing `-c` at that node's config file so it picks up the node's `log.dirs`. ## What format writes Into each directory listed in `log.dirs` (and `metadata.log.dir` if separate), format writes a **`meta.properties`** file. It contains: - `cluster.id` — the shared UUID you passed with `-t`. - `node.id` — taken from the config's `node.id`. - `version` — the on-disk format version (`version=1` for legacy, the newer KRaft layout uses a different version and a `directory.id`). ## Why it's mandatory The node treats `meta.properties` as proof that the directory was deliberately initialized for **this** cluster. On startup it reads the file: - **Missing** → the node exits with a fatal error telling you to run `kafka-storage.sh format`. - **cluster.id mismatch** between directories or against the quorum → `InconsistentClusterIdException`; the node refuses to start. This is a safety guard: it stops you from accidentally pooling disks or nodes from two different clusters. ## Edge cases - Re-running format on an already-formatted directory fails unless you pass `--ignore-formatted`, which skips already-formatted dirs (useful when adding a new log dir to an existing node). - Each node must use the **same** cluster.id but its **own** distinct `node.id`. - Combined-mode nodes (`process.roles=broker,controller`) are formatted exactly the same way.

  • Who used to assign the cluster.id before KRaft, and why does that matter now?
    ZooKeeper assigned it automatically when a broker first connected. In KRaft there is no ZooKeeper, so the operator generates it via kafka-storage.sh random-uuid and passes it to format. That's why formatting is a new mandatory step.
  • What happens if two nodes are formatted with different cluster.ids?
    They cannot form one cluster; the node detects the mismatch and fails with InconsistentClusterIdException, refusing to join the quorum.

saying these in an interview costs you the question

  • Saying KRaft still uses ZooKeeper to get the cluster.id.
  • Claiming format is optional or auto-runs on first boot.
  • Thinking each node should get a different cluster.id (it's the node.id that differs, not the cluster.id).
  • Confusing cluster.id with node.id.

context

open as a page

Show the kafka-storage.sh commands to bootstrap a KRaft cluster's storage and explain the key meta.properties fields and common formatting pitfalls.

level: middleimportance: must knowfreq 55%

basics

~10 s

Generate one cluster id with kafka-storage.sh random-uuid, then run kafka-storage.sh format -t <id> -c <config> on each node. That writes meta.properties holding cluster.id, node.id, and version. Use --ignore-formatted to skip already-formatted dirs.

open as a page

Where does KRaft store the cluster metadata on disk, and what is the on-disk layout of the __cluster_metadata log directory?

level: middleimportance: should knowfreq 45%

basics

~10 s

KRaft stores metadata in an internal single-partition topic, __cluster_metadata-0, written as a normal Kafka log (segments, indexes) plus snapshot files. By default it lives under log.dirs; you can move it with metadata.log.dir.

open as a page

Walk through what happens during first-boot quorum formation in a KRaft cluster: how do controllers find each other, elect a leader, and reach the point where brokers can register?

level: seniorimportance: should knowfreq 28%

basics

~20 s

On first boot each controller reads its formatted meta.properties (cluster.id, node.id) and the voter set (static voters config or bootstrap.checkpoint). Controllers contact each other on the controller listener, run a Raft election to pick an active controller (leader), and once a quorum agrees, brokers can register and the cluster is live.

open as a page

Explain how --initial-controllers and the bootstrap.checkpoint mechanism work for first-boot quorum bootstrap in newer Kafka versions, and how that differs from controller.quorum.voters.

level: seniorimportance: should knowfreq 30%

basics

~20 s

Older KRaft uses a static controller.quorum.voters list. Newer Kafka (KIP-853) supports dynamic quorums: at format time you pass --initial-controllers (or --standalone) so the formatter writes a bootstrap.checkpoint snapshot recording the initial voter set, and the quorum bootstraps from that instead of a static config line.

open as a page