skip to content

Why must you run kafka-storage.sh format before starting a KRaft-mode Kafka broker or controller, and what does it create?

level: juniorimportance: must knowfreq 70%

answer

  1. No ZooKeeper → no auto cluster.id
  2. random-uuid then format -t -c
  3. meta.properties = cluster.id + node.id + version
  4. Missing file → node won't start
  5. Mismatch → InconsistentClusterIdException

basics

~20 s

In KRaft mode you must format each node's log directory first. kafka-storage.sh format writes a meta.properties file containing a shared cluster.id and the node.id, so the node knows which cluster it belongs to. Without it, the node refuses to start.

solid answer

~40 s

KRaft removes ZooKeeper, which previously assigned the cluster.id automatically on first connect. In KRaft, you generate a cluster.id yourself (kafka-storage.sh random-uuid) and then run kafka-storage.sh format -t <cluster-id> -c server.properties on every node. Format writes a meta.properties file into each configured log directory (log.dirs / metadata.log.dir) holding the cluster.id, the node.id, and a version. This is a one-time bootstrap step per node before first boot. If a node starts with an unformatted directory, it logs 'No `meta.properties` found' and exits with a non-zero code. The same cluster.id must be used across all nodes of one cluster; a mismatch causes the node to refuse to join (InconsistentClusterIdException), which prevents accidentally mixing nodes from different clusters in the same quorum.

go deeper

for a junior

Know that you must format storage first and that it writes meta.properties with the cluster.id; without it the node won't start.

for a middle

Explain the random-uuid + format -t -c two-step and the contents of meta.properties.

for a senior

Discuss InconsistentClusterIdException, --ignore-formatted, per-node vs shared identifiers, and the ZK→KRaft rationale.

for a principal

Frame meta.properties as the durable identity contract that prevents disk/node cross-pollination across clusters and how this replaces ZK's implicit identity assignment.

## Background: what changed with KRaft Apache Kafka historically depended on **ZooKeeper**, a separate coordination service, to store cluster metadata (which brokers exist, topic configs, partition leaders, etc.). When a broker first connected to a ZooKeeper ensemble, ZooKeeper handed it a **cluster.id** — a unique identifier for that Kafka cluster — and the broker persisted it locally. You never had to format anything by hand. **KRaft** (Kafka Raft) removes ZooKeeper. Metadata is now stored inside Kafka itself, in an internal topic called **`__cluster_metadata`**, managed by a built-in Raft consensus quorum of **controller** nodes. Because there is no ZooKeeper to mint the cluster.id and bootstrap the storage, you must do that explicitly before the first start. ## The two-step bootstrap 1. **Generate a cluster id** (do this once for the whole cluster): ``` KAFKA_CLUSTER_ID="$(bin/kafka-storage.sh random-uuid)" ``` This prints a base64-encoded 128-bit UUID, e.g. `MkU3OEVBNTcwNTJENDM2Qk`. 2. **Format each node's storage** using that id: ``` bin/kafka-storage.sh format -t $KAFKA_CLUSTER_ID -c config/kraft/server.properties ``` You run this on **every** node (broker, controller, or combined), pointing `-c` at that node's config file so it picks up the node's `log.dirs`. ## What format writes Into each directory listed in `log.dirs` (and `metadata.log.dir` if separate), format writes a **`meta.properties`** file. It contains: - `cluster.id` — the shared UUID you passed with `-t`. - `node.id` — taken from the config's `node.id`. - `version` — the on-disk format version (`version=1` for legacy, the newer KRaft layout uses a different version and a `directory.id`). ## Why it's mandatory The node treats `meta.properties` as proof that the directory was deliberately initialized for **this** cluster. On startup it reads the file: - **Missing** → the node exits with a fatal error telling you to run `kafka-storage.sh format`. - **cluster.id mismatch** between directories or against the quorum → `InconsistentClusterIdException`; the node refuses to start. This is a safety guard: it stops you from accidentally pooling disks or nodes from two different clusters. ## Edge cases - Re-running format on an already-formatted directory fails unless you pass `--ignore-formatted`, which skips already-formatted dirs (useful when adding a new log dir to an existing node). - Each node must use the **same** cluster.id but its **own** distinct `node.id`. - Combined-mode nodes (`process.roles=broker,controller`) are formatted exactly the same way.

  • Who used to assign the cluster.id before KRaft, and why does that matter now?
    ZooKeeper assigned it automatically when a broker first connected. In KRaft there is no ZooKeeper, so the operator generates it via kafka-storage.sh random-uuid and passes it to format. That's why formatting is a new mandatory step.
  • What happens if two nodes are formatted with different cluster.ids?
    They cannot form one cluster; the node detects the mismatch and fails with InconsistentClusterIdException, refusing to join the quorum.

saying these in an interview costs you the question

  • Saying KRaft still uses ZooKeeper to get the cluster.id.
  • Claiming format is optional or auto-runs on first boot.
  • Thinking each node should get a different cluster.id (it's the node.id that differs, not the cluster.id).
  • Confusing cluster.id with node.id.

context