Why must you run kafka-storage.sh format before starting a KRaft-mode Kafka broker or controller, and what does it create?
answer
- No ZooKeeper → no auto cluster.id
- random-uuid then format -t -c
- meta.properties = cluster.id + node.id + version
- Missing file → node won't start
- Mismatch → InconsistentClusterIdException
basics
~20 sIn KRaft mode you must format each node's log directory first. kafka-storage.sh format writes a meta.properties file containing a shared cluster.id and the node.id, so the node knows which cluster it belongs to. Without it, the node refuses to start.
solid answer
~40 sKRaft removes ZooKeeper, which previously assigned the cluster.id automatically on first connect. In KRaft, you generate a cluster.id yourself (kafka-storage.sh random-uuid) and then run kafka-storage.sh format -t <cluster-id> -c server.properties on every node. Format writes a meta.properties file into each configured log directory (log.dirs / metadata.log.dir) holding the cluster.id, the node.id, and a version. This is a one-time bootstrap step per node before first boot. If a node starts with an unformatted directory, it logs 'No `meta.properties` found' and exits with a non-zero code. The same cluster.id must be used across all nodes of one cluster; a mismatch causes the node to refuse to join (InconsistentClusterIdException), which prevents accidentally mixing nodes from different clusters in the same quorum.
go deeper
Know that you must format storage first and that it writes meta.properties with the cluster.id; without it the node won't start.
Explain the random-uuid + format -t -c two-step and the contents of meta.properties.
Discuss InconsistentClusterIdException, --ignore-formatted, per-node vs shared identifiers, and the ZK→KRaft rationale.
Frame meta.properties as the durable identity contract that prevents disk/node cross-pollination across clusters and how this replaces ZK's implicit identity assignment.
## Background: what changed with KRaft Apache Kafka historically depended on **ZooKeeper**, a separate coordination service, to store cluster metadata (which brokers exist, topic configs, partition leaders, etc.). When a broker first connected to a ZooKeeper ensemble, ZooKeeper handed it a **cluster.id** — a unique identifier for that Kafka cluster — and the broker persisted it locally. You never had to format anything by hand. **KRaft** (Kafka Raft) removes ZooKeeper. Metadata is now stored inside Kafka itself, in an internal topic called **`__cluster_metadata`**, managed by a built-in Raft consensus quorum of **controller** nodes. Because there is no ZooKeeper to mint the cluster.id and bootstrap the storage, you must do that explicitly before the first start. ## The two-step bootstrap 1. **Generate a cluster id** (do this once for the whole cluster): ``` KAFKA_CLUSTER_ID="$(bin/kafka-storage.sh random-uuid)" ``` This prints a base64-encoded 128-bit UUID, e.g. `MkU3OEVBNTcwNTJENDM2Qk`. 2. **Format each node's storage** using that id: ``` bin/kafka-storage.sh format -t $KAFKA_CLUSTER_ID -c config/kraft/server.properties ``` You run this on **every** node (broker, controller, or combined), pointing `-c` at that node's config file so it picks up the node's `log.dirs`. ## What format writes Into each directory listed in `log.dirs` (and `metadata.log.dir` if separate), format writes a **`meta.properties`** file. It contains: - `cluster.id` — the shared UUID you passed with `-t`. - `node.id` — taken from the config's `node.id`. - `version` — the on-disk format version (`version=1` for legacy, the newer KRaft layout uses a different version and a `directory.id`). ## Why it's mandatory The node treats `meta.properties` as proof that the directory was deliberately initialized for **this** cluster. On startup it reads the file: - **Missing** → the node exits with a fatal error telling you to run `kafka-storage.sh format`. - **cluster.id mismatch** between directories or against the quorum → `InconsistentClusterIdException`; the node refuses to start. This is a safety guard: it stops you from accidentally pooling disks or nodes from two different clusters. ## Edge cases - Re-running format on an already-formatted directory fails unless you pass `--ignore-formatted`, which skips already-formatted dirs (useful when adding a new log dir to an existing node). - Each node must use the **same** cluster.id but its **own** distinct `node.id`. - Combined-mode nodes (`process.roles=broker,controller`) are formatted exactly the same way.
- Who used to assign the cluster.id before KRaft, and why does that matter now?ZooKeeper assigned it automatically when a broker first connected. In KRaft there is no ZooKeeper, so the operator generates it via kafka-storage.sh random-uuid and passes it to format. That's why formatting is a new mandatory step.
- What happens if two nodes are formatted with different cluster.ids?They cannot form one cluster; the node detects the mismatch and fails with InconsistentClusterIdException, refusing to join the quorum.
saying these in an interview costs you the question
- Saying KRaft still uses ZooKeeper to get the cluster.id.
- Claiming format is optional or auto-runs on first boot.
- Thinking each node should get a different cluster.id (it's the node.id that differs, not the cluster.id).
- Confusing cluster.id with node.id.