skip to content

Show the kafka-storage.sh commands to bootstrap a KRaft cluster's storage and explain the key meta.properties fields and common formatting pitfalls.

level: middleimportance: must knowfreq 55%

answer

  1. random-uuid once → format -t -c per node
  2. meta.properties: cluster.id, node.id, version, directory.id
  3. --ignore-formatted = idempotent / add a dir
  4. --release-version pins metadata.version
  5. --add-scram seeds first SCRAM user
  6. Same cluster.id, unique node.id

basics

~10 s

Generate one cluster id with kafka-storage.sh random-uuid, then run kafka-storage.sh format -t <id> -c <config> on each node. That writes meta.properties holding cluster.id, node.id, and version. Use --ignore-formatted to skip already-formatted dirs.

solid answer

~40 s

Bootstrap is two commands. First, once per cluster: KAFKA_CLUSTER_ID=$(bin/kafka-storage.sh random-uuid). Then on every node: bin/kafka-storage.sh format -t $KAFKA_CLUSTER_ID -c config/server.properties, which reads the node's log.dirs/metadata.log.dir and writes a meta.properties into each. Key fields: cluster.id (shared across the whole cluster), node.id (unique per node), version (on-disk layout; the newer KRaft layout adds directory.id), and for dynamic quorums a directory.id UUID. Pitfalls: reusing or mistyping the cluster.id across nodes causes InconsistentClusterIdException; re-running format on an initialized dir fails unless you pass --ignore-formatted (or --add-scram for SCRAM creds, --release-version to pin metadata.version). For dynamic quorums you also pass --standalone or --initial-controllers so the formatter writes bootstrap.checkpoint. Each node must use the same -t cluster id but its own node.id from its config.

go deeper

for a junior

Know the random-uuid then format -t -c sequence and that it writes meta.properties.

for a middle

Explain meta.properties fields, --ignore-formatted, and same-cluster-id/unique-node-id rules.

for a senior

Cover --release-version/metadata.version pinning, --add-scram, directory.id, and static-vs-dynamic format flags.

for a principal

Standardize the bootstrap as automation: idempotent formatting, version floors, credential seeding, and failure-mode diagnostics.

## The two commands **1. Generate the cluster id (once for the whole cluster):** ``` KAFKA_CLUSTER_ID="$(bin/kafka-storage.sh random-uuid)" ``` `random-uuid` prints a base64-encoded 128-bit UUID. This single value identifies the entire cluster and must be reused on every node. **2. Format each node's storage:** ``` bin/kafka-storage.sh format -t "$KAFKA_CLUSTER_ID" -c config/server.properties ``` - `-t/--cluster-id` — the shared cluster id. - `-c/--config` — the node's config file; the formatter reads `log.dirs` and `metadata.log.dir` from it and initializes each directory. Run this on **every** broker/controller, each pointing `-c` at its own config so it picks up that node's `node.id` and dirs. ## meta.properties fields After formatting, each directory has a `meta.properties` (Java properties format). Important keys: - **`cluster.id`** — the shared UUID; identical on all nodes/dirs of the cluster. - **`node.id`** — this node's id (from config); unique per node. - **`version`** — on-disk metadata format version. The newer KRaft layout (`version` for dynamic quorums) also includes **`directory.id`**. - **`directory.id`** — a per-directory UUID used by dynamic quorums to identify a voter's storage; also helps detect a swapped/replaced disk. (In legacy ZK-mode meta.properties you'd see `broker.id` instead of `node.id`; KRaft uses `node.id`.) ## Useful format flags - **`--ignore-formatted`** — skip directories that already have `meta.properties` instead of erroring; lets you add a new log dir to an existing node, or re-run idempotently. - **`--release-version <ver>`** (a.k.a. pinning `metadata.version`) — set the initial **metadata.version** feature level so a new cluster doesn't auto-pick the absolute latest IBP/feature level. - **`--add-scram 'SCRAM-SHA-256=[name=...,password=...]'`** — pre-create a SCRAM credential at format time (useful so the first admin user exists before any broker is up, since there's no ZooKeeper to add it to). - **`--standalone` / `--initial-controllers` / `--no-initial-controllers`** — choose the dynamic-quorum seed and trigger writing of `bootstrap.checkpoint`. ## Common pitfalls 1. **Different cluster.id per node** — typo or regenerating the UUID per node. Result: nodes can't join; `InconsistentClusterIdException`. Fix: generate once, export, reuse. 2. **Duplicate node.id** — two nodes share a `node.id`. Causes conflicts/fencing. Each node.id must be unique. 3. **Re-formatting a live dir** — running `format` again without `--ignore-formatted` errors out (protecting existing data). To wipe and start over you must delete the directory contents deliberately. 4. **Forgetting metadata.version** — letting the cluster default to the newest feature level can block a later downgrade; pin it with `--release-version` if you need a known floor. 5. **Mixing static and dynamic** — passing `--initial-controllers` while also setting `controller.quorum.voters` is contradictory; pick one membership model. ## Verifying You can inspect what was written with: ``` bin/kafka-storage.sh info -c config/server.properties ``` which reports each dir's formatted state and ids, and `kafka-metadata-quorum.sh ... describe` to see the quorum once running.

  • When would you use --ignore-formatted?
    When some directories are already formatted and you want format to initialize only the new/empty ones without erroring — e.g. adding a fresh log dir to an existing node, or making a format step idempotent in automation.
  • Why might you pass --release-version (pin metadata.version) at format time?
    A fresh cluster otherwise adopts the latest available metadata.version/feature level. Pinning a known version keeps a defined floor, preserves the ability to downgrade later, and avoids surprises from newly-enabled features.
  • How do you seed an initial admin SCRAM credential when there's no ZooKeeper?
    Pass --add-scram 'SCRAM-SHA-256=[name=admin,password=...]' to kafka-storage.sh format so the credential is written into the bootstrap metadata, making the user exist before any broker starts.

saying these in an interview costs you the question

  • Generating a new cluster.id per node instead of reusing one.
  • Saying meta.properties stores broker.id in KRaft (it's node.id).
  • Believing re-running format silently wipes and reinitializes a populated dir.
  • Forgetting that -c reads log.dirs/metadata.log.dir from the node's config.
  • Setting both controller.quorum.voters and --initial-controllers.

context