skip to content

How do you finalize a KIP-866 migration and verify the cluster is fully on KRaft?

level: middleimportance: should knowfreq 35%

answer

  1. finalize only after ALL brokers are KRaft
  2. remove migration flag + zookeeper.connect on controllers
  3. ZkMigrationState gauge -> completed
  4. kafka-metadata-quorum.sh describe --status
  5. decommission ZK last

basics

~20 s

After every broker is in KRaft mode, finalize by removing zookeeper.metadata.migration.enable (and zookeeper.connect) from the controllers and restarting them as a pure KRaft quorum. Verify via the migration/controller metrics showing the dual-write phase ended, then decommission ZooKeeper.

solid answer

~40 s

Finalization is the last step, done only after all brokers run as KRaft brokers. On each controller you remove zookeeper.metadata.migration.enable=true and the zookeeper.connect, then restart the controller quorum so it runs as pure KRaft and stops writing to ZooKeeper. You verify completion by checking the controller's migration-state metric (the KafkaController/ZkMigrationState gauge moves to the post-migration/'NONE' or completed state) and confirming no broker still references ZooKeeper. Functional checks: kafka-metadata-quorum.sh describe --status shows a healthy quorum, topics/ACLs/configs are all present, and clients operate normally. Only after these pass do you shut down and remove the ZooKeeper ensemble. Skipping the verification or decommissioning ZK before finalize is a common mistake.

go deeper

for a junior

Know finalize = remove the migration flag and restart controllers as pure KRaft.

for a middle

List the finalize steps, the verification checks, and the ZK-last decommission order.

for a senior

Explain the precondition (all brokers KRaft) and which metric/CLI confirm completion.

for a principal

Codify a verification runbook and gating criteria before irreversible finalize and ZK teardown.

## Precondition: all brokers must already be KRaft brokers Finalization is **not** a step you do mid-roll. Every broker must already be running as a pure KRaft broker (`process.roles=broker`, no `zookeeper.connect`). If any broker is still in ZK mode, finalizing orphans it because the controller stops mirroring to ZooKeeper. ## The finalize action On each KRaft **controller**: 1. Remove `zookeeper.metadata.migration.enable=true`. 2. Remove `zookeeper.connect` (the controller no longer needs ZooKeeper). 3. Restart the controller quorum. Now the controllers run as a normal KRaft quorum. The controller **stops dual-writing to ZooKeeper**, which formally completes the migration. This is the irreversible step. ## Verifying the migration is complete - **Migration-state metric**: the controller exposes a `ZkMigrationState` gauge (under the KafkaController metrics). During migration it reports a migrating/dual-write state; after finalize it reports the completed/non-migrating state. Confirm it has advanced. - **Quorum health**: `bin/kafka-metadata-quorum.sh --bootstrap-server <broker> describe --status` should show all controller voters, a stable leader, and the high-water mark advancing (no large lag). - **Metadata integrity**: list topics, ACLs, configs, and quotas — verify counts match what existed in ZK. - **No ZK references**: grep broker/controller configs; none should still contain `zookeeper.connect` or the migration flag. - **Client smoke test**: produce/consume on a canary topic; create/delete a topic to confirm the KRaft controller handles writes. ## Decommission ZooKeeper Only after verification: stop the ZooKeeper ensemble and remove it from automation/monitoring. Because finalize stopped dual-writes, ZK is already stale and safe to delete. ## Common mistakes - **Decommissioning ZK before finalize** — the controller is still trying to dual-write and will error. - **Leaving the migration flag set** on controllers — the cluster stays in dual-write mode indefinitely, paying the ZK-write cost for nothing. - **Not verifying the migration-state metric**, then assuming success. - **Forgetting to remove zookeeper.connect** from controllers after finalize.

  • What metric confirms the controller has left dual-write mode?
    The KafkaController ZkMigrationState gauge — it advances from a migrating/dual-write value to the completed/non-migrating state after finalize.
  • What happens if you decommission ZooKeeper before finalizing?
    The controller is still in migration mode and tries to dual-write to ZooKeeper; with ZK gone those writes fail, destabilizing the controller. Always finalize first, then remove ZK.

saying these in an interview costs you the question

  • Finalizing while some brokers are still in ZK mode — it orphans them.
  • Removing ZooKeeper before finalize, breaking the controller's dual-writes.
  • Leaving zookeeper.metadata.migration.enable set, keeping the cluster stuck in dual-write.
  • Declaring success without checking the ZkMigrationState metric or quorum status.

context