How do you finalize a KIP-866 migration and verify the cluster is fully on KRaft?
answer
- finalize only after ALL brokers are KRaft
- remove migration flag + zookeeper.connect on controllers
- ZkMigrationState gauge -> completed
- kafka-metadata-quorum.sh describe --status
- decommission ZK last
basics
~20 sAfter every broker is in KRaft mode, finalize by removing zookeeper.metadata.migration.enable (and zookeeper.connect) from the controllers and restarting them as a pure KRaft quorum. Verify via the migration/controller metrics showing the dual-write phase ended, then decommission ZooKeeper.
solid answer
~40 sFinalization is the last step, done only after all brokers run as KRaft brokers. On each controller you remove zookeeper.metadata.migration.enable=true and the zookeeper.connect, then restart the controller quorum so it runs as pure KRaft and stops writing to ZooKeeper. You verify completion by checking the controller's migration-state metric (the KafkaController/ZkMigrationState gauge moves to the post-migration/'NONE' or completed state) and confirming no broker still references ZooKeeper. Functional checks: kafka-metadata-quorum.sh describe --status shows a healthy quorum, topics/ACLs/configs are all present, and clients operate normally. Only after these pass do you shut down and remove the ZooKeeper ensemble. Skipping the verification or decommissioning ZK before finalize is a common mistake.
go deeper
Know finalize = remove the migration flag and restart controllers as pure KRaft.
List the finalize steps, the verification checks, and the ZK-last decommission order.
Explain the precondition (all brokers KRaft) and which metric/CLI confirm completion.
Codify a verification runbook and gating criteria before irreversible finalize and ZK teardown.
## Precondition: all brokers must already be KRaft brokers Finalization is **not** a step you do mid-roll. Every broker must already be running as a pure KRaft broker (`process.roles=broker`, no `zookeeper.connect`). If any broker is still in ZK mode, finalizing orphans it because the controller stops mirroring to ZooKeeper. ## The finalize action On each KRaft **controller**: 1. Remove `zookeeper.metadata.migration.enable=true`. 2. Remove `zookeeper.connect` (the controller no longer needs ZooKeeper). 3. Restart the controller quorum. Now the controllers run as a normal KRaft quorum. The controller **stops dual-writing to ZooKeeper**, which formally completes the migration. This is the irreversible step. ## Verifying the migration is complete - **Migration-state metric**: the controller exposes a `ZkMigrationState` gauge (under the KafkaController metrics). During migration it reports a migrating/dual-write state; after finalize it reports the completed/non-migrating state. Confirm it has advanced. - **Quorum health**: `bin/kafka-metadata-quorum.sh --bootstrap-server <broker> describe --status` should show all controller voters, a stable leader, and the high-water mark advancing (no large lag). - **Metadata integrity**: list topics, ACLs, configs, and quotas — verify counts match what existed in ZK. - **No ZK references**: grep broker/controller configs; none should still contain `zookeeper.connect` or the migration flag. - **Client smoke test**: produce/consume on a canary topic; create/delete a topic to confirm the KRaft controller handles writes. ## Decommission ZooKeeper Only after verification: stop the ZooKeeper ensemble and remove it from automation/monitoring. Because finalize stopped dual-writes, ZK is already stale and safe to delete. ## Common mistakes - **Decommissioning ZK before finalize** — the controller is still trying to dual-write and will error. - **Leaving the migration flag set** on controllers — the cluster stays in dual-write mode indefinitely, paying the ZK-write cost for nothing. - **Not verifying the migration-state metric**, then assuming success. - **Forgetting to remove zookeeper.connect** from controllers after finalize.
- What metric confirms the controller has left dual-write mode?The KafkaController ZkMigrationState gauge — it advances from a migrating/dual-write value to the completed/non-migrating state after finalize.
- What happens if you decommission ZooKeeper before finalizing?The controller is still in migration mode and tries to dual-write to ZooKeeper; with ZK gone those writes fail, destabilizing the controller. Always finalize first, then remove ZK.
saying these in an interview costs you the question
- Finalizing while some brokers are still in ZK mode — it orphans them.
- Removing ZooKeeper before finalize, breaking the controller's dual-writes.
- Leaving zookeeper.metadata.migration.enable set, keeping the cluster stuck in dual-write.
- Declaring success without checking the ZkMigrationState metric or quorum status.