skip to content

ZooKeeper-to-KRaft Migration

The dual-write migration path from ZooKeeper to KRaft: enabling migration, rolling brokers, finalizing, and the rollback window. A very practical question for anyone who has upgraded a real cluster.

part ofApache Kafkaoverview, primer and where to startread it →
on this pageshow

questions

6

What is the ZooKeeper-to-KRaft migration (KIP-866) and why does Kafka need it?

level: juniorimportance: must knowfreq 70%

answer

  1. ZK metadata -> KRaft log
  2. KIP-866 dual-write, online
  3. ZK removed in Kafka 4.0
  4. __cluster_metadata, controller quorum
  5. rollback window while ZK written

basics

~10 s

It is the online process that moves a Kafka cluster's metadata from ZooKeeper to KRaft (Kafka's built-in Raft controller) without downtime. Kafka needs it because ZooKeeper is deprecated and removed in Kafka 4.0.

solid answer

~40 s

KIP-866 defines a live, rolling migration that moves cluster metadata (topics, partitions, ACLs, configs) from an external ZooKeeper ensemble into KRaft, Kafka's self-managed Raft-based metadata quorum. Historically Kafka stored all metadata in ZooKeeper; KRaft replaces it with an internal controller quorum that persists metadata in the __cluster_metadata log. ZooKeeper was deprecated in Kafka 3.5 and removed entirely in 4.0, so every ZK-based cluster must migrate to stay on a supported version. The migration is designed to be online: you stand up a KRaft controller quorum, enter a dual-write phase where the controller writes to both KRaft and ZooKeeper, roll brokers from ZK mode to KRaft mode, then finalize. It avoids a flag-day cutover and preserves a rollback window while ZK is still being written.

go deeper

for a junior

Know it moves metadata from ZooKeeper to KRaft, is online, and ZK is gone in 4.0.

for a middle

Be able to name the dual-write phases and the __cluster_metadata log.

for a senior

Explain why online migration and the rollback window matter operationally, plus version gates.

for a principal

Frame the migration in upgrade-roadmap terms: version prerequisites, risk, and decommission strategy across a fleet.

## Background: what ZooKeeper did for Kafka For most of Kafka's history, the cluster relied on **Apache ZooKeeper** — a separate distributed coordination service — to store all *cluster metadata*: the list of topics, their partitions and replica assignments, in-sync replica (ISR) sets, controller election, broker registration, ACLs, and dynamic configs. One broker was elected **controller** and was responsible for reading/writing this metadata in ZooKeeper and propagating it to other brokers. This design had drawbacks: ZooKeeper is a second system to operate and secure, metadata propagation was slow at scale (the controller had to fetch and broadcast full state), and very large clusters hit scaling limits on partition counts and failover time. ## KRaft: Kafka's replacement for ZooKeeper **KRaft** (Kafka Raft) moves metadata management *inside* Kafka. A dedicated **controller quorum** (typically 3 or 5 nodes) runs the **Raft consensus protocol** and stores metadata as an event log in an internal topic, `__cluster_metadata`. Brokers consume this log to learn the cluster state. There is no external ZooKeeper. KRaft was introduced in KIP-500, became production-ready in Kafka 3.3, ZooKeeper mode was deprecated in **3.5**, and ZooKeeper support was **removed in Kafka 4.0**. So migration is mandatory for anyone who wants to upgrade past 3.x. ## What KIP-866 adds KIP-866 is the **dual-write migration** mechanism. The core problem: you cannot simply stop ZooKeeper and start KRaft, because the metadata lives in ZooKeeper and the brokers are running. KIP-866 lets you migrate *online* (no full downtime) by: 1. **Provisioning a KRaft controller quorum** configured in migration mode. 2. The KRaft controller **copies existing metadata out of ZooKeeper** into the KRaft log. 3. Entering a **dual-write phase**: the KRaft controller becomes the active controller but writes every metadata change to *both* KRaft *and* ZooKeeper. This keeps ZK in sync so you can still roll back. 4. **Rolling the brokers** one at a time from ZooKeeper mode into KRaft mode. 5. **Finalizing** the migration, after which ZooKeeper is no longer written to and can be decommissioned. ## Why online matters Kafka often backs critical pipelines; a multi-hour outage to migrate is unacceptable. The dual-write design means producers and consumers keep running throughout, and operators retain a **rollback window** while ZooKeeper is still receiving writes. ## Edge cases / caveats - All brokers and controllers must run a Kafka version that supports migration (3.4+ for the feature, with 3.5/3.6 strongly recommended for stability). - The cluster's **metadata version** (`inter.broker.protocol.version` in ZK mode; the `metadata.version` feature in KRaft) must be high enough. - Migration is a one-way *intent*: once finalized, you cannot return to ZooKeeper.

  • In which Kafka version was ZooKeeper removed entirely?
    Kafka 4.0. It was deprecated in 3.5, and KRaft became production-ready in 3.3.
  • Where does KRaft store metadata instead of ZooKeeper?
    In the internal __cluster_metadata log, replicated across the controller quorum via the Raft protocol.

saying these in an interview costs you the question

  • Saying migration requires full cluster downtime — it is an online rolling process.
  • Claiming KRaft still uses ZooKeeper under the hood — it removes ZooKeeper entirely.
  • Confusing KRaft with a ZooKeeper plugin; it is a built-in Raft-based controller quorum.

context

open as a page

How do you provision the KRaft controller quorum and enable migration mode for a KIP-866 migration?

level: middleimportance: must knowfreq 55%

basics

~10 s

Start a new KRaft controller quorum with process.roles=controller and zookeeper.metadata.migration.enable=true, pointing it at the existing ZooKeeper. It must reuse the cluster's existing cluster.id and connect to the same ZK.

open as a page

Explain the dual-write phase and how brokers are rolled from ZooKeeper mode to KRaft mode during migration.

level: seniorimportance: must knowfreq 50%

basics

~20 s

In dual-write, the active KRaft controller writes every metadata change to both KRaft and ZooKeeper, keeping them in sync. Then brokers are restarted one at a time into KRaft mode (process.roles=broker, no zookeeper.connect) until all are migrated.

open as a page

How do you finalize a KIP-866 migration and verify the cluster is fully on KRaft?

level: middleimportance: should knowfreq 35%

basics

~20 s

After every broker is in KRaft mode, finalize by removing zookeeper.metadata.migration.enable (and zookeeper.connect) from the controllers and restarting them as a pure KRaft quorum. Verify via the migration/controller metrics showing the dual-write phase ended, then decommission ZooKeeper.

open as a page

When can you roll back a KIP-866 migration, and what makes the rollback window close?

level: seniorimportance: should knowfreq 40%

basics

~20 s

You can roll back any time during the dual-write phase, because ZooKeeper is still a current mirror. The window closes once you finalize the migration by taking the controllers out of migration mode; after that ZooKeeper is no longer written and you cannot return.

open as a page

What version gates and prerequisites must be satisfied before attempting a ZK-to-KRaft migration, and what limitations apply?

level: principalimportance: should knowfreq 38%

basics

~20 s

All brokers and controllers must run a Kafka version supporting migration (3.4+, ideally 3.6+), with metadata.version/inter.broker.protocol.version at least 3.4-IV0. The cluster's brokers must be on a recent IBP, and certain features (e.g., older ZK-only behaviors) must be cleared first.

open as a page