skip to content

What is the ZooKeeper-to-KRaft migration (KIP-866) and why does Kafka need it?

level: juniorimportance: must knowfreq 70%

answer

  1. ZK metadata -> KRaft log
  2. KIP-866 dual-write, online
  3. ZK removed in Kafka 4.0
  4. __cluster_metadata, controller quorum
  5. rollback window while ZK written

basics

~10 s

It is the online process that moves a Kafka cluster's metadata from ZooKeeper to KRaft (Kafka's built-in Raft controller) without downtime. Kafka needs it because ZooKeeper is deprecated and removed in Kafka 4.0.

solid answer

~40 s

KIP-866 defines a live, rolling migration that moves cluster metadata (topics, partitions, ACLs, configs) from an external ZooKeeper ensemble into KRaft, Kafka's self-managed Raft-based metadata quorum. Historically Kafka stored all metadata in ZooKeeper; KRaft replaces it with an internal controller quorum that persists metadata in the __cluster_metadata log. ZooKeeper was deprecated in Kafka 3.5 and removed entirely in 4.0, so every ZK-based cluster must migrate to stay on a supported version. The migration is designed to be online: you stand up a KRaft controller quorum, enter a dual-write phase where the controller writes to both KRaft and ZooKeeper, roll brokers from ZK mode to KRaft mode, then finalize. It avoids a flag-day cutover and preserves a rollback window while ZK is still being written.

go deeper

for a junior

Know it moves metadata from ZooKeeper to KRaft, is online, and ZK is gone in 4.0.

for a middle

Be able to name the dual-write phases and the __cluster_metadata log.

for a senior

Explain why online migration and the rollback window matter operationally, plus version gates.

for a principal

Frame the migration in upgrade-roadmap terms: version prerequisites, risk, and decommission strategy across a fleet.

## Background: what ZooKeeper did for Kafka For most of Kafka's history, the cluster relied on **Apache ZooKeeper** — a separate distributed coordination service — to store all *cluster metadata*: the list of topics, their partitions and replica assignments, in-sync replica (ISR) sets, controller election, broker registration, ACLs, and dynamic configs. One broker was elected **controller** and was responsible for reading/writing this metadata in ZooKeeper and propagating it to other brokers. This design had drawbacks: ZooKeeper is a second system to operate and secure, metadata propagation was slow at scale (the controller had to fetch and broadcast full state), and very large clusters hit scaling limits on partition counts and failover time. ## KRaft: Kafka's replacement for ZooKeeper **KRaft** (Kafka Raft) moves metadata management *inside* Kafka. A dedicated **controller quorum** (typically 3 or 5 nodes) runs the **Raft consensus protocol** and stores metadata as an event log in an internal topic, `__cluster_metadata`. Brokers consume this log to learn the cluster state. There is no external ZooKeeper. KRaft was introduced in KIP-500, became production-ready in Kafka 3.3, ZooKeeper mode was deprecated in **3.5**, and ZooKeeper support was **removed in Kafka 4.0**. So migration is mandatory for anyone who wants to upgrade past 3.x. ## What KIP-866 adds KIP-866 is the **dual-write migration** mechanism. The core problem: you cannot simply stop ZooKeeper and start KRaft, because the metadata lives in ZooKeeper and the brokers are running. KIP-866 lets you migrate *online* (no full downtime) by: 1. **Provisioning a KRaft controller quorum** configured in migration mode. 2. The KRaft controller **copies existing metadata out of ZooKeeper** into the KRaft log. 3. Entering a **dual-write phase**: the KRaft controller becomes the active controller but writes every metadata change to *both* KRaft *and* ZooKeeper. This keeps ZK in sync so you can still roll back. 4. **Rolling the brokers** one at a time from ZooKeeper mode into KRaft mode. 5. **Finalizing** the migration, after which ZooKeeper is no longer written to and can be decommissioned. ## Why online matters Kafka often backs critical pipelines; a multi-hour outage to migrate is unacceptable. The dual-write design means producers and consumers keep running throughout, and operators retain a **rollback window** while ZooKeeper is still receiving writes. ## Edge cases / caveats - All brokers and controllers must run a Kafka version that supports migration (3.4+ for the feature, with 3.5/3.6 strongly recommended for stability). - The cluster's **metadata version** (`inter.broker.protocol.version` in ZK mode; the `metadata.version` feature in KRaft) must be high enough. - Migration is a one-way *intent*: once finalized, you cannot return to ZooKeeper.

  • In which Kafka version was ZooKeeper removed entirely?
    Kafka 4.0. It was deprecated in 3.5, and KRaft became production-ready in 3.3.
  • Where does KRaft store metadata instead of ZooKeeper?
    In the internal __cluster_metadata log, replicated across the controller quorum via the Raft protocol.

saying these in an interview costs you the question

  • Saying migration requires full cluster downtime — it is an online rolling process.
  • Claiming KRaft still uses ZooKeeper under the hood — it removes ZooKeeper entirely.
  • Confusing KRaft with a ZooKeeper plugin; it is a built-in Raft-based controller quorum.

context