skip to content

Why was ZooKeeper deprecated and removed from Kafka (KIP-500)? What concrete problems did the ZK dependency cause?

level: seniorimportance: must knowfreq 75%

answer

  1. second system / second failure domain
  2. controller failover = O(partitions) ZK reload
  3. partition-count ceiling
  4. two security/ACL models
  5. 3.5 deprecated, 4.0 removed by default

basics

~20 s

Running ZooKeeper meant operating a second distributed system alongside Kafka — extra ops, a separate failure domain, and a metadata-scaling bottleneck (slow controller failover and startup at high partition counts). KIP-500 replaced it so Kafka manages its own metadata, simplifying operations and improving scalability.

solid answer

~50 s

KIP-500 set out to remove ZooKeeper for four main reasons. (1) **Operational complexity**: ZK was a separate cluster to deploy, secure, monitor, upgrade, and back up — a second system with its own quirks and a second failure domain. (2) **Metadata scalability**: all metadata changes funneled through ZK, and controller failover required reloading all metadata from ZK, so recovery time and startup grew O(partitions). This capped practical partition counts per cluster. (3) **Consistency seams**: the controller cached ZK state in memory; the two-system split created divergence/propagation lag and complex code to keep them aligned. (4) **Security/config duplication**: ZK had its own auth and ACL model, separate from Kafka's. The replacement, **KRaft**, stores metadata as an internal Raft-replicated log managed by controller nodes, so Kafka becomes self-managed: faster failover, far higher partition limits, single security model, and one system to operate. ZK was deprecated in 3.x and removed by default in Kafka 4.0.

go deeper

for a junior

Know ZK was a separate system that added operational overhead and was replaced to simplify Kafka.

for a middle

List the main drivers: extra ops, separate failure domain, metadata-scaling limits.

for a senior

Quantify the failover/startup O(partitions) cost and the partition ceiling, plus the dual security model.

for a principal

Tie the motivations to the KRaft design choice, the deprecation/removal timeline, and the KIP-866 migration strategy and its risks.

## Background For over a decade Kafka required an external **ZooKeeper** ensemble to hold cluster metadata and run controller election. **KIP-500** ('Replace ZooKeeper with a Self-Managed Metadata Quorum') proposed eliminating that dependency. (The *replacement design* — KRaft, the metadata quorum — is covered by a sibling topic; here we focus on *why ZK had to go*.) ## The concrete pain points ### 1. Two systems to operate ZooKeeper is its own distributed system with its own deployment, configuration, JVM tuning, monitoring, upgrade cadence, and backup story. Operators had to be experts in *both* Kafka and ZK. ZK was also a **separate failure domain**: if the ZK ensemble lost quorum, Kafka couldn't elect leaders or change metadata even though the brokers themselves were fine. More moving parts = more ways to fail and more on-call burden. ### 2. Metadata scalability bottleneck Metadata changes (topic creates, leader changes, ISR updates) were serialized through the controller and ZK. Two costs dominated: - **Controller failover**: a new controller reloaded the *entire* metadata set from ZK to rebuild memory — an O(number of partitions) read storm. At hundreds of thousands of partitions this took a long time, during which leadership changes stalled. - **Broker startup / metadata propagation**: pushing full metadata via ZK and `UpdateMetadata`/`LeaderAndIsr` RPCs didn't scale incrementally. This effectively capped the number of partitions a cluster could safely run. ### 3. Consistency and code complexity The controller kept an in-memory cache of ZK state and synchronized the two. Keeping an external store and the in-memory model consistent created subtle bugs, propagation lag, and a lot of glue code. There was no single, ordered log of metadata changes — just point-in-time znode state plus watches. ### 4. Security and configuration duplication ZK had its own authentication (e.g. SASL) and ACL model, distinct from Kafka's. Securing a cluster meant securing two systems with two models, and misconfigured ZK ACLs were a real risk (anyone with ZK write access could corrupt cluster metadata). ## What replaced it (briefly) KRaft turns metadata itself into an internal, ordered, **Raft-replicated log** managed by dedicated controller nodes. Because the active controller is just the leader of that log, it already holds metadata in memory — failover becomes ~the time to elect a Raft leader instead of an O(partitions) reload. One system, one security model, incremental metadata propagation, far higher partition ceilings. ## Timeline / status - **KIP-500**: the umbrella proposal to remove ZK. - **Kafka 2.8**: early-access KRaft (not production). - **Kafka 3.3**: KRaft marked production-ready. - **Kafka 3.5**: ZooKeeper mode deprecated. - **Kafka 4.0**: ZooKeeper mode removed by default; clusters run KRaft. ## Edge cases / gotchas - 'Removing ZK' isn't *only* about speed — operational simplicity and a single security/config surface were equally motivating. - Migration from ZK to KRaft was non-trivial; KIP-866 provided a dual-write migration path so existing clusters could move without downtime. - ZK wasn't 'bad' — it was a sound choice early on; Kafka simply outgrew the two-system architecture as scale and operational expectations rose.

  • Name two operational benefits of removing ZooKeeper that are not about raw performance.
    (1) One system to deploy/secure/monitor/upgrade instead of two, eliminating a separate failure domain. (2) A single unified security and configuration model — no separate ZK auth/ACLs to manage and misconfigure.
  • Roughly when was ZK deprecated and removed?
    KRaft became production-ready in Kafka 3.3; ZooKeeper mode was deprecated in 3.5 and removed by default in Kafka 4.0. KIP-866 provided the online ZK-to-KRaft migration path.
  • Why did high partition counts specifically expose the ZK design?
    Controller failover and startup required loading all metadata from ZK, an O(partitions) operation, and metadata propagation wasn't incremental. So recovery time and resource use grew with partition count, capping how many partitions a cluster could safely run.

saying these in an interview costs you the question

  • Saying ZK was removed purely because 'it was slow' — operational simplicity and a single security model were equally central.
  • Claiming ZK stored message data and that's why it didn't scale — it only held metadata.
  • Asserting ZK was always a bad design — it was appropriate early; Kafka outgrew the two-system model.
  • Saying you can't migrate without downtime — KIP-866 enabled an online dual-write migration.

context