skip to content

ZooKeeper's Historical Role

What ZooKeeper actually did for Kafka: znodes, watches, controller election, and session-based liveness. Still asked because plenty of clusters remain on ZK and the design explains why KRaft exists.

part ofApache Kafkaoverview, primer and where to startread it →
on this pageshow

questions

5

Before KRaft, what did Kafka use ZooKeeper for, and what kinds of cluster metadata lived in it?

level: juniorimportance: must knowfreq 70%

answer

  1. ZK = metadata, not message data
  2. broker registry / topics / ISR / controller / ACLs
  3. ephemeral znodes for liveness
  4. controller watches ZK
  5. separate cluster, second failure domain

basics

~20 s

ZooKeeper was an external coordination service that stored Kafka's cluster metadata: which brokers are alive, the list of topics and partitions, partition leaders and ISR, controller election, and ACLs/configs. Kafka depended on it to run.

solid answer

~40 s

ZooKeeper (ZK) was a separate distributed coordination service Kafka required for cluster-wide metadata and consensus. It stored the broker registry (which brokers are alive), topic and partition definitions, partition leadership and the in-sync replica (ISR) set, dynamic configs, ACLs, and delegation tokens. It also handled controller election: the broker that won became the controller and managed leader assignment and failover. Brokers registered ephemeral znodes for liveness, and the controller watched ZK for membership changes. ZK was the durable source of truth for metadata, while the actual record data lived in Kafka log segments on the brokers. KIP-500 later replaced ZK with KRaft because the two-system architecture added operational complexity and a metadata-scaling bottleneck.

go deeper

for a junior

Know that ZK was an external service holding cluster metadata, and that records were never stored in it.

for a middle

Be able to enumerate concrete categories: broker registry, topic/partition assignment, leader+ISR, controller, configs/ACLs.

for a senior

Explain the control-plane vs data-plane split and how ZK liveness/watches drove leader failover.

for a principal

Frame ZK as a second failure domain and metadata bottleneck that motivated KIP-500, and reason about operational tradeoffs of a two-system design.

## What ZooKeeper is Apache ZooKeeper is a standalone distributed coordination service. It exposes a small filesystem-like tree of nodes called **znodes**, each holding a little data, and it provides strong consistency (linearizable writes) via its own consensus protocol (ZAB). Many distributed systems used it as a 'source of truth' for small but critical shared state. ## Why Kafka needed it A Kafka cluster is many independent broker processes. They must agree on facts that no single broker owns: *which brokers are alive, what topics exist, who leads each partition, which replicas are caught up*. Kafka (pre-2.8, fully removed by default in 4.0) delegated this agreement to ZooKeeper. Kafka stored **metadata** in ZK; the actual message data stayed in Kafka's own log segment files on the brokers' disks. ZK never held your records — only the control-plane state. ## What lived in ZooKeeper - **Broker registry** — `/brokers/ids/<id>`: an ephemeral znode per live broker with its host/port. Disappears when the broker's ZK session ends, giving liveness. - **Topics & partitions** — `/brokers/topics/<topic>`: replica assignment per partition. - **Partition state** — leader and the **ISR** (in-sync replica set) per partition, used during leader election/failover. - **Controller** — `/controller`: an ephemeral znode marking which broker is the elected controller. - **Configs / ACLs / quotas / delegation tokens** — under `/config`, `/kafka-acl`, etc. - **Consumer offsets (very old clients)** — older consumers stored offsets in ZK before the `__consumer_offsets` topic became the default. ## How it was used at runtime The **controller** broker watched ZK znodes (via **watches** — one-shot change notifications). When a broker died, its ephemeral znode vanished, ZK fired a watch, and the controller recomputed leaders for affected partitions and pushed updates to brokers. ## Edge cases / gotchas - ZK was a *separate cluster* to run, secure, and tune — a second failure domain. - Metadata changes funneled through ZK, which became a scaling bottleneck at high partition counts (slow controller failover, slow startup). - A ZK outage froze metadata changes (no new leaders elected) even if brokers were healthy. KIP-500 set the direction to remove ZK; KRaft (the replacement) folds metadata + consensus into Kafka itself.

  • Did Kafka store the actual messages/records in ZooKeeper?
    No. Records live in Kafka's own log segment files on broker disks. ZooKeeper only held control-plane metadata (broker registry, partition leadership/ISR, configs, ACLs, controller election).
  • Where did consumer offsets live — ZK or Kafka?
    Very old consumers committed offsets to ZK, but modern Kafka stores them in the internal compacted topic __consumer_offsets. Offsets-in-ZK was deprecated well before KIP-500.

saying these in an interview costs you the question

  • Saying Kafka stored message data in ZooKeeper (it never did — only metadata).
  • Claiming ZK was just for the controller; it held broad metadata (topics, ISR, ACLs, configs).
  • Confusing ZK with KRaft — ZK is the legacy external system, KRaft is the replacement.

context

open as a page

Why was ZooKeeper deprecated and removed from Kafka (KIP-500)? What concrete problems did the ZK dependency cause?

level: seniorimportance: must knowfreq 75%

basics

~20 s

Running ZooKeeper meant operating a second distributed system alongside Kafka — extra ops, a separate failure domain, and a metadata-scaling bottleneck (slow controller failover and startup at high partition counts). KIP-500 replaced it so Kafka manages its own metadata, simplifying operations and improving scalability.

open as a page

What is an ephemeral znode, and how did Kafka use ephemeral znodes and ZooKeeper sessions to detect a dead broker?

level: middleimportance: should knowfreq 55%

basics

~20 s

An ephemeral znode is a ZooKeeper node that exists only while the client's session is alive; it auto-deletes when the session ends. Each broker created one under /brokers/ids. When the broker's session timed out, the znode vanished, signaling the broker was dead.

open as a page

Walk through how the Kafka controller was elected via ZooKeeper, and what happened on controller failover.

level: seniorimportance: should knowfreq 50%

basics

~20 s

Brokers raced to create a single ephemeral znode at /controller; whoever created it first became the controller. If that broker died, its session expired, /controller was deleted, every other broker got a watch notification, and they raced again to elect a new controller.

open as a page

What were the key ZooKeeper znode paths Kafka used (e.g. /brokers, /controller, /admin), and how could you inspect them?

level: middleimportance: nice to knowfreq 35%

basics

~10 s

Kafka kept its metadata under well-known ZK paths: /brokers/ids (live brokers), /brokers/topics (topic/partition assignment), /controller (current controller), /controller_epoch, /admin (admin operations like deletes/reassignments), /config, and /kafka-acl. You inspected them with zookeeper-shell.sh.

open as a page