skip to content

ZooKeeper and KRaft Metadata Hardening

Securing the metadata plane itself, whether that is ZooKeeper ACLs and TLS or the KRaft controller listener. Interviewers ask because an open metadata store hands over the entire cluster.

part ofApache Kafkaoverview, primer and where to startread it →
on this pageshow

questions

5

Why does a Kafka cluster's ZooKeeper ensemble need to be secured, and what is the risk of leaving it open to unauthenticated access?

level: juniorimportance: must knowfreq 70%

answer

  1. ZK = metadata store: ACLs, SCRAM, topics, brokers
  2. No auth by default on 2181
  3. Bypasses Kafka authz entirely
  4. zookeeper.set.acl=true + SASL + TLS
  5. KRaft removes ZK attack surface

basics

~10 s

In ZooKeeper-based Kafka, ZooKeeper stores critical metadata: topic configs, ACLs, SCRAM credentials, and broker registrations. An unauthenticated client can read or delete this metadata directly, bypassing Kafka entirely, so ZooKeeper must require authentication.

solid answer

~40 s

Classic (pre-KRaft) Kafka keeps its cluster metadata in ZooKeeper: topic definitions and configs, partition assignments, controller election, broker registration znodes, ACLs, and SCRAM credential hashes. ZooKeeper has no auth by default, so anyone who can reach port 2181 can read or mutate that data directly, bypassing every Kafka-level authorization check. An attacker could delete topic znodes, alter ACLs to grant themselves access, or read SCRAM hashes for offline cracking. Hardening means enabling SASL authentication (typically Kerberos or DIGEST-MD5) between brokers and ZooKeeper, setting zookeeper.set.acl=true so Kafka writes its znodes with restrictive ACLs, and optionally TLS (zookeeper.ssl.client.enable=true) to encrypt the channel. Network isolation of ZooKeeper is also expected. KRaft removes ZooKeeper entirely, which is one of its security motivations.

go deeper

for a junior

Know that ZooKeeper stores Kafka metadata (ACLs, SCRAM, topics) and is open by default, so it must be authenticated/isolated.

for a middle

Be able to name the hardening knobs: SASL JAAS, zookeeper.set.acl=true, zookeeper.ssl.client.enable, and network isolation.

for a senior

Explain the bypass risk in detail, the migration-tool gap for existing znodes, and how SCRAM-hash exposure enables offline cracking.

for a principal

Frame ZooKeeper as a trust root whose compromise defeats all Kafka-side security, and position KRaft migration as removing that attack surface.

## Background **Apache Kafka** is a distributed event-streaming platform. Before version 3.x's KRaft mode, every Kafka cluster depended on **Apache ZooKeeper** — a separate distributed coordination service — to store the cluster's *metadata* (the data that describes the cluster itself, as opposed to the messages flowing through it). ## What ZooKeeper holds for Kafka ZooKeeper stores data in a tree of nodes called **znodes**. For Kafka this includes: - **Broker registrations** — which broker servers are alive (`/brokers/ids`). - **Topic and partition metadata** — topic configs, partition-to-broker assignments (`/brokers/topics`). - **Controller election** — which broker is the active controller. - **ACLs** — Kafka's authorization rules (who may produce/consume which topics) when using the ZooKeeper-based `AclAuthorizer`. - **SCRAM credentials** — salted password hashes for SCRAM SASL users (`/config/users`). ## The risk of leaving ZooKeeper open ZooKeeper ships with **no authentication by default**. Its client port (2181, or 2182 for TLS) accepts any connection. If that port is reachable by an attacker, they can use the `zookeeper-shell` or any ZK client to **read or write znodes directly**, completely bypassing Kafka's broker-side security. Concretely an attacker could: - **Read SCRAM credential hashes** and crack them offline. - **Modify or delete ACL znodes**, granting themselves access or denying others. - **Delete topic znodes**, causing data loss / cluster disruption. - **Forge broker registrations** or interfere with controller election. Because Kafka *trusts* whatever is in ZooKeeper, compromising ZooKeeper compromises the whole cluster, regardless of how well the Kafka listeners themselves are secured. ## Hardening measures 1. **SASL authentication broker↔ZooKeeper.** Configure a JAAS `Client` section so brokers authenticate to ZooKeeper (commonly Kerberos/GSSAPI or DIGEST-MD5). ZooKeeper then knows the broker's identity. 2. **`zookeeper.set.acl=true`** in the broker config. This tells Kafka to create its znodes with ZooKeeper ACLs that restrict write access to the authenticated Kafka principal, so other clients cannot mutate them. 3. **TLS**: `zookeeper.ssl.client.enable=true` plus keystore/truststore settings encrypt the broker↔ZooKeeper channel (ZK 3.5+), protecting credentials and metadata in transit. 4. **Network isolation** — put ZooKeeper on a private network segment, firewalled away from clients. ## Edge cases - Setting `zookeeper.set.acl=true` on an *existing* unsecured cluster does not retroactively secure already-written znodes; you run the `zookeeper-security-migration` tool to apply ACLs to existing znodes. - Even with ZK ACLs, the data is world-*readable* by default unless you also lock down read access — and SCRAM hashes being readable is itself a risk, which is why TLS + network isolation matter. ## KRaft relevance KRaft mode (KIP-500) eliminates ZooKeeper, folding metadata into an internal Kafka log managed by a controller quorum. This removes the ZooKeeper attack surface entirely — a major security driver for the migration — but introduces the need to secure the *controller quorum listener* instead.

  • What does zookeeper.set.acl=true actually do, and does it secure pre-existing znodes?
    It makes Kafka write new znodes with ZooKeeper ACLs restricting write access to the authenticated Kafka principal. It does NOT retroactively secure znodes already written when the cluster was unsecured — you run the zookeeper-security-migration tool to apply ACLs to existing nodes.
  • Why are SCRAM credentials in ZooKeeper especially sensitive?
    They are salted password hashes for SASL/SCRAM users. If readable, an attacker can crack them offline to recover client credentials, then authenticate to Kafka as a legitimate user. This is why ZK read access and channel encryption matter, not just write ACLs.

saying these in an interview costs you the question

  • Saying ZooKeeper authenticates clients by default — it does not; it is open unless SASL/TLS is configured.
  • Claiming Kafka's own ACLs protect ZooKeeper data — ZK access bypasses Kafka entirely.
  • Believing zookeeper.set.acl=true secures already-existing znodes without the migration tool.

context

open as a page

In KRaft mode, how do you secure the controller quorum, and why is the controller listener security distinct from the broker's client listeners?

level: seniorimportance: must knowfreq 45%

basics

~20 s

KRaft controllers form a Raft quorum that replicates the metadata log. You secure their listener (named in controller.listener.names) with its own SSL/SASL entry in listener.security.protocol.map, separate from client listeners, because that channel carries all cluster metadata.

open as a page

How do you configure SASL and TLS between Kafka brokers and ZooKeeper, and what do zookeeper.set.acl and zookeeper.ssl.client.enable control?

level: middleimportance: should knowfreq 50%

basics

~10 s

SASL is set up via a JAAS Client section so brokers authenticate to ZooKeeper. zookeeper.set.acl=true makes Kafka write znodes with restrictive ACLs. zookeeper.ssl.client.enable=true plus keystore/truststore turns on TLS encryption to ZooKeeper.

open as a page

How are sensitive metadata items like SCRAM credentials and ACLs protected from unauthenticated access in both ZooKeeper-mode and KRaft, and what's the bootstrap challenge for SCRAM in KRaft?

level: seniorimportance: should knowfreq 35%

basics

~20 s

In ZK-mode, SCRAM hashes and ACLs live in znodes protected by ZooKeeper ACLs (zookeeper.set.acl) plus SASL/TLS. In KRaft they live in the __cluster_metadata log, protected by securing the controller listener. KRaft's bootstrap challenge: seed the first SCRAM credential with kafka-storage --add-scram.

open as a page

What are the security implications of migrating a live cluster from ZooKeeper to KRaft, and how do you keep the dual-write/bridge period secure?

level: principalimportance: should knowfreq 25%

basics

~20 s

During ZK→KRaft migration (KIP-866) both systems run together and metadata is dual-written, so you must keep ZooKeeper hardened (SASL/TLS/ACLs) AND secure the new controller listener at the same time. The migration controllers still connect to ZK, so neither attack surface can be relaxed until ZK is removed.

open as a page