Why does a Kafka cluster's ZooKeeper ensemble need to be secured, and what is the risk of leaving it open to unauthenticated access?
answer
- ZK = metadata store: ACLs, SCRAM, topics, brokers
- No auth by default on 2181
- Bypasses Kafka authz entirely
- zookeeper.set.acl=true + SASL + TLS
- KRaft removes ZK attack surface
basics
~10 sIn ZooKeeper-based Kafka, ZooKeeper stores critical metadata: topic configs, ACLs, SCRAM credentials, and broker registrations. An unauthenticated client can read or delete this metadata directly, bypassing Kafka entirely, so ZooKeeper must require authentication.
solid answer
~40 sClassic (pre-KRaft) Kafka keeps its cluster metadata in ZooKeeper: topic definitions and configs, partition assignments, controller election, broker registration znodes, ACLs, and SCRAM credential hashes. ZooKeeper has no auth by default, so anyone who can reach port 2181 can read or mutate that data directly, bypassing every Kafka-level authorization check. An attacker could delete topic znodes, alter ACLs to grant themselves access, or read SCRAM hashes for offline cracking. Hardening means enabling SASL authentication (typically Kerberos or DIGEST-MD5) between brokers and ZooKeeper, setting zookeeper.set.acl=true so Kafka writes its znodes with restrictive ACLs, and optionally TLS (zookeeper.ssl.client.enable=true) to encrypt the channel. Network isolation of ZooKeeper is also expected. KRaft removes ZooKeeper entirely, which is one of its security motivations.
go deeper
Know that ZooKeeper stores Kafka metadata (ACLs, SCRAM, topics) and is open by default, so it must be authenticated/isolated.
Be able to name the hardening knobs: SASL JAAS, zookeeper.set.acl=true, zookeeper.ssl.client.enable, and network isolation.
Explain the bypass risk in detail, the migration-tool gap for existing znodes, and how SCRAM-hash exposure enables offline cracking.
Frame ZooKeeper as a trust root whose compromise defeats all Kafka-side security, and position KRaft migration as removing that attack surface.
## Background **Apache Kafka** is a distributed event-streaming platform. Before version 3.x's KRaft mode, every Kafka cluster depended on **Apache ZooKeeper** — a separate distributed coordination service — to store the cluster's *metadata* (the data that describes the cluster itself, as opposed to the messages flowing through it). ## What ZooKeeper holds for Kafka ZooKeeper stores data in a tree of nodes called **znodes**. For Kafka this includes: - **Broker registrations** — which broker servers are alive (`/brokers/ids`). - **Topic and partition metadata** — topic configs, partition-to-broker assignments (`/brokers/topics`). - **Controller election** — which broker is the active controller. - **ACLs** — Kafka's authorization rules (who may produce/consume which topics) when using the ZooKeeper-based `AclAuthorizer`. - **SCRAM credentials** — salted password hashes for SCRAM SASL users (`/config/users`). ## The risk of leaving ZooKeeper open ZooKeeper ships with **no authentication by default**. Its client port (2181, or 2182 for TLS) accepts any connection. If that port is reachable by an attacker, they can use the `zookeeper-shell` or any ZK client to **read or write znodes directly**, completely bypassing Kafka's broker-side security. Concretely an attacker could: - **Read SCRAM credential hashes** and crack them offline. - **Modify or delete ACL znodes**, granting themselves access or denying others. - **Delete topic znodes**, causing data loss / cluster disruption. - **Forge broker registrations** or interfere with controller election. Because Kafka *trusts* whatever is in ZooKeeper, compromising ZooKeeper compromises the whole cluster, regardless of how well the Kafka listeners themselves are secured. ## Hardening measures 1. **SASL authentication broker↔ZooKeeper.** Configure a JAAS `Client` section so brokers authenticate to ZooKeeper (commonly Kerberos/GSSAPI or DIGEST-MD5). ZooKeeper then knows the broker's identity. 2. **`zookeeper.set.acl=true`** in the broker config. This tells Kafka to create its znodes with ZooKeeper ACLs that restrict write access to the authenticated Kafka principal, so other clients cannot mutate them. 3. **TLS**: `zookeeper.ssl.client.enable=true` plus keystore/truststore settings encrypt the broker↔ZooKeeper channel (ZK 3.5+), protecting credentials and metadata in transit. 4. **Network isolation** — put ZooKeeper on a private network segment, firewalled away from clients. ## Edge cases - Setting `zookeeper.set.acl=true` on an *existing* unsecured cluster does not retroactively secure already-written znodes; you run the `zookeeper-security-migration` tool to apply ACLs to existing znodes. - Even with ZK ACLs, the data is world-*readable* by default unless you also lock down read access — and SCRAM hashes being readable is itself a risk, which is why TLS + network isolation matter. ## KRaft relevance KRaft mode (KIP-500) eliminates ZooKeeper, folding metadata into an internal Kafka log managed by a controller quorum. This removes the ZooKeeper attack surface entirely — a major security driver for the migration — but introduces the need to secure the *controller quorum listener* instead.
- What does zookeeper.set.acl=true actually do, and does it secure pre-existing znodes?It makes Kafka write new znodes with ZooKeeper ACLs restricting write access to the authenticated Kafka principal. It does NOT retroactively secure znodes already written when the cluster was unsecured — you run the zookeeper-security-migration tool to apply ACLs to existing nodes.
- Why are SCRAM credentials in ZooKeeper especially sensitive?They are salted password hashes for SASL/SCRAM users. If readable, an attacker can crack them offline to recover client credentials, then authenticate to Kafka as a legitimate user. This is why ZK read access and channel encryption matter, not just write ACLs.
saying these in an interview costs you the question
- Saying ZooKeeper authenticates clients by default — it does not; it is open unless SASL/TLS is configured.
- Claiming Kafka's own ACLs protect ZooKeeper data — ZK access bypasses Kafka entirely.
- Believing zookeeper.set.acl=true secures already-existing znodes without the migration tool.