How are sensitive metadata items like SCRAM credentials and ACLs protected from unauthenticated access in both ZooKeeper-mode and KRaft, and what's the bootstrap challenge for SCRAM in KRaft?
answer
- SCRAM = salted hash, still crackable offline → hide it
- ZK: znodes + set.acl + SASL + TLS + isolation
- KRaft: records in __cluster_metadata → secure controller listener
- KIP-900: kafka-storage format --add-scram seeds first creds
- Controller disk holds secrets — protect at rest
basics
~20 sIn ZK-mode, SCRAM hashes and ACLs live in znodes protected by ZooKeeper ACLs (zookeeper.set.acl) plus SASL/TLS. In KRaft they live in the __cluster_metadata log, protected by securing the controller listener. KRaft's bootstrap challenge: seed the first SCRAM credential with kafka-storage --add-scram.
solid answer
~40 sSCRAM credentials are salted/iterated password hashes, and ACLs are authorization rules — both are sensitive metadata. In ZooKeeper mode they sit in znodes (/config/users, ACL nodes); protection means zookeeper.set.acl=true (restrict write to the broker principal), SASL auth so ZK knows who's connecting, TLS to stop sniffing of the hashes, and network isolation since even readable hashes enable offline cracking. In KRaft, these records live in the replicated __cluster_metadata log; you protect them by securing the controller listener (mTLS/SASL_SSL) and isolating it — there are no separate znode ACLs. KRaft's bootstrap problem: SCRAM credentials needed to authenticate clients are stored in the metadata log, but you can't write to it before the cluster runs. KIP-900 solves this by letting kafka-storage.sh format --add-scram seed initial SCRAM credentials at format time.
go deeper
Know SCRAM creds and ACLs are sensitive metadata that must be kept from unauthenticated readers.
Map where they live (znodes vs __cluster_metadata) and the basic protections for each.
Explain offline-cracking risk of hashes, the read-vs-write ACL nuance, and the KIP-900 SCRAM bootstrap.
Treat controller hosts/ZK ensembles as secret stores with at-rest + in-transit + authz controls and design credential rotation.
## What counts as sensitive metadata - **SCRAM credentials** — for SASL/SCRAM auth, Kafka stores not the plaintext password but a *salted, iterated hash* (server key + stored key). Still, a leaked hash can be **brute-forced/dictionary-cracked offline**, so it must not be readable by attackers. - **ACLs** — rules of the form "principal P is allowed/denied operation O on resource R." If an attacker can edit them, they grant themselves access; if they can read them, they map the security model. - **Delegation tokens, feature flags, topic configs** — also sensitive to tampering. ## ZooKeeper-mode protection These items live in znodes: - SCRAM under `/config/users/<user>`. - ACLs under the authorizer's znode subtree. Protection layers: 1. **ZooKeeper ACLs** via `zookeeper.set.acl=true` — write access limited to the authenticated Kafka principal. (Read access may still be broad unless tightened, which is why the next layers matter.) 2. **SASL authentication** broker↔ZK so identities are known. 3. **TLS** (`zookeeper.ssl.client.enable`) so hashes aren't sniffable in transit. 4. **Network isolation** of the ensemble. The gap: turning these on doesn't retrofit existing znodes — run `zookeeper-security-migration.sh`. ## KRaft protection There is no ZooKeeper and no per-record ACL system inside the metadata store. Instead **all** of this lives as records in the `__cluster_metadata` Raft log. Protection therefore reduces to: 1. **Securing the controller listener** (mTLS/SASL_SSL) so only authenticated controllers/brokers can read/replicate the log. 2. **Authorizing** the principals (StandardAuthorizer, `super.users`). 3. **Network isolation** of the controller endpoints. 4. **At-rest** concerns: the metadata log files on controller disks contain these records — protect the disks/filesystem like any secret store. ## The SCRAM bootstrap problem in KRaft Classic chicken-and-egg: to authenticate the admin who would *create* SCRAM users, you may need a SCRAM credential — but credentials live in the metadata log that only exists once the cluster is up, and you may want SASL/SCRAM required from the very first connection. **KIP-900** addresses this: `kafka-storage.sh format` gains `--add-scram 'SCRAM-SHA-512=[name=admin,password=...]'`, writing initial SCRAM credentials directly into the metadata log *at format time*, before the cluster starts. This lets you stand up a cluster that requires SCRAM auth immediately, without a window of weak/no auth. ## Edge cases - In ZK-mode, even with write-ACLs, world-readable znodes leak SCRAM hashes — don't rely on write-ACLs alone. - In KRaft, anyone with filesystem access to a controller's metadata log can read SCRAM/ACL records — treat controller hosts as high-value. - Rotating SCRAM credentials: update via kafka-configs (writes new records); old hashes are superseded but ensure log compaction/retention doesn't leave stale secrets unexpectedly.
- If ZooKeeper znodes are world-readable but write-protected by set.acl, is the SCRAM data safe?No. Read access still leaks the salted SCRAM hashes, which can be cracked offline to recover passwords. You need read restrictions, TLS in transit, and network isolation, not just write ACLs.
- Where do SCRAM and ACL records physically live in KRaft, and what's the at-rest implication?They are records in the __cluster_metadata log, persisted on controller node disks. Anyone with filesystem access to a controller can read them, so controller hosts must be treated as high-value secret stores with disk/OS hardening.
saying these in an interview costs you the question
- Saying SCRAM stores plaintext passwords — it stores salted, iterated hashes.
- Believing write-only ACLs make readable SCRAM hashes safe.
- Assuming KRaft has ZooKeeper-style per-node ACLs for metadata — it doesn't; you secure the listener and disk instead.
- Not knowing the --add-scram format-time bootstrap (KIP-900).