How does metadata propagate from the active controller to brokers in KRaft, and how does this differ from the ZooKeeper-era controller model?
answer
- KRaft = brokers PULL the metadata log (observers)
- incremental deltas applied in offset order
- broker reports metadata offset → lag tracked, fencing
- ZK era = controller PUSHED LeaderAndIsr/UpdateMetadata RPCs
- ZK failover reloaded full state (slow); KRaft in-memory (fast)
basics
~20 sIn KRaft, brokers pull metadata changes by fetching the __cluster_metadata log from the active controller and applying records incrementally to a local metadata cache. In the ZooKeeper era, the controller pushed full LeaderAndIsr/UpdateMetadata RPCs to brokers, which was slower and harder to scale.
solid answer
~50 sKRaft replaces a push model with a pull, log-based model. Brokers are observers of the __cluster_metadata Raft log: they continuously fetch new records from the active controller and apply them in offset order to build an in-memory metadata image, advancing a per-broker metadata offset. Updates are incremental deltas (one record per change), and a broker reports the offset it has caught up to, so the controller knows each broker's metadata lag. On restart a broker loads the latest snapshot then the tail. The ZooKeeper-era controller instead watched ZooKeeper and pushed metadata to brokers via LeaderAndIsr and UpdateMetadata RPCs; on controller failover it reloaded all metadata from ZooKeeper (seconds in big clusters) and resent full state, which capped partition counts and slowed recovery. KRaft's deltas-over-a-shared-log design gives faster, near-instant failover and far higher partition scalability, and makes metadata a single ordered source of truth rather than fan-out RPCs.
go deeper
Know brokers in KRaft get metadata by reading the controller's log rather than ZooKeeper.
Explain the pull/observer model, incremental deltas, and that it replaced ZooKeeper-watch + push RPCs.
Detail offset tracking, fencing, heartbeats, snapshots, and the specific old RPCs (LeaderAndIsr/UpdateMetadata) replaced.
Reason about why deltas-over-a-log enables million-partition scale and near-instant failover, and the trade-offs vs fan-out RPC propagation.
## The KRaft propagation path (pull, log-based) 1. The **active controller** appends each metadata change as a **record** to the single-partition `__cluster_metadata` Raft log and commits it once a majority of voters acknowledge. 2. **Brokers are observers** (non-voting). Each broker runs a fetch loop that **pulls** committed records from the active controller, exactly as a follower fetches from a partition leader. 3. The broker **applies records in offset order** to its in-memory **metadata image** (the "MetadataDelta/MetadataImage" structures) — learning, for example, that it is now leader for partition X or that topic Y was created. 4. The broker tracks the **metadata offset** it has applied and reports it back (via heartbeats), so the controller knows how far behind each broker is (its **metadata lag**). A broker that falls too far behind can be **fenced** until it catches up. 5. On startup or reconnection, a broker loads the **latest snapshot** and then replays only the **tail** after that offset — bounded, fast catch-up. Key property: **incremental deltas**. Each change is one small record. Brokers never need a full metadata dump for routine changes; they just consume the next records. ## Broker registration and heartbeats Brokers **register** with the controller (a record in the log) and send periodic **heartbeats** (the `BrokerHeartbeat` RPC). The controller uses these to track liveness and metadata catch-up, replacing ZooKeeper ephemeral znodes for membership. A broker that stops heartbeating is **fenced** (excluded from leadership) and eventually unregistered. ## The ZooKeeper-era model (push, RPC-based) — the contrast - Membership/metadata lived in **ZooKeeper**. One broker was the **controller**, elected via a ZooKeeper ephemeral node. - The controller **watched** ZooKeeper for changes and **pushed** updates to brokers using **`LeaderAndIsr`** RPCs (telling a broker which partitions it leads/follows) and **`UpdateMetadata`** RPCs (cluster-wide topic/partition/leader maps). - On **controller failover**, the new controller had to **read all metadata from ZooKeeper** and **resend full state** to every broker — an O(metadata-size) operation taking many seconds in large clusters, during which the cluster was effectively unmanaged. - This fan-out-of-full-state design **limited partition counts** (practically the low hundreds of thousands) and made recovery slow. ## Why KRaft is better - **Near-instant failover**: every controller voter already has the full metadata log in memory; the new active controller is immediately up to date — no ZooKeeper reload. - **Incremental propagation**: brokers consume deltas, not full snapshots, so per-change cost is tiny and **millions of partitions** become feasible. - **Single ordered source of truth**: an offset-ordered log instead of fan-out RPCs that could arrive out of order or get lost. - **Observable lag**: per-broker metadata offset makes "is everyone caught up?" directly measurable. ## Edge cases - A broker with stale metadata (high lag) may briefly act on old leadership info; the offset tracking + fencing bound this window. - Snapshots are still needed: a brand-new or very-stale broker gets a snapshot rather than replaying from offset 0. - The data-plane RPCs that clients use (Produce/Fetch) are unchanged; only the *metadata* propagation mechanism changed. - `LeaderAndIsr`/`UpdateMetadata` RPCs are removed in KRaft — a frequent gotcha for people debugging based on old runbooks.
- Why is controller failover so much faster in KRaft than with ZooKeeper?Every controller voter already holds the complete metadata log in memory, so the newly elected active controller is instantly current. The ZooKeeper-era controller had to read all metadata from ZooKeeper and resend full state to brokers, which took seconds in large clusters.
- How does the controller know whether a broker has the latest metadata?Brokers report the metadata log offset they have applied (via heartbeats). The controller compares it to the committed high-watermark to compute each broker's metadata lag, and can fence brokers that fall too far behind.
- Which RPCs did the ZooKeeper-era controller use to push metadata, and what replaced them?LeaderAndIsr and UpdateMetadata (plus StopReplica) RPCs pushed state to brokers. KRaft removes them; brokers instead fetch the __cluster_metadata log and apply incremental records themselves.
saying these in an interview costs you the question
- Saying the active controller pushes full metadata to brokers in KRaft (brokers pull deltas)
- Claiming LeaderAndIsr/UpdateMetadata RPCs still drive metadata in KRaft (removed)
- Thinking failover requires reloading all metadata in KRaft (voters already hold it in memory)
- Confusing the metadata fetch path with client Produce/Fetch (those are unchanged)