Name some of the record types the active controller appends to the metadata log and explain what each represents.
answer
- TopicRecord = topic id+name
- PartitionRecord = replicas/leader/ISR/epoch
- PartitionChangeRecord = incremental delta
- RegisterBrokerRecord = id+endpoints+incarnation
- apiKey + version = evolvable
basics
~10 sMetadata changes are encoded as typed records, e.g. TopicRecord (a topic was created), PartitionRecord/PartitionChangeRecord (a partition's replicas/leader/ISR), and RegisterBrokerRecord plus BrokerRegistrationChangeRecord (a broker joined or changed state).
solid answer
~30 sThe active controller turns every cluster-state change into a typed, versioned record appended to the `__cluster_metadata` log. Key ones: `TopicRecord` (a new topic with its id and name), `PartitionRecord` (initial replica set, leader, ISR, and leader epoch for a partition), `PartitionChangeRecord` (incremental updates to leader/ISR/replicas without re-emitting the whole record), `RegisterBrokerRecord` (a broker registering with its id, endpoints, and incarnation), `BrokerRegistrationChangeRecord`/`FenceBrokerRecord`/`UnfenceBrokerRecord` (broker fencing state), `ConfigRecord` (a dynamic config change), `ProducerIdsRecord`, and ACL records. Each record has an apiKey and a version so the format can evolve. Brokers and follower controllers replay these in log order to reconstruct identical in-memory state.
go deeper
Know that metadata changes become typed records like TopicRecord and broker registration records.
List the main record types and what each captures, including incremental partition changes.
Explain versioning/apiKey for evolvability and why incremental records bound log growth.
Connect record design to schema evolution strategy, snapshotting, and long-term log-compaction/growth management.
In KRaft, metadata is *event-sourced*: rather than storing the current state directly, Kafka stores an ordered log of *change events*, and current state is computed by replaying them. Each event is a strongly-typed **record** with an `apiKey` (which record type) and a `version` (so the schema can evolve compatibly). The active controller is the only writer; it appends these to the internal single-partition topic `__cluster_metadata`. **Common record types:** - **TopicRecord** — emitted when a topic is created. Carries the topic name and a globally unique topic *id* (a UUID). Topic deletion is a `RemoveTopicRecord`. - **PartitionRecord** — describes one partition: its replica assignment (which broker ids hold replicas), the current leader, the in-sync replica set (ISR), the leader epoch, and partition epoch. Emitted when a partition is first created. - **PartitionChangeRecord** — an *incremental* update so the controller doesn't re-emit a whole PartitionRecord for every small change. It records only what changed: a new leader, a shrunk/grown ISR, a reassignment step, etc. This keeps the log compact. - **RegisterBrokerRecord** — written when a broker registers with the controller. Contains the broker id, listener endpoints, features, and an *incarnation id* (a UUID that distinguishes a fresh process start from a previous one). - **BrokerRegistrationChangeRecord / FenceBrokerRecord / UnfenceBrokerRecord** — track whether a broker is *fenced* (temporarily not eligible to host leaders, e.g. because it missed heartbeats) or active. - **ConfigRecord** — a dynamic configuration change (topic-level or broker-level). - **ProducerIdsRecord, AccessControlEntryRecord, FeatureLevelRecord, ClientQuotaRecord** — producer-id block allocation, ACLs, feature-flag levels, and quotas respectively. **Why typed records matter.** Because each is versioned, Kafka can add fields or new record types across releases while old replicas still parse what they understand. Because they are ordered and appended by a single writer, every node that replays the same log prefix derives the same state — the foundation of KRaft's consistency. **Edge cases:** records are only applied after they are *committed* (a quorum majority has them); the log is periodically *snapshotted* so a new node doesn't have to replay from the beginning of time; and `PartitionChangeRecord` is favored over re-emitting `PartitionRecord` specifically to bound log growth during churn.
- Why does Kafka use PartitionChangeRecord instead of re-emitting a full PartitionRecord?To keep the log compact: it records only the delta (leader/ISR change) rather than the entire partition state, bounding log growth during frequent leadership/ISR churn.
- How can these records evolve across Kafka versions without breaking older nodes?Each record carries an apiKey and a version; new fields are added in a backward-compatible way, so a node parses the versions it understands and tolerates additive changes.
saying these in an interview costs you the question
- Claiming the log stores current state snapshots only, not change events
- Saying brokers write these records
- Treating PartitionRecord and PartitionChangeRecord as identical
- Inventing record names that don't exist