Walk through how a KRaft broker or controller materializes the current cluster state at startup using snapshots and the log.
answer
- latest snapshot offset = X
- load snapshot -> state as of X
- replay records > X in order
- MetadataDelta -> MetadataImage on broker
- snapshot + tail == replay from 0
basics
~20 sIt loads the most recent snapshot to reach that offset instantly, then replays the metadata log records after the snapshot offset in order, applying each to its in-memory state, until it reaches the end of the log.
solid answer
~50 sState materialization is deterministic replay. On startup a node finds the latest valid snapshot, whose filename encodes the offset/epoch it covers. It loads that snapshot to reconstruct the full in-memory metadata image (topics, partitions, configs, registrations, ACLs) as of that offset. It then opens the `__cluster_metadata` log and replays every record after the snapshot offset, in strict order, feeding each into the metadata image-builder (`MetadataDelta`/`MetadataImage` on brokers) so the image advances record by record. When it reaches the high-water mark / log end it has the current committed state. Controllers do the same to rebuild the `QuorumController` state machine; brokers do it as observers to build their `MetadataImage` and publish it to internal listeners. Because replay is deterministic, every node that applies the same prefix arrives at the identical state, which is the correctness guarantee behind snapshots: snapshot at X plus log after X equals full replay from 0.
go deeper
Know the two-step shape: load snapshot, then replay the rest of the log.
Describe load-latest-snapshot then ordered replay of records after that offset into the metadata image.
Explain determinism, committed-only application, and the snapshot-install path for far-behind followers.
Reason about the replicated-state-machine model, epoch/offset ordering guarantees, and recovery correctness.
## Core idea: deterministic replay KRaft metadata is a **replicated state machine**. The 'program' is the ordered sequence of metadata records in the `__cluster_metadata` log; the 'state' is the materialized image (which topics exist, who leads each partition, what the configs are, etc.). Applying the same records in the same order from the same starting point always yields the same state. This determinism is what makes snapshots safe. ## Step 1 — find the latest snapshot Snapshot files are named by the (offset, epoch) they cover, e.g. `00000000000000005000-0000000000007.checkpoint`. The node selects the highest-offset complete snapshot. If none exists (brand-new cluster) it starts from offset 0 with empty state. ## Step 2 — load the snapshot The snapshot is deserialized into the in-memory metadata structures. - On a broker this builds a `MetadataImage` (an immutable snapshot of metadata, assembled via `MetadataDelta`). - On a controller it restores the `QuorumController`'s internal timeline data structures. After this step the node's state equals 'all records up to offset X applied'. ## Step 3 — replay the log tail The node opens the metadata log and reads records with offset > X in order. Each record is applied to a `MetadataDelta`, then committed into a new `MetadataImage`. Records are things like `TopicRecord`, `PartitionRecord`, `PartitionChangeRecord`, `ConfigRecord`, `RegisterBrokerRecord`, `FeatureLevelRecord`, etc. The node only applies ***committed*** records (those below the high-water mark agreed by the Raft quorum). ## Step 4 — publish and stay current Once caught up to the log end, the broker publishes the image to its internal listeners (e.g. the components that build the metadata cache the broker serves to clients). From then on the node keeps fetching new records and applying them incrementally — replay never stops, it just transitions from catch-up to steady-state tailing. ## Why snapshot + tail == full replay By construction the snapshot is the result of replaying offsets 0..X. Replaying X+1..end on top of it is identical to replaying 0..end. So a node with a snapshot reaches the same correct state as one without, but far faster. ## Edge cases 1. A node may discover a *newer* snapshot pushed by the leader if it has fallen so far behind that the leader already truncated the log past the follower's position — then the follower fetches and installs the snapshot instead of individual records. 2. Replay must respect leader epoch / offset ordering so that a node that crashed mid-apply re-derives consistent state. 3. Only committed records are applied; uncommitted tail records can be truncated on leader change.
- What classes assemble the broker-side metadata image during replay?On the broker, records are applied to a MetadataDelta which produces an immutable MetadataImage; the broker publishes that image to internal listeners that build the client-facing metadata cache.
- What if a follower has fallen so far behind that the records it needs were already deleted?The leader serves it the latest snapshot instead of individual records; the follower installs the snapshot to jump forward, then resumes tailing the log from the snapshot offset.
saying these in an interview costs you the question
- Saying the node replays from offset 0 even when a snapshot exists
- Applying uncommitted (above high-water-mark) records during replay
- Claiming order doesn't matter — replay must be in strict offset order
- Confusing MetadataImage (metadata) with a broker's user-data log segments