skip to content

How does KRaft replicate the metadata log, and why is it described as pull-based rather than push-based Raft?

level: seniorimportance: must knowfreq 55%

answer

  1. followers Fetch (pull), classic Raft pushes AppendEntries
  2. reuses Kafka replica-fetch machinery
  3. BeginQuorumEpoch/EndQuorumEpoch/Fetch/FetchSnapshot
  4. fetchOffset + lastFetchedEpoch drive commit + truncation
  5. leader learns majority position from Fetch offsets

basics

~20 s

In KRaft, followers and observers pull records from the leader using Fetch requests, just like Kafka consumers pull from a partition leader. Classic Raft instead has the leader push records via AppendEntries. KRaft reuses Kafka's existing fetch/replication machinery.

solid answer

~50 s

Classic Raft is push-based: the leader sends **AppendEntries** RPCs to followers. KRaft inverts this to be **pull-based**: each follower (voter) and observer (broker) periodically sends a **Fetch** request to the leader carrying its current `fetchOffset` and its `lastFetchedEpoch`, and the leader responds with the next batch of records. This deliberately mirrors how a normal Kafka consumer/replica fetches from a partition leader, so KRaft reuses Kafka's battle-tested log, fetch, and purgatory machinery. The leader uses the followers' reported fetch offsets to know how far each has replicated and therefore when a record reaches a majority and can be committed. The leader also runs **BeginQuorumEpoch** (to assert leadership) and processes **Vote** during elections; steady-state replication and heartbeating both flow over Fetch. A consequence: leadership liveness is detected by absence of successful Fetch exchanges rather than by leader-initiated heartbeats.

go deeper

for a junior

Know that in KRaft followers pull records via Fetch, like consumers fetch from a topic.

for a middle

Contrast pull-based Fetch with classic Raft's push AppendEntries and name BeginQuorumEpoch.

for a senior

Explain how fetchOffset/lastFetchedEpoch drive commit and truncation, plus EndQuorumEpoch and FetchSnapshot.

for a principal

Reason about reuse of Kafka's fetch machinery, liveness implications of pull, snapshotting for compaction, and observer scalability.

## Push vs pull, defined - **Push-based replication (classic Raft)**: the **leader initiates** by sending **AppendEntries** RPCs containing new log entries to each follower, and uses them as heartbeats too. - **Pull-based replication (KRaft)**: the **follower initiates** by sending a **Fetch** request asking 'give me records starting at offset X'; the leader replies with records. Nothing is replicated until a follower asks. ## Why KRaft chose pull Kafka already has a mature, high-throughput **replica fetcher** model: partition followers fetch from the partition leader. By making the metadata quorum also pull-based, KRaft reuses the same log segment format, the same fetch path, the same request **purgatory** (delayed-request handling), and the same offset/epoch bookkeeping. Less new code, more reuse of proven components. ## The core RPCs - **Vote**: sent by candidates during elections (covered separately). - **BeginQuorumEpoch**: the **leader** sends this to voters to announce 'I am leader for epoch E.' Note this is leader-initiated — it is how a freshly elected leader makes voters start fetching from it, since voters otherwise wouldn't know whom to Fetch from. - **Fetch**: the workhorse. Each follower/observer sends Fetch with `fetchOffset` and `lastFetchedEpoch`. The leader validates the epoch (divergence detection), returns records, and notes the follower's progress. - **EndQuorumEpoch**: sent by a leader that is gracefully resigning (e.g., controlled shutdown) so a new election starts promptly instead of waiting for timeouts. - **FetchSnapshot**: used when a follower is too far behind and the leader has compacted/snapshotted the metadata log; the follower fetches the snapshot, then resumes Fetch. ## How commit works under pull Because followers report their `fetchOffset` in each Fetch, the leader continuously learns each voter's replicated position. When a given offset has been fetched by a **majority of voters** (including the leader's own copy), that offset is **committed**, and the leader advances the **high watermark (HWM)**. The HWM is then propagated to followers in subsequent Fetch responses. ## Divergence / truncation The follower includes `lastFetchedEpoch` in Fetch. If the leader sees the follower's log diverges (the follower has records from an epoch the leader's history doesn't match at that offset), it tells the follower the **diverging offset/epoch**, and the follower **truncates** its log back to the last common point before continuing. This is how an old leader's uncommitted tail is cleaned up. ## Operational consequences - **Liveness detection**: a leader detects an unresponsive follower by the absence of Fetch requests; a follower detects a dead leader by Fetch failures/timeouts, which triggers a new election. - **Snapshots over infinite logs**: the metadata log is compacted into periodic snapshots so new/lagging nodes don't replay all history — they FetchSnapshot then tail with Fetch. - **Observers scale reads**: brokers as observers can fetch without affecting quorum, so adding brokers does not change commit latency. ## Pitfalls to avoid - Saying KRaft uses AppendEntries — it does not; that is classic Raft. - Thinking BeginQuorumEpoch carries the data — it only announces leadership; data flows via Fetch. - Assuming the leader pushes heartbeats — in steady state, followers drive the interaction by polling with Fetch.

  • How does a lagging KRaft follower catch up if the leader has already snapshotted and compacted old log records?
    It uses the FetchSnapshot RPC to download the latest metadata snapshot, installs it, then resumes normal Fetch from the snapshot's end offset to tail the live log.
  • What role does lastFetchedEpoch in a Fetch request play?
    It lets the leader detect log divergence: if the follower's epoch/offset history doesn't match the leader's, the leader returns a diverging offset and the follower truncates its uncommitted tail before continuing.

saying these in an interview costs you the question

  • Saying KRaft uses AppendEntries to push records
  • Claiming the leader pushes data to followers proactively
  • Confusing BeginQuorumEpoch (announce leadership) with the actual data-replication path (Fetch)
  • Ignoring snapshots/FetchSnapshot for catch-up

context