skip to content

Explain joint consensus and how it makes single-voter membership changes safe in KRaft.

level: principalimportance: should knowfreq 20%

answer

  1. majority of BOTH C_old and C_new
  2. prevents two disjoint majorities / two leaders
  3. single-server change = guaranteed overlap
  4. two-phase: C_old,new then C_new
  5. why 3→5 is two adds, not one

basics

~20 s

Joint consensus is a transitional Raft state where decisions need a majority of BOTH the old and the new voter sets at once. This overlap guarantees no two leaders can be elected from disjoint majorities during a membership change, keeping the cluster safe.

solid answer

~50 s

When KRaft changes its voter set, it cannot atomically switch every replica from the old configuration C-old to the new C-new — replicas adopt the change at different times, so for a window two configurations coexist. Joint consensus handles this: the leader appends a configuration entry that puts the cluster into a joint state C-old,new where every election and every commit must achieve a majority in BOTH C-old and C-new independently. Because the two majorities must each be satisfied, no candidate can win using only C-old while another wins using only C-new — the configurations can't diverge into two leaders. Once the joint entry commits, the leader appends a second entry for C-new alone, and after that commits the transition is done. KRaft restricts changes to one voter at a time, which is the simpler single-server special case where old and new majorities always overlap by construction, still preserving safety.

go deeper

for a junior

Recall that you can only change one controller at a time and that this is for safety.

for a middle

Explain that joint consensus needs a majority of both old and new sets so two leaders can't be elected.

for a senior

Give the disjoint-majority failure example and why single-server changes guarantee overlap.

for a principal

Reason about append-vs-commit timing, leader-crash mid-transition, and the two-phase C_old,new → C_new structure.

## The problem membership change creates Raft safety rests on **majority overlap**: any two majorities of the same voter set share at least one member, so a newly elected leader always sees the latest committed entry. But when the **voter set changes**, replicas switch from the old set (`C_old`) to the new set (`C_new`) at slightly different times. During that window, two *different* configurations are active across the cluster. If elections could complete using only `C_old` on some nodes and only `C_new` on others, you could elect **two leaders simultaneously** from **disjoint** majorities — a split brain that violates safety. ## Joint consensus (the general mechanism) The original Raft paper solves this with **joint consensus**. The leader appends a special configuration log entry that puts the cluster into a **joint configuration `C_old,new`**. While in this joint state: - **Log entries are committed** only when replicated to a majority of `C_old` **AND** a majority of `C_new`. - **Leader elections succeed** only with votes from a majority of `C_old` **AND** a majority of `C_new`. Because every decision requires *both* majorities, the old and new worlds cannot independently elect leaders or commit conflicting entries — the dual requirement forces overlap across the configurations. Once the `C_old,new` entry is committed, the leader appends a second entry for `C_new` alone; when that commits, `C_old` is retired and the system runs on the new set. ## KRaft's choice: one server at a time KRaft (KIP-853) deliberately limits reconfiguration to **adding or removing a single voter per change**. The single-server change is the well-known special case where `C_old` and `C_new` differ by exactly one member, so a majority of one **always intersects** a majority of the other — overlap is guaranteed without a full two-phase joint entry in the general sense. KRaft still models the transition as a controlled configuration change committed through the log (`VotersRecord`), and conceptually relies on the same joint-overlap safety property. The practical upshot for operators: **you may only change membership by one voter at a time**, which is why a 3→5 grow is two separate add-controller operations. ## Worked example: why two-at-once is unsafe Suppose `C_old = {1,2,3}` (majority 2) and you tried to jump directly to `C_new = {3,4,5}` (majority 2). Nodes {1,2} could form a majority of `C_old` and elect leader A; nodes {4,5} could form a majority of `C_new` and elect leader B — disjoint, two leaders. Single-server steps (`{1,2,3}` → `{1,2,3,4}` → `{1,2,3,4,5}` → ...) make every adjacent pair of configurations overlap, eliminating that gap. ## Edge cases and operational notes - A configuration change entry takes effect **as soon as a replica appends it**, not when it commits — so even an uncommitted change must be handled safely (KRaft accounts for this). - If a membership change is in flight and the leader crashes, the new leader continues or aborts the transition based on what is in its log; the single-server constraint keeps this safe. - Removing the current leader from the voter set is handled by stepping down after the change commits. - This is purely about the **controller voter set**, not broker membership, which is tracked separately as registrations. ## Takeaway Joint consensus / single-server changes guarantee that during any membership transition there is never a moment when two disjoint majorities can act independently — that overlap property is the whole point, and it is why KRaft forces one-voter-at-a-time reconfiguration.

  • Why does KRaft restrict reconfiguration to one voter at a time?
    A single-member difference guarantees majorities of the old and new sets always overlap, so no two disjoint majorities can elect separate leaders — the safe special case of joint consensus, without needing arbitrary multi-member jumps.
  • Give a concrete unsafe scenario if two voters were swapped atomically.
    C_old={1,2,3}, C_new={3,4,5}: {1,2} form a C_old majority and elect leader A while {4,5} form a C_new majority and elect leader B — two leaders from disjoint majorities, split brain.
  • When does a configuration change entry start affecting decisions — on append or on commit?
    On append. A replica must treat the new configuration as active as soon as it adds the entry to its log, even before it commits, which the protocol accounts for.

saying these in an interview costs you the question

  • Saying membership changes are atomic across all nodes — they propagate asynchronously, which is the whole reason joint consensus exists.
  • Claiming you can grow 3→5 in a single reconfiguration step.
  • Asserting a config change only matters once committed — it affects elections from the moment it is appended.
  • Confusing controller voter membership with broker registration.

context