skip to content

Walk through the JoinGroup and SyncGroup phases of a rebalance. Who computes the partition assignment, and how does it reach each member?

level: middleimportance: must knowfreq 65%

answer

  1. JoinGroup: leader gets all members, followers empty
  2. SyncGroup: leader sends assignment, coordinator fans out slices
  3. coordinator is assignment-agnostic
  4. PartitionAssignor runs on the leader (client-side)
  5. KIP-848 moves assignment server-side

basics

~20 s

All members send JoinGroup; the coordinator picks one member as group leader and returns the full member list to it. The leader computes assignments and sends them in SyncGroup; the coordinator fans them out to each member in their SyncGroup response.

solid answer

~40 s

A rebalance has two round-trips. In **JoinGroup**, every member sends its supported assignment protocols and subscriptions to the coordinator. The coordinator collects all members for one generation, elects the **group leader** (typically the first to join), and replies: the leader gets the entire member list plus everyone's metadata; followers get an empty member list. The coordinator itself does *not* compute assignment — it stays assignment-agnostic. In **SyncGroup**, the leader runs the configured `PartitionAssignor` (e.g. RangeAssignor, CooperativeStickyAssignor) to map partitions to members, then sends the full assignment map to the coordinator. Followers send empty SyncGroup requests. The coordinator stores the assignment and returns to each member only *their* slice. This client-side assignment design lets you plug in custom assignors without changing the broker.

go deeper

for a junior

Know the two phases exist and that one member (the leader) decides assignments.

for a middle

Explain leader election, assignor on the leader, and how the coordinator distributes slices.

for a senior

Discuss failure modes (leader death, protocol mismatch) and cooperative incremental rebalancing.

for a principal

Contrast the client-side leader model with KIP-848's server-side assignment and the tradeoffs.

## Why two phases A rebalance redistributes partitions across the current members of a group. Kafka splits the work between the **broker (coordinator)** — which knows *who* is in the group — and a **client (group leader)** — which decides *who gets what*. This keeps assignment logic on the client so it can be customized without broker changes. ## Phase 1 — JoinGroup 1. Each consumer sends a **JoinGroup** request to the coordinator containing its `group.id`, its current `member.id` (empty on first join), the list of **assignment protocols** it supports, and per-protocol metadata (its topic **subscription**, and for sticky assignors its currently owned partitions). 2. The coordinator waits up to `rebalance.timeout.ms` for all known members to (re)join, advancing the group to **PreparingRebalance** state. 3. The coordinator **elects the group leader** — normally the first member to join the new generation. 4. The coordinator picks a single assignment protocol that *all* members support (the protocol must be common to everyone). 5. JoinGroup responses go out: the **leader** receives the complete member list with each member's metadata; **followers** receive an empty member list. Everyone receives the new `generation.id` and their (possibly newly issued) `member.id`. ## Phase 2 — SyncGroup 1. The **leader** runs the chosen `org.apache.kafka.clients.consumer.ConsumerPartitionAssignor` (e.g. `RangeAssignor`, `RoundRobinAssignor`, `StickyAssignor`, `CooperativeStickyAssignor`) over the member list + subscriptions to build a member → partitions map. 2. The leader sends a **SyncGroup** request carrying the full assignment map. **Followers** send empty SyncGroup requests and simply wait. 3. The coordinator persists the assignment and replies to **each member with only their own assignment**. 4. Members install their partitions and begin (or resume) fetching; the group transitions to **Stable**. ## Key design points / edge cases - The **coordinator never computes assignment** — it only relays. Asymmetry (leader gets all, followers get nothing in JoinGroup) is intentional. - If members advertise incompatible protocol sets, JoinGroup can fail with `INCONSISTENT_GROUP_PROTOCOL`. - If the leader fails between JoinGroup and SyncGroup, the SyncGroup times out and a new rebalance (new generation) starts. - With **incremental cooperative rebalancing** (CooperativeStickyAssignor), assignment may take *two* rebalances: first revoke moved partitions, then assign them, so members keep partitions they retain instead of a stop-the-world revoke-all. - The new **consumer group protocol (KIP-848)** moves assignment to the broker-side coordinator and largely removes the client-side JoinGroup/SyncGroup leader role — worth naming as the modern alternative.

  • Why does Kafka run partition assignment on a client (the leader) instead of on the broker?
    So assignment strategy is pluggable client-side — you can supply a custom ConsumerPartitionAssignor without changing or redeploying brokers. The coordinator only needs to know membership, not assignment logic. (KIP-848 later moves this server-side for the new protocol.)
  • What happens if the group leader dies after JoinGroup but before SyncGroup completes?
    Its SyncGroup never arrives; the operation times out, the generation is abandoned, and the coordinator starts a fresh rebalance electing a new leader.

saying these in an interview costs you the question

  • Saying the coordinator computes the partition assignment (it does not in the classic protocol).
  • Claiming all members compute their own assignment independently (only the leader computes).
  • Forgetting that followers send empty JoinGroup/SyncGroup payloads and get back only their slice.

context