Walk through the JoinGroup and SyncGroup phases of a rebalance. Who computes the partition assignment, and how does it reach each member?
answer
- JoinGroup: leader gets all members, followers empty
- SyncGroup: leader sends assignment, coordinator fans out slices
- coordinator is assignment-agnostic
- PartitionAssignor runs on the leader (client-side)
- KIP-848 moves assignment server-side
basics
~20 sAll members send JoinGroup; the coordinator picks one member as group leader and returns the full member list to it. The leader computes assignments and sends them in SyncGroup; the coordinator fans them out to each member in their SyncGroup response.
solid answer
~40 sA rebalance has two round-trips. In **JoinGroup**, every member sends its supported assignment protocols and subscriptions to the coordinator. The coordinator collects all members for one generation, elects the **group leader** (typically the first to join), and replies: the leader gets the entire member list plus everyone's metadata; followers get an empty member list. The coordinator itself does *not* compute assignment — it stays assignment-agnostic. In **SyncGroup**, the leader runs the configured `PartitionAssignor` (e.g. RangeAssignor, CooperativeStickyAssignor) to map partitions to members, then sends the full assignment map to the coordinator. Followers send empty SyncGroup requests. The coordinator stores the assignment and returns to each member only *their* slice. This client-side assignment design lets you plug in custom assignors without changing the broker.
go deeper
Know the two phases exist and that one member (the leader) decides assignments.
Explain leader election, assignor on the leader, and how the coordinator distributes slices.
Discuss failure modes (leader death, protocol mismatch) and cooperative incremental rebalancing.
Contrast the client-side leader model with KIP-848's server-side assignment and the tradeoffs.
## Why two phases A rebalance redistributes partitions across the current members of a group. Kafka splits the work between the **broker (coordinator)** — which knows *who* is in the group — and a **client (group leader)** — which decides *who gets what*. This keeps assignment logic on the client so it can be customized without broker changes. ## Phase 1 — JoinGroup 1. Each consumer sends a **JoinGroup** request to the coordinator containing its `group.id`, its current `member.id` (empty on first join), the list of **assignment protocols** it supports, and per-protocol metadata (its topic **subscription**, and for sticky assignors its currently owned partitions). 2. The coordinator waits up to `rebalance.timeout.ms` for all known members to (re)join, advancing the group to **PreparingRebalance** state. 3. The coordinator **elects the group leader** — normally the first member to join the new generation. 4. The coordinator picks a single assignment protocol that *all* members support (the protocol must be common to everyone). 5. JoinGroup responses go out: the **leader** receives the complete member list with each member's metadata; **followers** receive an empty member list. Everyone receives the new `generation.id` and their (possibly newly issued) `member.id`. ## Phase 2 — SyncGroup 1. The **leader** runs the chosen `org.apache.kafka.clients.consumer.ConsumerPartitionAssignor` (e.g. `RangeAssignor`, `RoundRobinAssignor`, `StickyAssignor`, `CooperativeStickyAssignor`) over the member list + subscriptions to build a member → partitions map. 2. The leader sends a **SyncGroup** request carrying the full assignment map. **Followers** send empty SyncGroup requests and simply wait. 3. The coordinator persists the assignment and replies to **each member with only their own assignment**. 4. Members install their partitions and begin (or resume) fetching; the group transitions to **Stable**. ## Key design points / edge cases - The **coordinator never computes assignment** — it only relays. Asymmetry (leader gets all, followers get nothing in JoinGroup) is intentional. - If members advertise incompatible protocol sets, JoinGroup can fail with `INCONSISTENT_GROUP_PROTOCOL`. - If the leader fails between JoinGroup and SyncGroup, the SyncGroup times out and a new rebalance (new generation) starts. - With **incremental cooperative rebalancing** (CooperativeStickyAssignor), assignment may take *two* rebalances: first revoke moved partitions, then assign them, so members keep partitions they retain instead of a stop-the-world revoke-all. - The new **consumer group protocol (KIP-848)** moves assignment to the broker-side coordinator and largely removes the client-side JoinGroup/SyncGroup leader role — worth naming as the modern alternative.
- Why does Kafka run partition assignment on a client (the leader) instead of on the broker?So assignment strategy is pluggable client-side — you can supply a custom ConsumerPartitionAssignor without changing or redeploying brokers. The coordinator only needs to know membership, not assignment logic. (KIP-848 later moves this server-side for the new protocol.)
- What happens if the group leader dies after JoinGroup but before SyncGroup completes?Its SyncGroup never arrives; the operation times out, the generation is abandoned, and the coordinator starts a fresh rebalance electing a new leader.
saying these in an interview costs you the question
- Saying the coordinator computes the partition assignment (it does not in the classic protocol).
- Claiming all members compute their own assignment independently (only the leader computes).
- Forgetting that followers send empty JoinGroup/SyncGroup payloads and get back only their slice.