Describe the consumer group states the coordinator maintains (Empty, PreparingRebalance, CompletingRebalance, Stable, Dead) and the transitions between them.
answer
- Empty -> PreparingRebalance -> CompletingRebalance -> Stable
- JoinGroup = Preparing->Completing; SyncGroup = Completing->Stable
- Stable falls back to Preparing on join/leave/timeout
- Empty can hold offsets but no members
- Dead = metadata removed / deleted
basics
~10 sThe coordinator tracks a group's lifecycle: Empty (no members), PreparingRebalance (waiting for members to join), CompletingRebalance (awaiting SyncGroup), Stable (assignment done, consuming), and Dead (group removed). Membership changes drive transitions through these states.
solid answer
~40 sThe coordinator runs each group through a state machine. **Empty**: the group exists (may hold committed offsets) but has no active members. **PreparingRebalance**: triggered when a member joins/leaves or times out; the coordinator collects JoinGroup requests until all known members rejoin or `rebalance.timeout.ms` elapses. **CompletingRebalance** (a.k.a. AwaitingSync): the leader has been chosen and the coordinator is waiting for the leader's SyncGroup with the assignment. **Stable**: assignment is distributed and members are consuming, sending heartbeats. **Dead**: the group has no members and no offsets (or was explicitly deleted) and its metadata is removed. **PreparingRebalance** can also be re-entered from Stable on membership change. You can inspect the current state with `kafka-consumer-groups.sh --describe --state`. These states map directly to the JoinGroup (Preparing→Completing) and SyncGroup (Completing→Stable) protocol steps.
go deeper
Recognize the main states (Empty, PreparingRebalance, Stable) and that rebalances move between them.
Order all states and map JoinGroup/SyncGroup/heartbeat to the transitions.
Explain timeouts (rebalance.timeout.ms, session.timeout.ms) and rebalance-storm flapping.
Tie state recovery to __consumer_offsets log replay and reason about operational diagnosis via --state.
## Why a state machine The coordinator must serialize concurrent member activity (joins, leaves, heartbeats, commits) and know which protocol step the group is in. It models this with an explicit **group state machine** stored in memory (and recoverable from the `__consumer_offsets` log). ## The states - **Empty** — The group is known to the coordinator and may still hold **committed offsets**, but there are currently **no live members**. A brand-new group and a group whose consumers all left both sit here. Offsets in Empty are retained until `offsets.retention.minutes` expires. - **PreparingRebalance** — Entered when membership must change: a new member sends JoinGroup, an existing member leaves (LeaveGroup) or misses its `session.timeout.ms` heartbeat, or a member's subscription changes. The coordinator now **collects JoinGroup requests** from all members, waiting up to `rebalance.timeout.ms` (which the client advertises, derived from `max.poll.interval.ms`). Members that fail to rejoin in time are evicted. - **CompletingRebalance** (historically **AwaitingSync**) — All expected members have rejoined, the **leader is elected**, and JoinGroup responses are sent. The coordinator now **waits for the leader's SyncGroup** carrying the computed assignment. - **Stable** — The leader's assignment has been received and each member has been sent its slice. Members are actively **fetching and heartbeating**. This is the steady state. - **Dead** — The group has no members and no offsets, or was explicitly deleted (`kafka-consumer-groups.sh --delete`). Its metadata is removed; subsequent requests get `UNKNOWN_MEMBER_ID` / coordinator-level errors. ## Transitions (the common cycle) ``` Empty ──(first JoinGroup)──> PreparingRebalance PreparingRebalance ──(all members joined / timeout)──> CompletingRebalance CompletingRebalance ──(leader SyncGroup received)──> Stable Stable ──(member join/leave/timeout/subscription change)──> PreparingRebalance Stable/Empty ──(all members gone + offsets expired or --delete)──> Empty/Dead ``` ## Protocol mapping - **JoinGroup** drives PreparingRebalance → CompletingRebalance. - **SyncGroup** drives CompletingRebalance → Stable. - **Heartbeats** keep a member alive in Stable; a missed heartbeat (session timeout) kicks the group back to PreparingRebalance. The coordinator also signals an in-progress rebalance to other members by returning `REBALANCE_IN_PROGRESS` on their heartbeats so they rejoin. ## Edge cases - A group can flap between Stable and PreparingRebalance under churn (rebalance storms) — often caused by slow processing exceeding `max.poll.interval.ms`. - An Empty group still answers OffsetFetch, which is how tools read committed offsets/lag for a stopped group. - Inspect live state: `kafka-consumer-groups.sh --bootstrap-server ... --describe --group g --state`.
- Which protocol request transitions a group from CompletingRebalance to Stable?The group leader's SyncGroup request: once the coordinator receives the assignment and distributes each member's slice, the group becomes Stable.
- Can a group in the Empty state still have committed offsets, and why does that matter?Yes — Empty means no live members but offsets persist until offsets.retention.minutes expires. This lets monitoring tools read committed offsets and lag for a stopped consumer group.
saying these in an interview costs you the question
- Saying Empty means the group is deleted (that's Dead; Empty can still hold offsets).
- Putting SyncGroup as the Preparing->Completing trigger (that's JoinGroup; SyncGroup is Completing->Stable).
- Claiming a missed heartbeat moves the group straight to Dead instead of back to PreparingRebalance.