Walk through how the coordinator detects a dead member via missed heartbeats and what happens next.
answer
- coordinator tracks last heartbeat
- no HB for session.timeout.ms -> dead
- PreparingRebalance + generation bump
- survivors see REBALANCE_IN_PROGRESS in HB response
- LeaveGroup on clean close = instant
- static membership waits out the timeout
basics
~20 sThe coordinator tracks each member's last heartbeat. If a member sends none for session.timeout.ms, the coordinator declares it dead, removes it from the group, and starts a rebalance that reassigns the dead member's partitions to the remaining consumers.
solid answer
~40 sEvery member's background thread heartbeats the group coordinator on heartbeat.interval.ms. The coordinator stores the timestamp of each member's most recent heartbeat. A reaper checks for expiry: if a member's session has gone session.timeout.ms without a heartbeat, the coordinator considers it dead. It evicts that member, transitions the group into a rebalance (PreparingRebalance state), and bumps the generation id. Surviving members discover the rebalance through their next heartbeat response (which signals REBALANCE_IN_PROGRESS), then rejoin via JoinGroup; the assignor recomputes assignments, and the dead member's partitions are handed to live consumers via SyncGroup. A clean shutdown short-circuits this: the consumer sends an explicit LeaveGroup so the coordinator rebalances immediately instead of waiting out the full session timeout. So detection latency is up to session.timeout.ms for crashes, but near-instant for graceful exits.
go deeper
Know: no heartbeats for the timeout window -> kicked out -> partitions reassigned.
Add the rebalance trigger and that survivors rejoin; know clean shutdown is faster.
Explain generation bumps, REBALANCE_IN_PROGRESS signaling, LeaveGroup, and detection-latency = session.timeout.ms.
Design for it: static membership for rolling restarts, cooperative rebalancing to limit stop-the-world, and tuning session.timeout.ms to the recovery SLA.
## Cast of characters - **Group coordinator:** the broker managing a given consumer group (selected by hashing `group.id` to a `__consumer_offsets` partition; the leader of that partition's broker is the coordinator). - **Member:** one consumer instance, identified by a `member.id` the coordinator assigns at join time. - **Generation:** a monotonically increasing integer that versions the group's membership; it increments on each rebalance to fence stale members. ## Step 1 — steady state Each member's background thread sends `Heartbeat` requests every `heartbeat.interval.ms`. The coordinator records the last-seen time per member. ## Step 2 — detection The coordinator runs expiry checks. If `now - lastHeartbeat(member) > session.timeout.ms`, the member's session is expired. The coordinator declares it **dead** and removes it from the group's member list. (A consumer can also be removed for exceeding `max.poll.interval.ms` — that path is signaled differently, via the member failing to rejoin, but the result, eviction + rebalance, is similar.) ## Step 3 — rebalance Losing a member changes group membership, so the coordinator moves the group to **PreparingRebalance**, increments the **generation id**, and waits for members to rejoin. Surviving members learn a rebalance is underway because their next `Heartbeat` response returns `REBALANCE_IN_PROGRESS`. They then: 1. send `JoinGroup` (the coordinator picks a **group leader** — one of the consumers — to run the assignor), 2. the leader runs the configured **partition assignor** (e.g., `CooperativeStickyAssignor`) to compute who gets which partitions, including the dead member's, 3. members fetch their new assignment via `SyncGroup`, 4. the group returns to **Stable** and consumption resumes. The dead member's partitions are now owned by live consumers, so processing continues. ## Graceful vs ungraceful exit - **Crash / network loss:** no `LeaveGroup` is sent, so detection waits up to the full `session.timeout.ms`. This is the worst-case detection latency. - **Clean shutdown / `consumer.close()`:** the client proactively sends a **`LeaveGroup`** request, so the coordinator evicts immediately and rebalances without waiting for the timeout. (Static membership via `group.instance.id` can intentionally suppress this so a quick restart avoids a rebalance — see below.) ## Static membership nuance (KIP-345) With `group.instance.id` set, a member is **static**: on a brief restart it reclaims its old identity and assignment, and the coordinator does **not** immediately rebalance on its disappearance — it waits out `session.timeout.ms`, which is the point (avoid churn during rolling restarts). Without static membership, every leave/join causes a rebalance. ## Why this matters - Detection latency = up to `session.timeout.ms` (tune it for your recovery SLA vs false-positive tolerance). - During the rebalance, with the older eager protocol, consumers stop consuming ('stop-the-world'); the cooperative/incremental protocol (KIP-429) limits disruption to only the moved partitions. ## Edge cases - A GC pause longer than `session.timeout.ms` looks identical to a crash and triggers eviction. - After eviction, if the 'dead' consumer revives and tries to use its old generation, it is fenced (`ILLEGAL_GENERATION`) and must rejoin.
- How do surviving consumers find out a rebalance has started?Their next heartbeat response from the coordinator carries REBALANCE_IN_PROGRESS, prompting them to rejoin via JoinGroup.
- Why might a graceful shutdown rebalance faster than a crash?A graceful close sends an explicit LeaveGroup, so the coordinator evicts immediately. A crash sends nothing, so detection waits up to the full session.timeout.ms.
- How does static membership (group.instance.id) change this?A static member's brief absence does not trigger an immediate rebalance; the coordinator waits out session.timeout.ms, letting a fast restart reclaim the same assignment and avoid churn.
saying these in an interview costs you the question
- Saying a crashed consumer is detected instantly — without LeaveGroup it takes up to session.timeout.ms.
- Claiming the coordinator pushes a notification to evict members proactively, rather than survivors learning via heartbeat responses.
- Ignoring that a long GC pause is indistinguishable from a crash and causes eviction.
- Confusing missed-heartbeat eviction with max.poll.interval.ms eviction (different trigger, similar outcome).