skip to content

What are the generation ID and member.id in the consumer group protocol, and how do they protect against stale/zombie members?

level: seniorimportance: should knowfreq 45%

answer

  1. generation = fencing token, bumped per rebalance
  2. ILLEGAL_GENERATION rejects stale members
  3. member.id issued by coordinator, UNKNOWN_MEMBER_ID rejects zombies
  4. empty member.id on first join -> MEMBER_ID_REQUIRED (KIP-394)
  5. group.instance.id = static membership (KIP-345)

basics

~20 s

The generation ID is a monotonically increasing counter the coordinator bumps on each successful rebalance; member.id uniquely identifies a member within the group. Requests carrying a stale generation or unknown member.id are rejected, fencing zombies.

solid answer

~50 s

Each completed rebalance produces a new **generation** — a monotonically increasing integer the coordinator assigns to fence out members from previous rounds. Every member also has a **member.id**, a coordinator-issued unique identifier (UUID-like) for that member within the group. After joining, members stamp their generation.id and member.id on Heartbeat, SyncGroup, and OffsetCommit requests. If a member sends a request with a **stale generation**, the coordinator returns `ILLEGAL_GENERATION`; if the member.id is unrecognized it returns `UNKNOWN_MEMBER_ID`. Either forces the member to rejoin. This prevents a slow or partitioned 'zombie' consumer from committing offsets or acting on a partition it no longer owns after a rebalance moved that partition elsewhere — the generation acts like a fencing token. KIP-345 added `group.instance.id` (static membership) so a restarting member keeps its identity across short outages without bumping the generation.

go deeper

for a junior

Know generation.id increases each rebalance and member.id identifies a member.

for a middle

Explain ILLEGAL_GENERATION / UNKNOWN_MEMBER_ID and that the coordinator issues member.id.

for a senior

Articulate the zombie-fencing role and the MEMBER_ID_REQUIRED / static-membership nuances.

for a principal

Relate the fencing-token model to exactly-once concerns and contrast with KIP-848 epochs.

## The zombie problem Distributed consumers can stall (long GC pause, network partition, slow `poll()` loop). While one is frozen, the coordinator may rebalance and hand its partitions to another member. If the frozen consumer wakes up and naively commits offsets or processes records for a partition it no longer owns, you get **double processing and corrupted offsets**. Generation IDs and member IDs are the fencing mechanism that prevents this. ## generation.id - An **integer counter** maintained by the coordinator for the group. - It is **incremented on every successful rebalance** (every time the group reaches Stable through PreparingRebalance/CompletingRebalance). - Members learn their current generation in the JoinGroup response and must include it on subsequent **Heartbeat, SyncGroup, and OffsetCommit** requests. - A request bearing an *older* generation is rejected with **`ILLEGAL_GENERATION`**, forcing the member to rejoin and pick up the new assignment. The generation thus behaves like a **fencing token** — only the current generation can act. ## member.id - A **unique identifier** for a member within the group, **assigned by the coordinator** (the client sends an empty member.id on its first JoinGroup; the coordinator issues one, often returning `MEMBER_ID_REQUIRED` first so the member retries with it — this guards against rebalance storms from misbehaving clients, per KIP-394). - The coordinator uses it to track each member's heartbeats and session timeout. - A request with an unknown/expired member.id gets **`UNKNOWN_MEMBER_ID`**, forcing a full rejoin. ## How they work together to fence zombies A zombie consumer that missed a rebalance carries a stale generation and possibly an expired member.id. Its OffsetCommit is rejected, so it cannot overwrite the offsets of the member that legitimately took over the partition. It must rejoin (new generation, fresh assignment) before it can do anything. ## Static membership (KIP-345) Setting `group.instance.id` makes a member **static**: on a clean restart within `session.timeout.ms` it rejoins with the same identity and **keeps its assignment without triggering a rebalance or bumping the generation**. This reduces churn for rolling restarts but means a truly dead static member is only detected after the session timeout. ## Edge cases - A rebalance that results in the *same* assignment still bumps the generation. - During failover, the new coordinator restores the last generation from the `__consumer_offsets` log. - The newer KIP-848 protocol replaces the JoinGroup/SyncGroup generation handshake with an epoch-based reconciliation but keeps the same fencing intent.

  • What error does a zombie consumer get when it tries to commit after missing a rebalance, and what must it do?
    ILLEGAL_GENERATION (or UNKNOWN_MEMBER_ID if its member.id expired). The commit is rejected and the consumer must rejoin the group, getting a fresh assignment for the new generation before it can proceed.
  • How does group.instance.id (static membership) change generation behavior on restart?
    A static member restarting within session.timeout.ms rejoins with the same identity and keeps its existing assignment without forcing a rebalance, so the generation is not bumped — reducing churn during rolling restarts.

saying these in an interview costs you the question

  • Claiming the client picks its own member.id (the coordinator assigns it).
  • Saying generation only changes when membership changes — it bumps on every successful rebalance, even with identical assignment.
  • Confusing generation.id with offset numbers — it's a rebalance epoch, not a record position.

context