skip to content

Static Membership

Giving consumers a stable group.instance.id so a short restart does not trigger a rebalance. Interviewers ask it as the fix for rebalance churn during rolling deploys.

part ofApache Kafkaoverview, primer and where to startread it →
on this pageshow

questions

5

What is static membership in Kafka consumer groups, and which config enables it?

level: juniorimportance: must knowfreq 60%

answer

  1. group.instance.id = stable identity
  2. KIP-345, Kafka 2.3+
  3. restart within session.timeout.ms = no rebalance
  4. static member skips LeaveGroup
  5. must be unique per instance

basics

~20 s

Static membership gives a consumer a stable identity via the group.instance.id config (KIP-345). With it, a brief consumer restart does not trigger a group rebalance, because the broker recognizes the returning member by its fixed id.

solid answer

~40 s

Static membership (KIP-345, Kafka 2.3+) is enabled by setting group.instance.id to a unique, persistent value per consumer instance. Normally each consumer gets an ephemeral member ID from the broker every time it joins, so leaving and rejoining forces a rebalance and partition reassignment. With a static group.instance.id, the broker remembers the member across restarts: when the same instance leaves and rejoins within session.timeout.ms, the group coordinator hands it back the same partitions without rebalancing the whole group. This avoids stop-the-world reassignments during rolling restarts, deploys, and transient crashes. The id must be globally unique within the group; reuse causes one instance to be fenced.

go deeper

for a junior

Know that group.instance.id enables static membership and that it avoids a rebalance on quick restarts.

for a middle

Explain the role of session.timeout.ms and the no-LeaveGroup-on-shutdown behavior.

for a senior

Discuss uniqueness/fencing, tuning timeouts, and the rolling-restart benefit quantitatively.

for a principal

Reason about fleet-wide operational impact, StatefulSet id derivation, and failure modes vs cooperative rebalancing.

## The problem static membership solves A Kafka **consumer group** is a set of consumer processes that cooperatively read a topic's partitions; each partition is owned by exactly one member. The **group coordinator** (a broker) tracks membership and runs **rebalances** — the protocol that (re)assigns partitions to members. A rebalance is disruptive: members stop consuming, revoke their partitions, and wait for a new assignment ("stop-the-world"). By default, membership is **dynamic**: every time a consumer calls `poll()`/joins, the coordinator assigns it a fresh, ephemeral **member ID**. So if a consumer process restarts (deploy, crash, k8s pod reschedule), the coordinator sees the old member leave and a brand-new member join — triggering **two** rebalances (one on leave, one on join). For large groups this is costly and causes consumer lag spikes. ## What static membership does **KIP-345** (Apache Kafka 2.3) introduced **static membership**. You assign each consumer instance a stable identity via the consumer config: ``` group.instance.id=worker-3 ``` This value must be **unique per instance** and **persistent across restarts** (e.g. derived from a k8s StatefulSet ordinal or a fixed hostname). When a consumer with a `group.instance.id` joins, the coordinator records it as a **static member**. - On a **clean shutdown**, a static member does **not** send a `LeaveGroup` request (unlike a dynamic member), so the coordinator keeps its slot. - When the same instance **rejoins within `session.timeout.ms`**, the coordinator recognizes the id and **returns the same partition assignment** — no rebalance, no reassignment. - Only if the instance stays gone **longer than `session.timeout.ms`** does the coordinator expire it and rebalance. ## Key terms - **`group.instance.id`**: the static, unique identity string. Its presence is what makes a member static. - **`session.timeout.ms`**: how long the coordinator waits without a heartbeat before declaring a member dead. For static membership you typically **raise** this (e.g. from the default 45s) so transient restarts fit inside the window. - **Fencing**: if two live instances share the same `group.instance.id`, the broker keeps only one and rejects the other with `FencedInstanceIdException`. ## Why it matters Static membership turns a routine deploy of an N-instance consumer fleet from N stop-the-world rebalances into (ideally) **zero** rebalances, dramatically cutting consumer lag and processing pauses during rolling restarts.

  • What happens if you forget to make group.instance.id unique across instances?
    Two live instances with the same id collide; the broker fences one of them, throwing FencedInstanceIdException, so only one keeps consuming.
  • Does static membership require any broker-side configuration?
    No broker config toggle is needed — it's a consumer-side feature, but brokers and clients must both be 2.3+ to support the protocol.

saying these in an interview costs you the question

  • Saying static membership 'disables rebalancing entirely' — it only skips rebalance for transient restarts within the session timeout.
  • Claiming group.instance.id can be the same for all instances in a group.
  • Confusing group.instance.id with group.id (the group name).

context

open as a page

How does static membership skip a rebalance when a consumer restarts, and what is session.timeout.ms's role?

level: middleimportance: must knowfreq 55%

basics

~20 s

A static member keeps its slot during a brief restart because it doesn't send LeaveGroup and the coordinator hasn't expired its session yet. If it rejoins within session.timeout.ms with the same group.instance.id, it gets the same partitions back — no rebalance.

open as a page

What is the FencedInstanceIdException, and when does a static consumer get fenced?

level: seniorimportance: should knowfreq 40%

basics

~20 s

If two live consumers join the same group with the same group.instance.id, the broker treats them as duplicates of one static member and fences the older one, throwing FencedInstanceIdException so only one instance keeps consuming each set of partitions.

open as a page

How do you tune static membership for zero-rebalance rolling restarts, and what trade-off does the session timeout introduce?

level: seniorimportance: should knowfreq 45%

basics

~20 s

Give each instance a stable group.instance.id and set session.timeout.ms longer than a single instance's worst-case restart time. Then rolling-restart one instance at a time so each returns inside its window — no rebalances. The trade-off: a longer timeout slows detection of genuinely crashed instances.

open as a page

How does static membership differ from cooperative (incremental) rebalancing, and when would you use each or both?

level: principalimportance: should knowfreq 35%

basics

~20 s

Static membership (group.instance.id) avoids rebalances entirely for transient restarts. Cooperative-sticky rebalancing makes rebalances that DO happen incremental — only affected partitions move instead of stopping the whole group. They solve different problems and are best used together.

open as a page