How does static membership skip a rebalance when a consumer restarts, and what is session.timeout.ms's role?
answer
- static = no LeaveGroup on close
- heartbeat stops → session.timeout.ms countdown
- rejoin before timeout = same partitions, no rebalance
- tune session.timeout.ms > restart duration
- trade-off: slower crash detection
basics
~20 sA static member keeps its slot during a brief restart because it doesn't send LeaveGroup and the coordinator hasn't expired its session yet. If it rejoins within session.timeout.ms with the same group.instance.id, it gets the same partitions back — no rebalance.
solid answer
~40 sWith dynamic membership, a clean shutdown sends a LeaveGroup request and an ephemeral member ID, so the coordinator immediately rebalances and the restarted process joins as a new member. Static members behave differently: on shutdown they do NOT send LeaveGroup, so the coordinator simply stops receiving heartbeats. session.timeout.ms is the grace window — the coordinator only declares the member dead and triggers a rebalance if no heartbeat arrives for that long. If the same instance restarts and rejoins (re-presenting its group.instance.id) before that timer fires, the coordinator matches the persisted member, returns its previous assignment, and skips the rebalance entirely. So you tune session.timeout.ms to be longer than your worst-case restart duration to cover the whole rolling deploy.
go deeper
Know that quick restart within the timeout avoids a rebalance.
Explain LeaveGroup suppression, heartbeat/session-timeout mechanics, and the timing constraint.
Quantify timeout tuning vs crash-detection latency and broker min/max bounds.
Frame the lag-vs-availability trade-off across a fleet and integrate with deploy tooling timing.
## Heartbeats and session expiry A consumer keeps its membership alive by sending **heartbeats** to the **group coordinator** on a background thread, every `heartbeat.interval.ms` (default 3s). The coordinator considers a member alive as long as it hears a heartbeat within **`session.timeout.ms`** (default 45s in modern Kafka). If that window lapses with no heartbeat, the coordinator declares the member dead and **rebalances** the group. ## Dynamic vs static shutdown **Dynamic member** restart: 1. On graceful close (`consumer.close()`), the member sends an explicit **`LeaveGroup`** request. 2. The coordinator immediately removes it and rebalances (assignment redistributed to survivors). 3. The restarted process joins fresh, gets a **new ephemeral member ID**, triggering a **second** rebalance. Net: two rebalances per restart, partitions bounce around. **Static member** restart: 1. A static member (has `group.instance.id`) **suppresses the `LeaveGroup`** on close. The coordinator keeps the member entry and its assignment. 2. Heartbeats stop, so the coordinator starts the `session.timeout.ms` countdown — but takes no action yet. 3. The process restarts and rejoins, presenting the **same `group.instance.id`**. The coordinator recognizes it, **re-binds the existing assignment**, and the member resumes its old partitions. 4. **No rebalance occurs** — provided step 3 happened within `session.timeout.ms`. ## The timing constraint The entire restart — process death, JVM startup, Spring/app init, reconnect, rejoin — must complete inside `session.timeout.ms`. If it overruns, the coordinator expires the member and rebalances; when the slow instance finally returns it's a fresh join (another rebalance). That's why static membership usually pairs with a **raised** `session.timeout.ms` (e.g. 2-5 minutes) sized to your restart time plus margin. ## Caveats - `session.timeout.ms` must stay within the broker bounds `group.min.session.timeout.ms` / `group.max.session.timeout.ms` (defaults 6s / 30min), or the broker rejects the join. - A **longer** session timeout also means a **genuinely crashed** static instance is detected more slowly — its partitions stay unowned (lag grows) until the timeout fires. This is the central trade-off. - Subscription/topic-metadata changes still trigger rebalances regardless of static membership.
- What's the downside of raising session.timeout.ms a lot?A truly dead static instance is detected only after the (now longer) timeout, so its partitions sit unconsumed and lag grows during that window — slower failure recovery.
- Does a static member ever send LeaveGroup?Not on a normal close. Modern clients can send LeaveGroup on explicit administrative removal, but the default graceful shutdown of a static member suppresses it to preserve the slot.
saying these in an interview costs you the question
- Saying the restart can take any amount of time — it must fit within session.timeout.ms.
- Thinking heartbeat.interval.ms is the expiry window — it's session.timeout.ms.
- Believing static membership eliminates rebalances caused by topic/subscription changes.