skip to content

How do you tune static membership for zero-rebalance rolling restarts, and what trade-off does the session timeout introduce?

level: seniorimportance: should knowfreq 45%

answer

  1. stable id + session.timeout.ms > restart time
  2. heartbeat ≈ session/3
  3. one instance at a time
  4. long timeout = slow crash detection / growing lag
  5. cap = group.max.session.timeout.ms (30 min)

basics

~20 s

Give each instance a stable group.instance.id and set session.timeout.ms longer than a single instance's worst-case restart time. Then rolling-restart one instance at a time so each returns inside its window — no rebalances. The trade-off: a longer timeout slows detection of genuinely crashed instances.

solid answer

~40 s

To get zero-rebalance rolling restarts: (1) assign each instance a stable, unique group.instance.id (e.g. from a StatefulSet ordinal); (2) raise session.timeout.ms above the worst-case time for one instance to restart and rejoin — including JVM/app startup, often 2-5 minutes; (3) keep heartbeat.interval.ms at roughly a third of the session timeout; (4) ensure it stays within the broker's group.max.session.timeout.ms (default 30 min). Restart instances one at a time so each returns before its session expires; the coordinator hands back the same partitions and the group never rebalances. The trade-off is failure-detection latency: with a long session timeout, a real crash leaves that instance's partitions unconsumed until the timeout elapses, so lag accumulates longer before another consumer (after a rebalance) takes over. Size the timeout to balance smooth deploys against acceptable recovery time.

go deeper

for a junior

Know that a stable id plus a long-enough timeout avoids deploy rebalances.

for a middle

Set session.timeout.ms above restart time and heartbeat ~1/3 of it; restart one at a time.

for a senior

Articulate the timeout vs crash-detection trade-off and broker bounds; measure real restart time.

for a principal

Co-design with cooperative-sticky, deploy batching, and lag SLOs across the whole fleet.

## Goal: a deploy with no stop-the-world rebalances A **rolling restart** replaces consumer instances one (or a few) at a time. With dynamic membership each replacement triggers rebalances, redistributing partitions repeatedly and spiking lag. Static membership lets a returning instance reclaim its exact partitions with **no rebalance**, if it comes back fast enough. ## The knobs 1. **`group.instance.id`** — stable + unique per instance. Without it, nothing else matters. Derive from a deterministic per-instance source (Kubernetes StatefulSet pod ordinal, fixed slot). 2. **`session.timeout.ms`** — the grace window the coordinator waits for a missing heartbeat before declaring the member dead and rebalancing. This must exceed the **end-to-end restart time** of one instance: process stop → container/JVM start → app/Spring init → broker reconnect → rejoin. For a JVM service that's often **60-300s**, well above the 45s default. 3. **`heartbeat.interval.ms`** — keep around `session.timeout.ms / 3` so liveness is detected promptly while a member is up. 4. **Broker bounds** — `group.max.session.timeout.ms` (default 30 min) caps what you can request; the broker rejects joins outside `[group.min.session.timeout.ms, group.max.session.timeout.ms]`. 5. **`max.poll.interval.ms`** is separate — it governs eviction for slow *processing* between polls, not restarts; leave it sized to your processing time. ## Procedure - Restart **one instance at a time** (or a bounded batch smaller than what your timeout can absorb). - Each restarted instance must rejoin within `session.timeout.ms`; the coordinator matches the static id and returns the same assignment. Group stays stable: **zero rebalances** across the whole deploy. ## The core trade-off A larger `session.timeout.ms` makes deploys smooth but **delays crash detection**. If an instance genuinely dies (hardware, OOM-kill, network partition), the coordinator won't reassign its partitions until the full timeout elapses. During that window those partitions are **unowned** and their **lag grows** — no consumer is making progress on them. So: - **Too short**: deploys/restarts overrun the window → rebalances return, defeating the purpose. - **Too long**: real failures take longer to recover, hurting availability/freshness. Pick the smallest timeout that still comfortably covers your real restart time (measure it), and keep restart batches small. ## Related correctness notes - Subscription or topic-partition metadata changes still rebalance regardless. - Pair static membership with the **cooperative-sticky** assignor so that the rare unavoidable rebalance is incremental rather than stop-the-world — they're complementary, not mutually exclusive. - Auto-commit + a long timeout means a crashed instance's last commits may lag; understand your offset-commit cadence when reasoning about reprocessing after the eventual rebalance.

  • Why combine static membership with the cooperative-sticky assignor?
    Static membership avoids rebalances on transient restarts; cooperative-sticky makes any rebalance that does happen incremental (only affected partitions move) instead of stop-the-world. Together they minimize both rebalance frequency and rebalance cost.
  • If session.timeout.ms is set too high, what symptom appears on a real crash?
    The dead instance's partitions stay unassigned and their consumer lag grows until the timeout elapses and a rebalance reassigns them — slow failure recovery.

saying these in an interview costs you the question

  • Setting session.timeout.ms huge with no regard for crash-detection latency.
  • Restarting all instances at once and expecting no rebalance.
  • Confusing session.timeout.ms with max.poll.interval.ms.
  • Thinking heartbeat.interval.ms should equal or exceed session.timeout.ms.

context