skip to content

How does static membership differ from cooperative (incremental) rebalancing, and when would you use each or both?

level: principalimportance: should knowfreq 35%

answer

  1. static = reduce rebalance FREQUENCY (KIP-345)
  2. cooperative-sticky = reduce rebalance COST (KIP-429)
  3. eager = revoke all, stop-the-world
  4. cooperative = two-phase, only moved partitions revoked
  5. use both; orthogonal levers

basics

~20 s

Static membership (group.instance.id) avoids rebalances entirely for transient restarts. Cooperative-sticky rebalancing makes rebalances that DO happen incremental — only affected partitions move instead of stopping the whole group. They solve different problems and are best used together.

solid answer

~40 s

Static membership and incremental cooperative rebalancing both reduce rebalance pain but at different points. Static membership (KIP-345) prevents a rebalance from being triggered at all when a known instance restarts within session.timeout.ms — the returning member reclaims its partitions. It does nothing to soften a rebalance that does occur. Cooperative-sticky rebalancing (KIP-429), set via partition.assignment.strategy=CooperativeStickyAssignor, changes how a rebalance executes: instead of every member revoking all partitions (eager, stop-the-world), only the partitions that must move are revoked, in a two-phase protocol, so unaffected members keep consuming. Use static membership for stable fleets with predictable restarts (rolling deploys, k8s StatefulSets). Use cooperative-sticky whenever rebalances are unavoidable (scaling, topic metadata changes). Best practice is to enable both: static membership minimizes rebalance frequency, cooperative-sticky minimizes the cost of the rare rebalance that still happens.

go deeper

for a junior

Know that both reduce rebalance pain but in different ways.

for a middle

Static avoids rebalance on restart; cooperative-sticky makes a rebalance incremental.

for a senior

Map each to its config/KIP and decide which fits a given workload.

for a principal

Architect both together, handle assignor migration, and reason about elastic vs fixed fleets and stateful Streams.

## Two distinct levers on rebalance pain Rebalances hurt in two dimensions: **how often** they happen and **how disruptive each one is**. Static membership and cooperative rebalancing attack these separately. ### Static membership (KIP-345) — reduce frequency Enabled by a stable, unique **`group.instance.id`** per instance. The coordinator remembers a static member across a restart; if it returns within **`session.timeout.ms`**, the coordinator returns its prior assignment and **no rebalance is triggered**. This targets the most common churn: rolling deploys and transient restarts. It does **not** change what happens when a rebalance genuinely must run. ### Cooperative / incremental rebalancing (KIP-429) — reduce cost The **assignor**, chosen via `partition.assignment.strategy`, controls rebalance *execution*: - **Eager assignors** (legacy default `RangeAssignor`/`RoundRobinAssignor`): on any rebalance, **every** member **revokes all** its partitions, then the coordinator reassigns from scratch. This is **stop-the-world** — the whole group pauses. - **`CooperativeStickyAssignor`** (KIP-429): a **two-phase, incremental** protocol. In the first phase members keep their partitions and only compute what needs to move; in the second phase **only the partitions that change owner** are revoked and reassigned. Members not affected by the change **keep consuming throughout**. "Sticky" means it tries to preserve existing assignments to minimize movement. ### Why they're complementary | Concern | Static membership | Cooperative-sticky | |---|---|---| | Restart of a known instance | **No rebalance** | Would still rebalance (eagerly or incrementally) | | Scaling up/down, topic metadata change | Still rebalances | **Incremental** rebalance, minimal disruption | | Mechanism | stable `group.instance.id` | `partition.assignment.strategy` | | KIP | 345 | 429 | Enabling **both** gives the strongest result: static membership absorbs the routine restart churn so a rebalance rarely fires, and when one is genuinely needed (a new instance is added, a partition count changes, an instance is permanently lost), cooperative-sticky keeps it cheap. ## Operational guidance - **Stable, fixed-size fleet with frequent deploys** (e.g. a k8s StatefulSet): static membership is the big win; derive ids from pod ordinals and size `session.timeout.ms` to restart time. - **Elastic/autoscaling workloads**: cooperative-sticky matters more because scaling events are unavoidable rebalances; static membership still helps for the restart subset. - **Kafka Streams**: similar logic; Streams has long supported static membership and uses cooperative rebalancing by default in modern versions, plus warm-up replicas to move stateful tasks gracefully. ## Subtlety: migrating assignors Switching from an eager to the cooperative assignor across a running group requires a **two-step rolling upgrade** (a transitional period where the list contains both assignors), because all members must agree on the protocol. Static membership has no such migration constraint — it's purely per-instance. ## Don't conflate - Static membership is **not** an assignor; it's an identity feature orthogonal to `partition.assignment.strategy`. - Cooperative rebalancing does **not** prevent rebalances; it makes them incremental. Only static membership *avoids* the rebalance on a known restart.

  • Does cooperative-sticky rebalancing make static membership unnecessary?
    No. Cooperative-sticky only makes a rebalance incremental; static membership avoids the rebalance entirely for known restarts. They address different dimensions and are best combined.
  • What's special about switching a live group to the CooperativeStickyAssignor?
    It needs a two-phase rolling upgrade — all members must agree on the protocol, so you temporarily list both the old and cooperative assignors, then drop the old one in a second roll.

saying these in an interview costs you the question

  • Saying static membership and cooperative rebalancing are the same thing or alternatives.
  • Claiming cooperative-sticky prevents rebalances (it only makes them incremental).
  • Treating group.instance.id as an assignment strategy.
  • Forgetting that eager assignors are stop-the-world.

context