skip to content

You need deterministic, static partition ownership (each service instance always owns the same partitions) without rebalance pauses. How do you design this, and what do you give up?

level: principalimportance: should knowfreq 30%

answer

  1. assign() + stable ordinal (StatefulSet pod index)
  2. no rebalance, deterministic, but you own failover
  3. scaling = recompute modulo + roll all
  4. static membership group.instance.id = KIP-345 middle ground
  5. rejoin within session.timeout.ms = no rebalance

basics

~20 s

Use assign() to pin a fixed partition set to each instance based on a stable ordinal (e.g. pod index). You avoid rebalances and get deterministic ownership, but you give up automatic load balancing and failover — you must handle scaling and dead-instance takeover yourself.

solid answer

~50 s

Drop group management: each instance calls assign() with a deterministic partition set derived from a stable identity — e.g. a Kubernetes StatefulSet ordinal mapping pod-i to partitions {i, i+N, ...}. Benefits: zero rebalances, no stop-the-world pauses, fully predictable ownership ideal for stateful local processing or co-locating partition state. What you give up: (1) automatic failover — if an instance dies, nobody picks up its partitions until you redeploy/re-shard, so you need external orchestration (StatefulSet rescheduling, a coordination service) to restore coverage; (2) automatic load balancing on scale-up/down — adding instances requires recomputing the static mapping and restarting; (3) the safety net of the coordinator detecting dead members. A common middle ground is static group membership (group.instance.id, KIP-345): you keep subscribe() and the group, but members have stable IDs so a quick restart within session.timeout.ms avoids a rebalance entirely. Choose assign() only when you truly need deterministic ownership; otherwise static membership gives most of the stability with failover intact.

go deeper

for a junior

Know that assign() pins fixed partitions and avoids rebalances.

for a middle

Explain that you lose automatic failover and must handle scaling manually.

for a senior

Map deterministic ownership to stable identities (pod ordinals) and contrast with the cost of orphaned partitions.

for a principal

Weigh assign() vs static group membership (KIP-345), design failover/scaling orchestration, and tie partition ownership to stateful local data placement.

## Designing static partition ownership The goal: each running instance **always owns the same partitions**, and bouncing or scaling instances does **not** trigger expensive rebalances. There are two main approaches with very different tradeoffs. ### Approach 1 — Manual assignment with assign() Abandon the consumer group's automatic assignment. Each instance computes its partition set from a **stable identity**: - In Kubernetes, a **StatefulSet** gives each pod a stable ordinal (`app-0`, `app-1`, …). Map pod ordinal `i` of `N` replicas to partitions `{ p : p mod N == i }` (or a contiguous range). Each pod calls `consumer.assign(thoseTopicPartitions)`. - The consumer joins **no group** for assignment purposes (you may still set `group.id` to commit offsets to `__consumer_offsets`, or store offsets externally). **Gains:** - **No rebalances** — no stop-the-world pauses, no `onPartitionsRevoked` thrash. Latency is predictable. - **Deterministic ownership** — perfect for **stateful** processing where local state (RocksDB, caches, in-memory aggregates) is co-located with specific partitions; ownership never moves unexpectedly. **What you give up (must engineer yourself):** 1. **Failover.** If pod-2 dies, partitions {2, 2+N, …} are **orphaned** until pod-2 is rescheduled. There is no coordinator to hand them to a survivor. You rely on the orchestrator (StatefulSet reschedules the same ordinal) to restore coverage. During the gap, those partitions lag. 2. **Elastic scaling.** Changing replica count `N` changes the modulo mapping for *everyone* — you must recompute and roll all instances. There's no smooth incremental move. 3. **Liveness detection.** No group coordinator watching heartbeats; you need health checks/orchestration to notice and replace a dead instance. 4. **Over/under-provisioning risk.** If instances > partitions, extras sit idle (same as groups); if instances < partitions, the mapping must cover all partitions or some go unconsumed. ### Approach 2 — Static group membership (KIP-345) Often you don't truly need manual assignment — you just want to **avoid rebalances on routine restarts**. Set a stable `group.instance.id` per instance and keep `subscribe()`: - Members are now **static**: when a member with a known `group.instance.id` leaves and **rejoins within `session.timeout.ms`**, the coordinator gives it back its **same partitions without a rebalance**. Rolling restarts and brief crashes become rebalance-free. - You **keep** automatic failover: if a static member stays gone past the timeout, a rebalance still redistributes its partitions to survivors. This is the recommended middle ground: stable assignment across restarts **and** retained failover. Pair with `CooperativeStickyAssignor` to make any unavoidable rebalances incremental. ### Decision guide | Need | Use | |---|---| | Deterministic ownership tied to instance identity, accept manual failover | assign() (Approach 1) | | Avoid restart rebalances but keep auto failover/balancing | static membership group.instance.id (Approach 2) | | Standard elastic worker pool | plain subscribe() + cooperative assignor | ### Key takeaway Manual assign() buys determinism and zero rebalances at the cost of building your own failover and scaling. For most "stable parallelism" needs, **static group membership** delivers the stability while the coordinator still handles the hard parts.

  • What does static group membership (group.instance.id / KIP-345) buy you over plain assign()?
    It avoids rebalances on routine restarts (a member rejoining within session.timeout.ms keeps its partitions) while still retaining automatic failover and load balancing via the coordinator. assign() gives determinism but makes you build failover yourself.
  • With assign() pinned by pod ordinal, what happens when an instance crashes?
    Its partitions are orphaned and lag until the orchestrator reschedules the same ordinal. There's no coordinator to hand them to a survivor, so you depend on external orchestration for recovery.

saying these in an interview costs you the question

  • Claiming assign() still gives automatic failover.
  • Confusing static group membership (still uses subscribe/group) with manual assign().
  • Saying scaling an assign()-based deployment is seamless — it requires recomputing the mapping and rolling instances.
  • Forgetting you can still commit offsets with assign() if group.id is set.

context