skip to content

Rebalancing and Assignment Strategies

The available assignors and the difference between stop-the-world eager rebalancing and incremental cooperative rebalancing. A high-signal question, since cooperative sticky assignment is the fix for most rebalance pain.

part ofApache Kafkaoverview, primer and where to startread it →
on this pageshow

questions

5

What is a consumer group rebalance in Kafka, and what triggers one?

level: juniorimportance: must knowfreq 78%

answer

  1. one partition -> one consumer per group
  2. join / leave / evict / subscription change
  3. coordinator orchestrates, leader computes
  4. eager = stop-the-world pause
  5. max.poll.interval.ms eviction

basics

~10 s

A rebalance is when Kafka redistributes a topic's partitions among the consumers in a group. It is triggered when a consumer joins or leaves the group, or when the set of subscribed partitions changes.

solid answer

~40 s

A consumer group rebalance is the process by which Kafka decides which consumer instance reads which partitions, so every partition has exactly one owner within the group. It is coordinated by a broker acting as the group coordinator. Triggers include: a new consumer joining (sending a JoinGroup request), an existing consumer leaving gracefully (LeaveGroup) or being evicted after missing heartbeats (session.timeout.ms) or exceeding max.poll.interval.ms, and changes to the subscribed partition set (e.g. a new partition added, or a regex subscription matching a new topic). During an eager rebalance all consumers stop processing, give up their partitions, and receive a fresh assignment, which is why frequent rebalances hurt throughput.

go deeper

for a junior

Know the one-line definition: partitions get redistributed among group members, triggered by join/leave/subscription changes.

for a middle

Distinguish the trigger types and know that eager rebalancing is a stop-the-world pause that hurts throughput.

for a senior

Explain the coordinator-vs-leader split, heartbeat vs poll-interval eviction, and how static membership avoids needless rebalances.

for a principal

Reason about rebalance storms in large groups, tuning session/poll timeouts and static membership, and how the protocol choice shapes availability during deploys.

## What a consumer group is In Kafka, a **topic** is split into **partitions** — ordered, independent logs. A **consumer group** is a set of consumer processes sharing a `group.id` that cooperate to read a topic. Kafka guarantees that, within a group, **each partition is assigned to exactly one consumer** at a time. This is how Kafka scales reads horizontally: add more consumers (up to the partition count) and the load spreads. ## What a rebalance is A **rebalance** is the protocol that (re)computes the mapping of partitions to consumers. The result is a *partition assignment*. A broker is elected as the **group coordinator** for each group; it tracks membership and orchestrates rebalances. One consumer is chosen as the **group leader** and it actually runs the assignment algorithm (the assignor) and reports the result back through the coordinator. ## What triggers a rebalance 1. **A consumer joins** — on startup it sends a `JoinGroup` request; the coordinator triggers a rebalance so the newcomer gets partitions. 2. **A consumer leaves gracefully** — calling `consumer.close()` sends a `LeaveGroup`, so its partitions can be reassigned immediately. 3. **A consumer is evicted** — if it stops sending heartbeats within `session.timeout.ms`, or if processing a batch takes longer than `max.poll.interval.ms` (so it never calls `poll()` again in time), the coordinator considers it dead and removes it. 4. **The subscription changes** — partitions are added to a subscribed topic, or a pattern/regex subscription (`subscribe(Pattern)`) starts matching a new topic. ## Why it matters Under the classic **eager** protocol, every member revokes ALL its partitions at the start of a rebalance and stops consuming until a new assignment arrives — a 'stop-the-world' pause. So a flapping consumer (e.g. one that keeps exceeding `max.poll.interval.ms`) can cause repeated rebalances that stall the whole group. Newer **incremental cooperative** rebalancing (KIP-429) reduces this by only moving the partitions that actually need to move. ## Edge cases - A rebalance is **group-wide**: it affects every member, not just the one that changed. - Static membership (`group.instance.id`, KIP-345) lets a consumer restart within `session.timeout.ms` *without* triggering a rebalance, keeping its previous assignment. - Manual partition assignment via `assign()` (no `group.id` coordination) bypasses rebalancing entirely.

  • What is the difference between session.timeout.ms and max.poll.interval.ms as rebalance triggers?
    session.timeout.ms governs the heartbeat thread — miss heartbeats for that long and you're evicted as 'dead'. max.poll.interval.ms governs the application thread — if you don't call poll() again within that window (e.g. slow processing), you're evicted even though heartbeats were fine. Since KIP-62 heartbeats run on a background thread, so the two are decoupled.
  • Who actually computes the partition assignment — the broker or a consumer?
    A consumer. The coordinator (a broker) picks one member as the group leader; that leader runs the configured PartitionAssignor and sends the computed assignment back via the coordinator's SyncGroup response. The broker just relays it.

saying these in an interview costs you the question

  • Saying the broker computes the assignment (the group leader consumer does it; the broker only coordinates).
  • Claiming a rebalance affects only the consumer that changed — it is group-wide.
  • Confusing session.timeout.ms (heartbeat) with max.poll.interval.ms (processing time).
  • Saying a single partition can be read by multiple consumers in the same group.

context

open as a page

Compare RangeAssignor, RoundRobinAssignor, and StickyAssignor. When would you pick each?

level: middleimportance: must knowfreq 70%

basics

~10 s

RangeAssignor assigns contiguous partition ranges per topic (can imbalance with multiple topics). RoundRobinAssignor spreads all partitions evenly across consumers. StickyAssignor balances evenly while trying to keep previous assignments stable across rebalances.

open as a page

Explain the difference between eager (stop-the-world) and incremental cooperative rebalancing (KIP-429).

level: seniorimportance: must knowfreq 62%

basics

~20 s

In eager rebalancing every consumer revokes all its partitions and stops processing until a new assignment arrives (stop-the-world). Incremental cooperative rebalancing (KIP-429) only revokes the partitions that actually need to move, so most partitions keep being processed throughout.

open as a page

How does a ConsumerRebalanceListener's callback behavior differ between eager and cooperative rebalancing, and where should you commit offsets?

level: seniorimportance: should knowfreq 40%

basics

~10 s

ConsumerRebalanceListener has onPartitionsRevoked, onPartitionsAssigned, and onPartitionsLost. Commit offsets in onPartitionsRevoked before giving partitions up. Under eager you're given all partitions; under cooperative you only get the ones actually moving.

open as a page

Your consumer group experiences frequent, expensive rebalances during deploys and under load. What assignor and configs would you tune, and why?

level: principalimportance: should knowfreq 38%

basics

~10 s

Switch to CooperativeStickyAssignor so most partitions keep running during rebalances, enable static membership (group.instance.id) so restarts don't trigger rebalances, and tune session.timeout.ms / heartbeat.interval.ms / max.poll.interval.ms so consumers aren't falsely evicted.

open as a page