What is a consumer group rebalance in Kafka, and what triggers one?
answer
- one partition -> one consumer per group
- join / leave / evict / subscription change
- coordinator orchestrates, leader computes
- eager = stop-the-world pause
- max.poll.interval.ms eviction
basics
~10 sA rebalance is when Kafka redistributes a topic's partitions among the consumers in a group. It is triggered when a consumer joins or leaves the group, or when the set of subscribed partitions changes.
solid answer
~40 sA consumer group rebalance is the process by which Kafka decides which consumer instance reads which partitions, so every partition has exactly one owner within the group. It is coordinated by a broker acting as the group coordinator. Triggers include: a new consumer joining (sending a JoinGroup request), an existing consumer leaving gracefully (LeaveGroup) or being evicted after missing heartbeats (session.timeout.ms) or exceeding max.poll.interval.ms, and changes to the subscribed partition set (e.g. a new partition added, or a regex subscription matching a new topic). During an eager rebalance all consumers stop processing, give up their partitions, and receive a fresh assignment, which is why frequent rebalances hurt throughput.
go deeper
Know the one-line definition: partitions get redistributed among group members, triggered by join/leave/subscription changes.
Distinguish the trigger types and know that eager rebalancing is a stop-the-world pause that hurts throughput.
Explain the coordinator-vs-leader split, heartbeat vs poll-interval eviction, and how static membership avoids needless rebalances.
Reason about rebalance storms in large groups, tuning session/poll timeouts and static membership, and how the protocol choice shapes availability during deploys.
## What a consumer group is In Kafka, a **topic** is split into **partitions** — ordered, independent logs. A **consumer group** is a set of consumer processes sharing a `group.id` that cooperate to read a topic. Kafka guarantees that, within a group, **each partition is assigned to exactly one consumer** at a time. This is how Kafka scales reads horizontally: add more consumers (up to the partition count) and the load spreads. ## What a rebalance is A **rebalance** is the protocol that (re)computes the mapping of partitions to consumers. The result is a *partition assignment*. A broker is elected as the **group coordinator** for each group; it tracks membership and orchestrates rebalances. One consumer is chosen as the **group leader** and it actually runs the assignment algorithm (the assignor) and reports the result back through the coordinator. ## What triggers a rebalance 1. **A consumer joins** — on startup it sends a `JoinGroup` request; the coordinator triggers a rebalance so the newcomer gets partitions. 2. **A consumer leaves gracefully** — calling `consumer.close()` sends a `LeaveGroup`, so its partitions can be reassigned immediately. 3. **A consumer is evicted** — if it stops sending heartbeats within `session.timeout.ms`, or if processing a batch takes longer than `max.poll.interval.ms` (so it never calls `poll()` again in time), the coordinator considers it dead and removes it. 4. **The subscription changes** — partitions are added to a subscribed topic, or a pattern/regex subscription (`subscribe(Pattern)`) starts matching a new topic. ## Why it matters Under the classic **eager** protocol, every member revokes ALL its partitions at the start of a rebalance and stops consuming until a new assignment arrives — a 'stop-the-world' pause. So a flapping consumer (e.g. one that keeps exceeding `max.poll.interval.ms`) can cause repeated rebalances that stall the whole group. Newer **incremental cooperative** rebalancing (KIP-429) reduces this by only moving the partitions that actually need to move. ## Edge cases - A rebalance is **group-wide**: it affects every member, not just the one that changed. - Static membership (`group.instance.id`, KIP-345) lets a consumer restart within `session.timeout.ms` *without* triggering a rebalance, keeping its previous assignment. - Manual partition assignment via `assign()` (no `group.id` coordination) bypasses rebalancing entirely.
- What is the difference between session.timeout.ms and max.poll.interval.ms as rebalance triggers?session.timeout.ms governs the heartbeat thread — miss heartbeats for that long and you're evicted as 'dead'. max.poll.interval.ms governs the application thread — if you don't call poll() again within that window (e.g. slow processing), you're evicted even though heartbeats were fine. Since KIP-62 heartbeats run on a background thread, so the two are decoupled.
- Who actually computes the partition assignment — the broker or a consumer?A consumer. The coordinator (a broker) picks one member as the group leader; that leader runs the configured PartitionAssignor and sends the computed assignment back via the coordinator's SyncGroup response. The broker just relays it.
saying these in an interview costs you the question
- Saying the broker computes the assignment (the group leader consumer does it; the broker only coordinates).
- Claiming a rebalance affects only the consumer that changed — it is group-wide.
- Confusing session.timeout.ms (heartbeat) with max.poll.interval.ms (processing time).
- Saying a single partition can be read by multiple consumers in the same group.