In a Kafka-style consumer group, what is a 'rebalance,' and what effect can it have on message ordering and processing guarantees while it's happening?
answer
- group coordinator reassigns partition ownership
- eager = stop-the-world, cooperative = incremental
- log order untouched; commit timing is the risk
- uncommitted processed records get reprocessed
- rebalance storms from slow poll loops
basics
~20 sA rebalance is when the group reshuffles which consumer reads which partition (e.g., because someone joined or left). During it, processing pauses briefly, and a partition can move to a different consumer that has no memory of what was being processed, so you can see duplicate or stalled processing, though the log order itself isn't changed.
solid answer
~40 sA rebalance is triggered when group membership changes (a consumer joins, crashes, or is considered dead by the group coordinator) or partition count changes, and it reassigns partition ownership among the live consumer instances. With the classic eager rebalance protocol, all consumers stop consuming, revoke all their partitions, and only resume once reassignment completes - a stop-the-world pause across the whole group. Ordering within each partition's log is untouched; what's at risk is exactly-once/at-least-once processing semantics at the boundary: if a consumer had processed messages but not yet committed offsets when a partition is revoked from it, the new owner re-reads from the last committed offset and reprocesses those messages, producing duplicates downstream unless the consumer is idempotent. Newer cooperative/incremental rebalancing narrows the pause to only the partitions actually being reassigned.
go deeper
Should know a rebalance reassigns partitions between consumers and that it can cause a brief processing pause.
Should explain why duplicates can occur around a rebalance (uncommitted offsets) and that the partition log itself isn't reordered.
Should compare eager vs cooperative rebalancing, know common triggers (deploys, crashes, timeouts), and design consumers to be idempotent around rebalance boundaries.
Should diagnose and prevent rebalance storms at the fleet level (poll-interval/session-timeout tuning, assignor choice) and set organization-wide consumer defaults that keep routine deploys from causing visible lag or duplicate-processing incidents.
## What a rebalance is A **consumer group** is Kafka's mechanism for letting multiple consumer instances split the work of reading a topic's partitions, with the constraint that each partition is owned by exactly one consumer in the group at a time (so within a partition there's still a single reader maintaining order). A **rebalance** is the process by which the **group coordinator** (a broker) reassigns which consumer owns which partitions. Rebalances are triggered by group membership changes: - a new consumer instance joins (e.g., during a rolling deploy or scale-out); - an existing one leaves gracefully; - one is declared dead because it missed its heartbeat/session-timeout window (e.g., a long GC pause or a slow poll loop); - or the topic's partition count changes. ## The two protocols Mechanically, Kafka's newer cooperative/incremental rebalancing protocol (via `CooperativeStickyAssignor`) improves on the older eager rebalance protocol: | Eager | Cooperative, incremental | |---|---| | every member of the group revokes all of its currently assigned partitions before the coordinator computes and hands out a brand-new assignment | only the partitions that are actually moving get revoked and reassigned, while consumers keep reading their unaffected partitions throughout | | this is a stop-the-world event for the entire group — no partition in the topic is being consumed by anyone for the duration of the rebalance, even partitions whose ownership doesn't actually change | the assignor tries to keep partitions with their previous owner where possible ("stickiness") to minimize churn | ## Why the mechanism exists This mechanism exists because you need dynamic membership to make consumer groups elastic and fault-tolerant: instances should be able to scale in and out, and a crashed instance's partitions need to be picked up by a survivor without manual intervention, or the group would silently stop making progress on those partitions. Rebalancing is the coordination protocol that keeps "every partition owned by exactly one live consumer" true as the group's membership changes. ## The trade-off The trade-off is availability/liveness during the rebalance window versus complexity. - **Eager rebalancing** is simple to reason about but costs a full-group pause — with a group of many consumers and a slow rebalance (e.g., large state to restore, or consumers with slow `onPartitionsAssigned` callbacks), this pause can stretch to seconds or tens of seconds, directly adding to end-to-end latency and consumer lag across the entire topic, not just the partitions that moved. - **Cooperative rebalancing** reduces blast radius but is more complex to implement correctly (it requires two rounds of the protocol) and doesn't eliminate the pause for the partitions that do move. ## What is actually at risk: offset commits The ordering-relevant risk isn't that the log itself gets reordered — offsets and the physical order of records in a partition are completely unaffected by a rebalance; that data is durable and immutable regardless of which consumer is reading it. The risk is at the processing boundary: **offset commits**. A consumer typically processes a batch of records and then commits the offset marking "I've handled up through here." If a rebalance revokes a partition from a consumer after it processed some records but before it committed the offset for them, the new owner resumes from the last committed offset, which is earlier than what was actually processed — so those records get delivered and processed a second time. This is why Kafka's default delivery semantics are "at-least-once": duplicates around rebalance boundaries are expected, and consumers that need correctness under this must either - be naturally idempotent (e.g., upserts keyed by a business ID), - or track processed-record IDs explicitly to detect and skip duplicates. ## Rebalance storms A second, subtler failure mode is a **rebalance storm**: if consumer instances are misconfigured with an aggressive `max.poll.interval.ms` relative to how long each poll's processing actually takes, a slow batch causes the coordinator to declare the consumer dead, triggering a rebalance, which itself adds load and latency, which can cause the next consumer to also miss its window, cascading into repeated rebalances that make almost no forward progress — a classic production incident pattern often diagnosed by seeing heartbeat-failure log lines correlating with lag spikes. ## Where it shows up A concrete real-world example: rolling deployments of a consumer service (say, a Kubernetes rolling update of 6 pods in a group reading a 12-partition topic) intentionally trigger a rebalance per pod restart; teams commonly tune session.timeout.ms/max.poll.interval.ms and adopt cooperative sticky assignment specifically to keep these routine, expected rebalances from causing visible lag spikes or duplicate-processing storms during normal deploys, distinguishing "ordering is preserved because it's a log property" from "processing exactly-once is not, because that's a consumer-offset-commit property."
- Does a rebalance ever change the order of records within a partition?No - partition order is a property of the durable log itself and is completely independent of which consumer is reading it; a rebalance only changes ownership/assignment, never the stored order of records.
- How does cooperative rebalancing reduce the impact compared to eager rebalancing?Cooperative (incremental) rebalancing only revokes and reassigns the specific partitions that are actually changing owners, letting consumers keep processing their unaffected partitions throughout, whereas eager rebalancing revokes every partition from every consumer first, pausing the entire group even for assignments that don't change.
- What's a practical way to reduce duplicate processing caused by rebalances?Make processing idempotent (e.g., upsert by business key rather than append) so reprocessing a record after a rebalance produces the same end state, and/or commit offsets frequently and synchronously right after processing each small batch rather than in large infrequent batches, shrinking the reprocessing window.
It's like reshuffling which cashier is assigned to which checkout lane mid-shift: the items already in each lane stay in the same order, but if a cashier had scanned some items and not yet finalized the receipt when they got reassigned, the next cashier might rescan those same items from the last finalized point.
saying these in an interview costs you the question
- Thinks a rebalance reorders or loses messages in the log
- Assumes exactly-once delivery is guaranteed by default during rebalances
- Doesn't know the difference between eager and cooperative rebalancing
- Can't explain why uncommitted offsets lead to duplicate processing
- Unaware that overly aggressive poll-interval settings can trigger rebalance storms