Your consumer group experiences frequent, expensive rebalances during deploys and under load. What assignor and configs would you tune, and why?
answer
- CooperativeSticky = cheaper rebalances
- group.instance.id = static membership, restart w/o rebalance
- max.poll.interval.ms / max.poll.records for slow processing
- heartbeat ~ 1/3 session.timeout.ms
- group.initial.rebalance.delay.ms batches startup
basics
~10 sSwitch to CooperativeStickyAssignor so most partitions keep running during rebalances, enable static membership (group.instance.id) so restarts don't trigger rebalances, and tune session.timeout.ms / heartbeat.interval.ms / max.poll.interval.ms so consumers aren't falsely evicted.
solid answer
~40 sFrequent expensive rebalances usually come from (a) the eager protocol stopping the whole group and (b) members being evicted unnecessarily. First, set partition.assignment.strategy to CooperativeStickyAssignor (migrating in two phases) so only moving partitions pause — huge win during rolling deploys and autoscaling. Second, enable static membership with a stable group.instance.id (KIP-345): a consumer that restarts within session.timeout.ms keeps its prior assignment and triggers NO rebalance, ideal for k8s rolling restarts. Third, prevent false evictions: raise max.poll.interval.ms if processing batches is slow, or lower max.poll.records so each poll loop finishes in time; keep heartbeat.interval.ms at roughly a third of session.timeout.ms. On the broker, group.initial.rebalance.delay.ms batches startup joins so a fleet starting together rebalances once. Together these cut both rebalance frequency and per-rebalance cost.
go deeper
Know that CooperativeStickyAssignor and longer timeouts can reduce rebalance pain.
Pick the right knob per symptom: assignor for cost, static membership + timeouts for frequency.
Separate per-rebalance cost from rebalance frequency and tune each independently, including max.poll.interval.ms vs session.timeout.ms.
Design a fleet-wide deploy/scaling strategy balancing rebalance cost, false-eviction avoidance, and failure-recovery latency; own the two-phase migration.
## Diagnose the two cost drivers Expensive rebalances hurt in two independent ways: 1. **Per-rebalance cost** — under the **eager** protocol the whole group stops processing. Fix: a cooperative assignor. 2. **Rebalance frequency** — members keep joining/leaving/being evicted. Fix: avoid needless membership churn (static membership, timeout tuning, batching). ## Lever 1: CooperativeStickyAssignor (per-rebalance cost) Set `partition.assignment.strategy=org.apache.kafka.clients.consumer.CooperativeStickyAssignor`. Now a rebalance only revokes partitions that must move; everything else keeps flowing (KIP-429). This is the single biggest win for deploy-time pain because each restarting instance no longer freezes the whole group. **Migrate in two phases** (list both cooperative+old, then cooperative only) — never flip in one deploy. ## Lever 2: Static membership (rebalance frequency) Give each consumer a stable `group.instance.id` (KIP-345). When such a 'static' member disconnects and reconnects within `session.timeout.ms`, the coordinator recognizes it and **returns its previous assignment without triggering a rebalance at all**. This is tailor-made for Kubernetes rolling restarts where pods bounce but identity (via StatefulSet ordinal) is stable. Pair it with a `session.timeout.ms` long enough to cover a pod restart. ## Lever 3: Stop false evictions (rebalance frequency) Members get evicted (triggering rebalances) two ways: - **Heartbeat timeout:** background heartbeat thread misses `session.timeout.ms`. Tune `heartbeat.interval.ms` to ~1/3 of `session.timeout.ms`; raise `session.timeout.ms` for flaky networks (within broker-allowed `group.min/max.session.timeout.ms`). - **Poll-interval timeout:** the app thread doesn't call `poll()` within `max.poll.interval.ms` because a batch took too long. Either **raise `max.poll.interval.ms`**, or **lower `max.poll.records`** / speed up processing / move slow work off the poll thread so loops finish in time. This is the most common cause of 'mysterious' rebalances under load. ## Lever 4: Batch startup joins (broker-side) `group.initial.rebalance.delay.ms` (default 3000) makes the coordinator wait briefly on first group formation so a fleet starting simultaneously joins in one rebalance rather than N. Useful when scaling up many consumers at once. ## Lever 5: Graceful shutdown Always `consumer.close()` (sends LeaveGroup) so departures are immediate and clean rather than waiting for a heartbeat timeout — unless you're using static membership and *want* the slot held for a quick restart (in which case a SIGTERM that skips close, or `internal.leave.group.on.close=false`, keeps the assignment). ## Putting it together - Rolling deploy of a stateful group on k8s: CooperativeStickyAssignor + static membership + session.timeout.ms comfortably above restart time = near-zero rebalance impact. - Slow-processing group under load: raise max.poll.interval.ms / lower max.poll.records first; that alone often stops the rebalance storm. ## Trade-offs / edge cases - Long `session.timeout.ms` + static membership means a truly dead consumer's partitions stay unowned (lag accumulates) until the timeout expires — balance recovery time vs churn. - Cooperative adds an extra rebalance round-trip per change; fine in exchange for availability. - Static membership requires unique, stable IDs; duplicate `group.instance.id` causes fencing.
- A consumer is being evicted under load even though its network and heartbeats are healthy. What's the likely cause and fix?Its processing of a batch is exceeding max.poll.interval.ms, so it doesn't call poll() in time and the coordinator evicts it. Fix by lowering max.poll.records (smaller batches per poll), speeding up or offloading the slow processing, or raising max.poll.interval.ms. Heartbeats are decoupled (KIP-62) so they stay healthy while this happens.
- How does static membership avoid a rebalance during a pod restart, and what's the downside?With a stable group.instance.id, a member that rejoins within session.timeout.ms is recognized and handed back its old assignment with no rebalance. The downside: if that member is genuinely dead, its partitions stay unassigned (and lag builds) until session.timeout.ms expires, so you trade faster restarts for slower failure recovery.
saying these in an interview costs you the question
- Reaching only for longer timeouts while ignoring the assignor — that hides churn but leaves the stop-the-world cost.
- Recommending static membership without acknowledging the slower dead-consumer recovery trade-off.
- Switching to CooperativeStickyAssignor in a single deploy (must be two-phase).
- Raising session.timeout.ms to fix slow-processing evictions — those are governed by max.poll.interval.ms, not session.timeout.ms.