What is the Group Coordinator in Kafka, and how is the coordinator broker for a particular consumer group chosen?
answer
- hash(group.id) % __consumer_offsets partitions
- coordinator = leader of that partition
- FindCoordinator request first
- default 50 offset partitions
- failover via partition leader election
basics
~10 sThe Group Coordinator is a broker that manages a consumer group: it handles members joining/leaving, triggers rebalances, and stores offsets. The coordinator is the broker that leads the __consumer_offsets partition the group maps to.
solid answer
~40 sEvery consumer group is managed by one broker acting as its Group Coordinator. Kafka derives which broker by hashing the group.id and taking it modulo the number of partitions in the internal __consumer_offsets topic: hash(group.id) % numPartitions (default 50). The broker that is the leader of that partition is the coordinator. A consumer first sends a FindCoordinator request to any broker to discover the coordinator's host, then directs JoinGroup, SyncGroup, Heartbeat, and OffsetCommit requests to it. Because the coordinator is tied to a __consumer_offsets partition leader, if that broker fails the partition leadership moves to another broker, which then becomes the new coordinator and reloads group/offset state from the partition log.
go deeper
Know it's a broker that manages a group and that the coordinator is found via FindCoordinator.
Explain the hash(group.id) % numPartitions selection and FindCoordinator discovery flow.
Discuss failover, log replay to rebuild state, and co-location of offsets with the coordinator.
Reason about why piggybacking on partition-leader election is a good design and its operational implications (e.g. __consumer_offsets replication factor).
## The problem it solves A consumer group is a set of consumer processes that cooperatively read a topic so that each partition is consumed by exactly one member. Something must decide *which member reads which partition*, detect when members come and go, and remember *how far* each group has read (committed offsets). That coordination authority is the **Group Coordinator**. ## What the coordinator is The Group Coordinator is not a separate process — it is a role played by an ordinary **broker**. Every group is assigned exactly one coordinator broker at a time. The coordinator: - Accepts members into the group (JoinGroup) and assigns each a `member.id`. - Drives **rebalances** (recomputing partition assignments when membership changes). - Tracks liveness via **heartbeats**. - Persists and serves **committed offsets**. ## How the coordinator broker is selected Kafka stores group metadata and offsets in an internal compacted topic called **`__consumer_offsets`** (50 partitions by default, controlled by `offsets.topic.num.partitions`). For a given group: ``` partition = Utils.abs(group.id.hashCode()) % offsets.topic.num.partitions ``` The broker that is the **leader of that `__consumer_offsets` partition** is the coordinator for the group. This neatly piggybacks coordinator election on Kafka's existing partition-leader election — no extra election protocol is needed. ## Discovery flow 1. A consumer sends a **FindCoordinator** request (keyed by group.id) to any bootstrap broker. 2. The broker computes the partition and replies with the coordinator's host/port. 3. The consumer sends all group requests (JoinGroup, SyncGroup, Heartbeat, OffsetCommit/Fetch) to that coordinator. ## Failover / edge cases - If the coordinator broker dies, the `__consumer_offsets` partition's leadership fails over to another in-sync replica; that broker becomes the new coordinator and **replays the partition log** to rebuild group + offset state. - During failover, consumers get a `COORDINATOR_NOT_AVAILABLE` / `NOT_COORDINATOR` error and re-issue FindCoordinator. - Because the same hashing also routes **offset commits**, the coordinator and the offset storage always live on the same broker — keeping commits local and fast.
- What happens to the group if the coordinator broker crashes?The __consumer_offsets partition it led fails over to another replica; that broker becomes the new coordinator, reloads group/offset state from the partition log, and consumers re-issue FindCoordinator after a NOT_COORDINATOR error.
- Why is offset storage co-located with the coordinator?Both are keyed by hash(group.id) onto the same __consumer_offsets partition, so the coordinator broker also holds the group's committed offsets — keeping OffsetCommit/Fetch local and consistent with group lifecycle.
saying these in an interview costs you the question
- Saying there is a single global coordinator for the whole cluster (each group has its own).
- Claiming the coordinator is the controller broker — it is not; it's the __consumer_offsets partition leader.
- Thinking ZooKeeper picks the coordinator (it's derived from group.id hashing onto __consumer_offsets).