skip to content

Heartbeat and Session Liveness

How the background heartbeat thread and session.timeout.ms decide whether a group member is still alive. Comes up constantly in questions about why a group keeps rebalancing.

part ofApache Kafkaoverview, primer and where to startread it →
on this pageshow

questions

5

What is a consumer heartbeat in Kafka, and what does it accomplish?

level: juniorimportance: must knowfreq 70%

answer

  1. background thread, not poll()
  2. tells coordinator 'I'm alive'
  3. miss it for session.timeout.ms -> dead
  4. death triggers rebalance
  5. KIP-62 split heartbeat from poll

basics

~20 s

A heartbeat is a small periodic signal a consumer sends to the group coordinator broker to prove it is alive and still part of its consumer group. If heartbeats stop, the broker assumes the consumer died and reassigns its partitions.

solid answer

~40 s

A heartbeat is a lightweight message a Kafka consumer periodically sends to the group coordinator (the broker that manages its consumer group) to signal liveness and confirm continued membership. In modern clients a dedicated background heartbeat thread sends them on the interval set by heartbeat.interval.ms, independent of the application's poll() calls. The coordinator tracks the last heartbeat per member; if none arrives within session.timeout.ms, it marks the member dead, removes it from the group, and triggers a rebalance so the dead member's partitions get reassigned to surviving members. Heartbeats are how the group stays aware of which members are present, which underpins partition assignment and failure detection.

go deeper

for a junior

Know it's a periodic 'I'm alive' signal to the broker; missing it gets you kicked out of the group.

for a middle

Know the background-thread model, that it goes to the group coordinator, and that timeout triggers a rebalance.

for a senior

Distinguish heartbeat liveness from processing progress (max.poll.interval.ms), and explain LeaveGroup on graceful shutdown.

for a principal

Reason about failure-detection latency vs false-positive evictions and how heartbeat design choices affect group stability at scale.

## What problem heartbeats solve A Kafka **consumer group** is a set of consumer instances that cooperatively read from a topic's partitions, with each partition assigned to exactly one member. To keep that assignment correct, the cluster must know which members are alive. The **group coordinator** — a specific broker chosen to manage a given group — needs a continuous signal that each member is healthy. That signal is the **heartbeat**. ## Mechanism A heartbeat is a small RPC (a `Heartbeat` request) the consumer sends to the coordinator. In modern Java clients (since KIP-62, Kafka 0.10.1) heartbeats are sent by a **dedicated background thread** in the consumer, not on the application thread. This thread fires roughly every `heartbeat.interval.ms` (default 3000 ms). The coordinator records the timestamp of each member's last heartbeat. ## What it accomplishes - **Liveness / failure detection:** if the coordinator does not hear from a member within `session.timeout.ms` (default 45000 ms in recent versions), it declares the member dead, evicts it from the group, and starts a **rebalance** to reassign that member's partitions to the survivors. - **Rebalance signaling:** the heartbeat response also carries a flag telling the consumer when a rebalance is in progress, so the consumer knows it must rejoin the group. - **Membership awareness:** heartbeats are the heartbeat (pun intended) of the membership protocol that drives partition assignment. ## Edge cases / nuances - Heartbeats prove the *network/thread* is alive, not that the application is making progress. A consumer can heartbeat fine while being stuck in slow processing; that case is covered by a separate timer, `max.poll.interval.ms`, not by the heartbeat path. - If the background thread dies or the consumer is GC-paused longer than the session timeout, the member is evicted even though the process technically exists. - A graceful shutdown sends a `LeaveGroup` request so the coordinator can rebalance immediately rather than waiting for the session to expire.

  • Which broker receives the heartbeats?
    The group coordinator — the specific broker elected to manage that consumer group (chosen by hashing the group.id onto a __consumer_offsets partition).
  • Does a heartbeat prove the application is processing records?
    No. It only proves the consumer's background thread and network are alive. Stuck application processing is detected by max.poll.interval.ms, a separate mechanism.

saying these in an interview costs you the question

  • Saying heartbeats are sent inside poll() in modern clients (KIP-62 moved them to a background thread).
  • Claiming a healthy heartbeat means the app is making processing progress.
  • Confusing the heartbeat with the produce/fetch data path.

context

open as a page

Explain the relationship between heartbeat.interval.ms and session.timeout.ms and how you would tune them.

level: middleimportance: must knowfreq 75%

basics

~20 s

heartbeat.interval.ms is how often the consumer sends heartbeats; session.timeout.ms is how long the coordinator waits without one before declaring the consumer dead. The interval must be well below the timeout — the rule of thumb is heartbeat.interval.ms <= session.timeout.ms / 3.

open as a page

Walk through how the coordinator detects a dead member via missed heartbeats and what happens next.

level: seniorimportance: must knowfreq 60%

basics

~20 s

The coordinator tracks each member's last heartbeat. If a member sends none for session.timeout.ms, the coordinator declares it dead, removes it from the group, and starts a rebalance that reassigns the dead member's partitions to the remaining consumers.

open as a page

What are group.min.session.timeout.ms and group.max.session.timeout.ms, and what happens if a consumer requests a session timeout outside them?

level: seniorimportance: should knowfreq 45%

basics

~20 s

They are broker-side limits on the session.timeout.ms a consumer may request: group.min.session.timeout.ms (default 6000 ms) and group.max.session.timeout.ms (default 1800000 ms). A consumer asking for a value outside this range is rejected and cannot join the group.

open as a page

Since KIP-62 split heartbeating from poll(), how do session.timeout.ms and max.poll.interval.ms differ, and why was the split necessary?

level: seniorimportance: should knowfreq 55%

basics

~20 s

session.timeout.ms detects a crashed or network-isolated consumer via missed heartbeats from the background thread. max.poll.interval.ms detects a consumer that is alive but stuck/slow in processing because it hasn't called poll() in time. Before KIP-62 both were conflated into the single session timeout.

open as a page