skip to content

Since KIP-62 split heartbeating from poll(), how do session.timeout.ms and max.poll.interval.ms differ, and why was the split necessary?

level: seniorimportance: should knowfreq 55%

answer

  1. KIP-62, Kafka 0.10.1, background thread
  2. session.timeout.ms = crash/network
  3. max.poll.interval.ms = stuck processing
  4. default max.poll = 300000ms
  5. exceed max.poll -> proactive LeaveGroup
  6. old conflation forced inflated session timeout

basics

~20 s

session.timeout.ms detects a crashed or network-isolated consumer via missed heartbeats from the background thread. max.poll.interval.ms detects a consumer that is alive but stuck/slow in processing because it hasn't called poll() in time. Before KIP-62 both were conflated into the single session timeout.

solid answer

~40 s

KIP-62 (Kafka 0.10.1) decoupled liveness from processing progress by moving heartbeats to a background thread. After it, two independent timers exist. session.timeout.ms is the heartbeat-based liveness window: the background thread keeps it satisfied regardless of how slow the app is, so it really detects crashes, GC pauses, and network partitions. max.poll.interval.ms (default 300000 ms) is the processing window: if the application thread doesn't call poll() again within it, the consumer proactively leaves the group, assuming it is stuck (e.g., a slow handler or a deadlock). The split was necessary because, before it, a slow consumer that hadn't polled would stop heartbeating and get falsely evicted — forcing people to inflate session.timeout.ms, which then delayed real crash detection. Now you tune crash detection (session.timeout.ms) and processing tolerance (max.poll.interval.ms) separately.

go deeper

for a junior

Know there are two timers: one for 'is it alive' and one for 'is it processing in time.'

for a middle

Name both configs and their defaults; know slow processing is a max.poll.interval.ms concern.

for a senior

Explain KIP-62's motivation and the proactive LeaveGroup on max.poll.interval.ms breach; diagnose the CommitFailedException.

for a principal

Guide teams on batch sizing, async processing patterns, and separating crash-detection SLAs from processing tolerances.

## The problem before KIP-62 In old clients (< 0.10.1) the *only* place heartbeats were sent was inside `poll()`. So 'is the consumer alive?' was answered by 'how recently did it call poll()?'. A consumer doing slow processing between polls would stop heartbeating and be evicted as dead — a false positive. The only workaround was to raise `session.timeout.ms` high enough to cover the slowest batch, but that also pushed out the time to detect a *genuinely* crashed consumer. Liveness and processing-progress were conflated into one knob, and you couldn't optimize both. ## What KIP-62 changed (Kafka 0.10.1) It introduced a **background heartbeat thread** so heartbeats fire on `heartbeat.interval.ms` independent of `poll()`. This freed two distinct concerns into two configs: ### `session.timeout.ms` — liveness / crash detection - Satisfied by the background thread. - Expires only if the thread stops (crash, long GC pause, network partition). - Bounded by broker `group.min/max.session.timeout.ms`. - This is your **crash/network failure detector**. ### `max.poll.interval.ms` — processing-progress detection - Default **300000 ms (5 min)**. - Measures time between successive `poll()` calls on the application thread. - If exceeded, the consumer concludes the application is stuck and **proactively sends LeaveGroup**, leaving the group so its partitions are reassigned. The background thread also stops heartbeating in that state. - This is your **stuck-processing detector**. ## How they interact - Background thread keeps `session.timeout.ms` happy → a slow-but-progressing consumer is no longer falsely evicted for liveness reasons. - But if the app stops calling poll() for longer than `max.poll.interval.ms`, the consumer leaves voluntarily — slow processing is still caught, just by the *right* timer. - Practical tuning: if you process large batches, raise `max.poll.interval.ms` and/or lower `max.poll.records` so each poll cycle finishes in time — **do not** touch `session.timeout.ms` for that purpose. ## Edge cases - A `CommitFailedException` saying the consumer 'is no longer part of the group' usually means `max.poll.interval.ms` was exceeded, not a heartbeat/session issue — a frequent diagnostic confusion. - Offloading processing to another thread to keep poll() cadence can keep the consumer in the group but risks committing offsets for unprocessed records; pause/resume partitions to manage this. - A GC pause that exceeds `session.timeout.ms` still evicts via the heartbeat path, independent of poll().

  • A consumer logs 'cannot commit because it is no longer part of the group.' Which timer most likely fired?
    max.poll.interval.ms — processing took longer than the poll interval, so the consumer left the group and its partitions were reassigned, invalidating the commit.
  • You process huge batches slowly. Which configs do you adjust, and which do you leave alone?
    Raise max.poll.interval.ms and/or lower max.poll.records. Leave session.timeout.ms alone — it's for crash detection, and the background thread already keeps it satisfied.

saying these in an interview costs you the question

  • Saying you fix slow processing by raising session.timeout.ms — that's the pre-KIP-62 anti-pattern.
  • Claiming heartbeats still only go out inside poll() in modern clients.
  • Conflating the two timers — they have different triggers (background thread vs poll cadence).

context