skip to content

How does max.poll.interval.ms relate to session.timeout.ms and heartbeat.interval.ms? Why were poll liveness and heartbeat liveness decoupled?

level: middleimportance: should knowfreq 64%

answer

  1. KIP-62 split poll liveness from heartbeat liveness
  2. background heartbeat thread vs single-threaded poll
  3. heartbeat ≈ session.timeout / 3
  4. session 45s, hb 3s, poll-interval 5min (defaults)
  5. expiry sends proactive LeaveGroup

basics

~10 s

session.timeout.ms / heartbeat.interval.ms detect a dead consumer via a background heartbeat thread. max.poll.interval.ms detects a live-but-stuck consumer via the poll() call. They were split (KIP-62) so slow processing doesn't get confused with a crash.

solid answer

~40 s

Before KIP-62 (Kafka 0.10.1), heartbeats were sent only inside poll(), so session.timeout.ms had to cover both connection liveness AND processing time — forcing you to set a huge session timeout for slow consumers, which made real crash detection slow. KIP-62 moved heartbeats to a dedicated background thread and introduced max.poll.interval.ms. Now: heartbeat.interval.ms (default 3 s) is the cadence of background heartbeats; session.timeout.ms (default 45 s) is how long the coordinator waits for a heartbeat before declaring the member dead (set heartbeat.interval to roughly 1/3 of it); and max.poll.interval.ms (default 5 min) independently bounds processing time. A crash is detected fast (session timeout), while slow processing gets its own generous budget without blocking crash detection. The two evictions report differently — heartbeat failure vs. poll-interval expiry — and both trigger a rebalance.

go deeper

for a junior

Know there are two different timers: one for 'is it alive' (heartbeat/session) and one for 'is it making progress' (poll interval).

for a middle

Explain the three configs, their defaults, and the heartbeat ≈ session/3 rule.

for a senior

Explain KIP-62's motivation (coupling forced huge session timeouts), the background heartbeat thread, and the proactive LeaveGroup on interval expiry.

for a principal

Tune all three for fast crash failover plus long-processing tolerance, account for GC pauses and tail latency, and know static-membership interactions.

## Three configs, two concerns | Config | Default | Thread | Detects | |---|---|---|---| | `heartbeat.interval.ms` | 3000 ms | background | how *often* heartbeats are sent | | `session.timeout.ms` | 45000 ms | (coordinator side) | a *dead/disconnected* consumer | | `max.poll.interval.ms` | 300000 ms | application | a *live-but-stuck* consumer | ## Before KIP-62: one timer for two jobs In early Kafka, the consumer only sent heartbeats *inside* `poll()`. That coupled liveness to processing: if your processing took 4 minutes, no heartbeat went out for 4 minutes, so `session.timeout.ms` had to be > 4 minutes. But a large session timeout means a genuinely crashed consumer takes minutes to be detected — slow failover. You couldn't have both fast crash detection and long processing. ## KIP-62 (Kafka 0.10.1): decouple them KIP-62 introduced a **background heartbeat thread** and the new `max.poll.interval.ms` config. Now there are two independent liveness signals: - **Heartbeat liveness** — the background thread heartbeats every `heartbeat.interval.ms`. If the coordinator sees no heartbeat for `session.timeout.ms`, the member is declared dead. This catches crashes, network partitions, long GC pauses, and `kill -9` quickly. - **Poll liveness** — proven by calling `poll()` again within `max.poll.interval.ms`. This catches livelock: the JVM and connection are fine (heartbeats flowing) but the application thread is wedged. When the interval expires, the background thread proactively sends a **LeaveGroup** request so the partitions are reassigned promptly. ## Tuning relationships - Rule of thumb: `heartbeat.interval.ms` ≈ `session.timeout.ms` / 3, so a few missed heartbeats are tolerated. - `session.timeout.ms` must fall within the broker's allowed range `group.min.session.timeout.ms` … `group.max.session.timeout.ms` (defaults 6 s … 30 min). The broker rejects out-of-range values. - `max.poll.interval.ms` is unrelated to those bounds and is set purely by your processing time. ## Why it matters operationally - Set `session.timeout.ms` small enough for fast failover but large enough to survive GC pauses. - Set `max.poll.interval.ms` to comfortably exceed your worst-case (p99/tail) batch processing time, not the average — one slow batch evicts you. - Static membership (`group.instance.id`, KIP-345) lets a consumer restart within `session.timeout.ms` without a rebalance, but does NOT extend `max.poll.interval.ms`.

  • Which KIP introduced this decoupling and the max.poll.interval.ms config?
    KIP-62, shipped in Kafka 0.10.1, which moved heartbeats to a background thread and added max.poll.interval.ms.
  • Should you tune max.poll.interval.ms to your average or worst-case processing time?
    Worst-case (tail/p99). A single slow batch that exceeds the interval evicts the consumer, so you must budget for the slowest plausible batch, not the average.

saying these in an interview costs you the question

  • Saying session.timeout.ms governs processing time — since KIP-62 it governs only heartbeat liveness.
  • Claiming heartbeats are sent inside poll() — that was the pre-0.10.1 behavior; they now run on a background thread.
  • Setting session.timeout.ms huge to accommodate slow processing — that is exactly the anti-pattern KIP-62 removed; raise max.poll.interval.ms instead.

context