skip to content

Explain the relationship between heartbeat.interval.ms and session.timeout.ms and how you would tune them.

level: middleimportance: must knowfreq 75%

answer

  1. interval <= timeout / 3
  2. interval = how often; timeout = how patient
  3. defaults 3000 / 45000
  4. low timeout = fast detect but flappy
  5. bounded by broker min/max

basics

~20 s

heartbeat.interval.ms is how often the consumer sends heartbeats; session.timeout.ms is how long the coordinator waits without one before declaring the consumer dead. The interval must be well below the timeout — the rule of thumb is heartbeat.interval.ms <= session.timeout.ms / 3.

solid answer

~40 s

These two settings work as a pair. heartbeat.interval.ms (default 3000 ms) controls how frequently the consumer's background thread sends heartbeats. session.timeout.ms (default 45000 ms in recent releases) is the coordinator-side window: if no heartbeat arrives within it, the member is evicted and a rebalance starts. The interval must be small enough that several heartbeats fit inside one session window, so a couple of dropped or delayed heartbeats don't cause a false eviction — Kafka recommends heartbeat.interval.ms be no more than one-third of session.timeout.ms. Tuning trade-off: a lower session.timeout.ms detects real failures faster but risks spurious rebalances under GC pauses or network jitter; a higher value is more tolerant but slows failure detection. session.timeout.ms must also fall within the broker's group.min.session.timeout.ms / group.max.session.timeout.ms bounds, or the join is rejected.

go deeper

for a junior

Know interval = send frequency, timeout = patience window, and interval must be smaller.

for a middle

State the <= timeout/3 rule, the defaults, and the fast-detection vs false-eviction trade-off.

for a senior

Tie tuning to GC pauses/network jitter and keep it separate from max.poll.interval.ms; mention broker bounds.

for a principal

Set fleet-wide policy balancing rebalance churn vs recovery time, accounting for GC profiles and SLAs.

## The two knobs - **`heartbeat.interval.ms`** (client-side, default 3000 ms): how often the consumer's background heartbeat thread sends a heartbeat to the group coordinator. - **`session.timeout.ms`** (client-side, but bounded by the broker, default 45000 ms in recent Kafka; was 10000 ms historically): the maximum time the coordinator will wait for a heartbeat before declaring the member dead and triggering a rebalance. ## Why the interval must be a fraction of the timeout Networks drop packets and JVMs pause for GC. If you sent only one heartbeat per session window, a single hiccup would evict a healthy consumer. By sending heartbeats several times per window, you tolerate transient losses: the member survives as long as *at least one* heartbeat lands inside the window. Kafka's documented guidance is: ``` heartbeat.interval.ms <= session.timeout.ms / 3 ``` With defaults (3000 vs 45000) you get ~15 heartbeats per window — very tolerant. A common tighter pairing is 3000 / 10000 (~3 heartbeats per window). ## Tuning trade-offs - **Lower `session.timeout.ms`** → faster detection of genuinely dead consumers (partitions reassigned sooner, less consumer lag during failures) but **more false-positive evictions** and rebalances under GC pauses, CPU starvation, or network jitter. - **Higher `session.timeout.ms`** → robust against transient stalls, but a truly dead consumer's partitions sit unprocessed longer. - **`heartbeat.interval.ms`** mostly affects responsiveness to rebalance signals (the consumer learns about an in-progress rebalance via heartbeat responses) and the redundancy described above. Lower = faster rebalance reaction + more RPC overhead. ## Broker bounds `session.timeout.ms` is validated against the broker configs `group.min.session.timeout.ms` (default 6000 ms) and `group.max.session.timeout.ms` (default 1800000 ms = 30 min). A consumer requesting a value outside that range gets an `InvalidSessionTimeout` error and cannot join. ## Separation from processing time Crucially, neither setting governs how long your application may take to process a batch — that is `max.poll.interval.ms`. Since KIP-62 the heartbeat thread keeps liveness alive even when processing is slow, so you tune session.timeout.ms for *crash/network* detection and max.poll.interval.ms for *processing* tolerance, independently. ## Edge cases - Setting `heartbeat.interval.ms` too close to `session.timeout.ms` defeats redundancy and causes flapping. - Setting `session.timeout.ms` below `group.min.session.timeout.ms` is silently impossible — the broker rejects it.

  • What's the recommended ratio between the two?
    heartbeat.interval.ms should be at most session.timeout.ms / 3, so multiple heartbeats fit in one session window and a single drop won't cause eviction.
  • If processing each batch takes 5 minutes, should you raise session.timeout.ms?
    No. Since KIP-62 the background thread keeps heartbeating during processing. Raise max.poll.interval.ms instead; session.timeout.ms is for crash/network detection.

saying these in an interview costs you the question

  • Raising session.timeout.ms to accommodate slow processing — that's max.poll.interval.ms's job since KIP-62.
  • Setting heartbeat.interval.ms equal to or near session.timeout.ms (no redundancy).
  • Forgetting that the broker min/max bounds can reject the chosen value.

context