Explain the relationship between heartbeat.interval.ms and session.timeout.ms and how you would tune them.
answer
- interval <= timeout / 3
- interval = how often; timeout = how patient
- defaults 3000 / 45000
- low timeout = fast detect but flappy
- bounded by broker min/max
basics
~20 sheartbeat.interval.ms is how often the consumer sends heartbeats; session.timeout.ms is how long the coordinator waits without one before declaring the consumer dead. The interval must be well below the timeout — the rule of thumb is heartbeat.interval.ms <= session.timeout.ms / 3.
solid answer
~40 sThese two settings work as a pair. heartbeat.interval.ms (default 3000 ms) controls how frequently the consumer's background thread sends heartbeats. session.timeout.ms (default 45000 ms in recent releases) is the coordinator-side window: if no heartbeat arrives within it, the member is evicted and a rebalance starts. The interval must be small enough that several heartbeats fit inside one session window, so a couple of dropped or delayed heartbeats don't cause a false eviction — Kafka recommends heartbeat.interval.ms be no more than one-third of session.timeout.ms. Tuning trade-off: a lower session.timeout.ms detects real failures faster but risks spurious rebalances under GC pauses or network jitter; a higher value is more tolerant but slows failure detection. session.timeout.ms must also fall within the broker's group.min.session.timeout.ms / group.max.session.timeout.ms bounds, or the join is rejected.
go deeper
Know interval = send frequency, timeout = patience window, and interval must be smaller.
State the <= timeout/3 rule, the defaults, and the fast-detection vs false-eviction trade-off.
Tie tuning to GC pauses/network jitter and keep it separate from max.poll.interval.ms; mention broker bounds.
Set fleet-wide policy balancing rebalance churn vs recovery time, accounting for GC profiles and SLAs.
## The two knobs - **`heartbeat.interval.ms`** (client-side, default 3000 ms): how often the consumer's background heartbeat thread sends a heartbeat to the group coordinator. - **`session.timeout.ms`** (client-side, but bounded by the broker, default 45000 ms in recent Kafka; was 10000 ms historically): the maximum time the coordinator will wait for a heartbeat before declaring the member dead and triggering a rebalance. ## Why the interval must be a fraction of the timeout Networks drop packets and JVMs pause for GC. If you sent only one heartbeat per session window, a single hiccup would evict a healthy consumer. By sending heartbeats several times per window, you tolerate transient losses: the member survives as long as *at least one* heartbeat lands inside the window. Kafka's documented guidance is: ``` heartbeat.interval.ms <= session.timeout.ms / 3 ``` With defaults (3000 vs 45000) you get ~15 heartbeats per window — very tolerant. A common tighter pairing is 3000 / 10000 (~3 heartbeats per window). ## Tuning trade-offs - **Lower `session.timeout.ms`** → faster detection of genuinely dead consumers (partitions reassigned sooner, less consumer lag during failures) but **more false-positive evictions** and rebalances under GC pauses, CPU starvation, or network jitter. - **Higher `session.timeout.ms`** → robust against transient stalls, but a truly dead consumer's partitions sit unprocessed longer. - **`heartbeat.interval.ms`** mostly affects responsiveness to rebalance signals (the consumer learns about an in-progress rebalance via heartbeat responses) and the redundancy described above. Lower = faster rebalance reaction + more RPC overhead. ## Broker bounds `session.timeout.ms` is validated against the broker configs `group.min.session.timeout.ms` (default 6000 ms) and `group.max.session.timeout.ms` (default 1800000 ms = 30 min). A consumer requesting a value outside that range gets an `InvalidSessionTimeout` error and cannot join. ## Separation from processing time Crucially, neither setting governs how long your application may take to process a batch — that is `max.poll.interval.ms`. Since KIP-62 the heartbeat thread keeps liveness alive even when processing is slow, so you tune session.timeout.ms for *crash/network* detection and max.poll.interval.ms for *processing* tolerance, independently. ## Edge cases - Setting `heartbeat.interval.ms` too close to `session.timeout.ms` defeats redundancy and causes flapping. - Setting `session.timeout.ms` below `group.min.session.timeout.ms` is silently impossible — the broker rejects it.
- What's the recommended ratio between the two?heartbeat.interval.ms should be at most session.timeout.ms / 3, so multiple heartbeats fit in one session window and a single drop won't cause eviction.
- If processing each batch takes 5 minutes, should you raise session.timeout.ms?No. Since KIP-62 the background thread keeps heartbeating during processing. Raise max.poll.interval.ms instead; session.timeout.ms is for crash/network detection.
saying these in an interview costs you the question
- Raising session.timeout.ms to accommodate slow processing — that's max.poll.interval.ms's job since KIP-62.
- Setting heartbeat.interval.ms equal to or near session.timeout.ms (no redundancy).
- Forgetting that the broker min/max bounds can reject the chosen value.