skip to content

Max Poll Interval and Long-Processing Pitfalls

Why slow processing gets a consumer evicted from its group, and the pause/resume and batch-size fixes. One of the most common real-world Kafka bugs, so interviewers like the debugging story.

part ofApache Kafkaoverview, primer and where to startread it →
on this pageshow

questions

5

What is max.poll.interval.ms, and what happens if your consumer takes longer than this value to process a batch of records?

level: juniorimportance: must knowfreq 78%

answer

  1. time between two poll() calls
  2. default 5 minutes (300000 ms)
  3. livelock — heartbeat alive, processing stuck
  4. eviction → rebalance → CommitFailedException
  5. single-threaded poll loop

basics

~20 s

max.poll.interval.ms is the maximum time allowed between two poll() calls. If processing a batch takes longer, Kafka assumes the consumer is stuck, kicks it out of the group, and rebalances its partitions to other members.

solid answer

~50 s

max.poll.interval.ms (default 300000 ms = 5 minutes) bounds the time between successive poll() calls on a consumer. The KafkaConsumer's poll loop is single-threaded: you poll, then process, then poll again. If processing the returned batch takes longer than this interval, the consumer fails to call poll() in time. The group coordinator then treats the member as failed (a 'livelock' — heartbeats may still flow on the background thread, but the application isn't making progress), evicts it from the group, and triggers a rebalance to reassign its partitions. When the slow member finally calls poll() or commitSync(), it gets a CommitFailedException / RebalanceInProgress error because it no longer owns those partitions, and any commit it attempts is rejected. The fix is to make processing faster, raise the interval, lower max.poll.records, or move work off the poll thread.

go deeper

for a junior

Know the one-line definition: max time between poll() calls, default 5 min, exceed it and you get kicked from the group.

for a middle

Explain the single-threaded poll loop and why slow processing starves poll(), and that commit then fails.

for a senior

Distinguish heartbeat liveness from poll() liveness, explain livelock, and name CommitFailedException plus the rebalance-thrash failure mode.

for a principal

Reason about how this config interacts with batch sizing, downstream latency tails, and overall group stability; design processing so the poll thread is never the bottleneck.

## The poll loop A Kafka consumer is fundamentally a single-threaded loop run by your application: ``` while (running) { records = consumer.poll(timeout); // fetch a batch process(records); // your business logic consumer.commitSync(); // record progress } ``` The consumer does NOT process records on a background thread. It hands you a batch and waits for you to come back and call `poll()` again. So the time between two `poll()` calls is essentially the time it takes to process one batch. ## Two different liveness signals Kafka distinguishes two kinds of 'is this consumer alive?': 1. **Heartbeats** — sent by a *background* thread on a `session.timeout.ms` / `heartbeat.interval.ms` cadence. These prove the consumer's network connection and JVM are alive. 2. **poll() liveness** — proven only when your application calls `poll()` again, governed by `max.poll.interval.ms` (default 5 minutes). This proves the application is actually *making progress*, not just breathing. The second signal exists because of **livelock**: a consumer whose background heartbeat thread keeps sending heartbeats, but whose application thread is wedged (stuck in a slow DB call, an infinite loop, a huge batch). Without `max.poll.interval.ms`, such a zombie would hold its partitions forever and no one would consume them. ## What eviction looks like When you exceed `max.poll.interval.ms`: - The group coordinator marks the member as failed and reassigns its partitions to other members (a rebalance). - When your slow code finally finishes and calls `poll()` or `commitSync()`, the consumer discovers it's been kicked out. `commitSync()` throws **CommitFailedException** ("Commit cannot be completed since the group has already rebalanced and assigned the partitions to another member"). - The records you just processed may be re-delivered to whichever member now owns the partition — duplicate processing. - The evicted consumer rejoins the group on its next poll, triggering *another* rebalance. If processing is chronically slow, you get a thrashing loop of evict → rejoin → evict. ## Edge cases - This is separate from `session.timeout.ms`. You can have a healthy heartbeat but still be evicted for slow processing. - A single oversized record (e.g. a 1 MB message that triggers a 30-second downstream call) can blow the interval even if average throughput is fine. - Static group membership (`group.instance.id`) does NOT exempt you from `max.poll.interval.ms` — it only suppresses the rebalance on a *clean* leave/rejoin within `session.timeout.ms`. ## Fixes (in order of preference) 1. Make processing faster. 2. Reduce `max.poll.records` so each batch is smaller and the loop comes back to `poll()` sooner. 3. Increase `max.poll.interval.ms` if processing is legitimately long. 4. Offload processing to a worker pool and use `pause()`/`resume()` to keep the poll thread responsive.

  • What is the default value of max.poll.interval.ms?
    300000 ms, i.e. 5 minutes.
  • Why does Kafka need both heartbeats and max.poll.interval.ms?
    Heartbeats prove the consumer process/connection is alive (sent on a background thread), but a consumer can heartbeat fine while its application thread is wedged in slow processing — a livelock. max.poll.interval.ms proves actual progress by requiring poll() to be called again.

saying these in an interview costs you the question

  • Saying heartbeats and max.poll.interval.ms are the same thing — they are separate liveness signals on separate threads.
  • Claiming the consumer processes records on a background thread — the poll loop is single-threaded; you process on the calling thread.
  • Thinking a healthy heartbeat protects you from eviction for slow processing — it does not.

context

open as a page

Explain the pause()/resume() backpressure pattern for offloading long-running record processing to worker threads while keeping the consumer in its group.

level: seniorimportance: must knowfreq 56%

basics

~20 s

Hand records to a worker pool for slow processing. To avoid fetching more than you can handle, call consumer.pause() on the assigned partitions and keep calling poll() (which now returns nothing but proves liveness). When workers catch up, call resume(). Commit only completed offsets.

open as a page

How does max.poll.interval.ms relate to session.timeout.ms and heartbeat.interval.ms? Why were poll liveness and heartbeat liveness decoupled?

level: middleimportance: should knowfreq 64%

basics

~10 s

session.timeout.ms / heartbeat.interval.ms detect a dead consumer via a background heartbeat thread. max.poll.interval.ms detects a live-but-stuck consumer via the poll() call. They were split (KIP-62) so slow processing doesn't get confused with a crash.

open as a page

A consumer is being evicted because each batch takes too long to process. How does reducing max.poll.records help, and what are the tradeoffs?

level: middleimportance: should knowfreq 58%

basics

~20 s

max.poll.records caps how many records each poll() returns. Lowering it means smaller batches that process faster, so poll() is called again sooner and you stay under max.poll.interval.ms. The tradeoff is more poll round-trips and potentially lower throughput.

open as a page

As a principal engineer, how would you decide between raising max.poll.interval.ms, lowering max.poll.records, and offloading processing for a consumer with slow, variable per-record processing? What are the systemic risks of just cranking the interval up?

level: principalimportance: should knowfreq 40%

basics

~20 s

Match the lever to the cost model: lower max.poll.records if cost scales with count; offload with pause()/resume() if individual records are slow or variable; raise max.poll.interval.ms only for legitimately long, bounded work. Cranking the interval high delays detection of truly stuck consumers, slowing partition reassignment and recovery.

open as a page