skip to content

A consumer group has growing lag but CPU on the instances is low and the network isn't saturated. Walk through how you'd diagnose and which fetch/parallelism configs you'd suspect.

level: seniorimportance: should knowfreq 45%

answer

  1. lag up + CPU idle = blocked or starved, not compute-bound
  2. decompose per-partition lag first (skew?)
  3. poll-idle-ratio high -> starved fetch or blocking I/O
  4. check instances vs partitions (ceiling)
  5. rebalance death spiral via max.poll.interval.ms

basics

~20 s

Low CPU + growing lag usually means the consumers are blocked waiting, not working: a slow downstream call, partition skew (one hot partition), too few partitions to parallelize, or fetches that are too small/infrequent. Check per-partition lag, then look at max.poll.records, fetch sizes, and partition assignment.

solid answer

~50 s

Growing lag with idle CPU points to blocking or under-utilization, not raw compute limits. First, break lag down per partition (kafka-consumer-groups.sh --describe): if one partition lags and the rest are fine, it's key skew or a hot partition — no consumer config fixes that, you need a better partitioning key. If lag is even but consumers are slow, profile where the poll loop spends time: a synchronous downstream call (DB/HTTP) blocks the thread, so adding consumers up to partition count or parallelizing in-process helps. Check whether you're partition-bound (instances == partitions already). On the fetch side, tiny fetch.max.bytes / max.partition.fetch.bytes or low max.poll.records can throttle how much each poll pulls, adding round-trip overhead. Also confirm you're not rebalancing repeatedly (each rebalance pauses consumption) due to max.poll.interval.ms being exceeded. Metrics: records-lag-max, fetch-latency, poll-idle-ratio, and rebalance rate.

go deeper

for a junior

Know to check consumer lag with kafka-consumer-groups.sh and that low CPU doesn't mean healthy.

for a middle

Decompose lag per partition and connect symptoms to max.poll.records / fetch sizes / rebalances.

for a senior

Run a structured diagnosis using metrics (poll-idle-ratio, fetch-latency, rebalance-rate) and pick the right lever among skew, parallelism, fetch sizing.

for a principal

Establish lag SLOs, dashboards, and runbooks; teach teams to avoid the add-more-consumers anti-pattern and design partitioning for skew resilience.

## What 'low CPU + rising lag' tells you Lag = (log end offset) - (committed consumer offset). If lag grows while the consumer CPU is idle, the consumers are **not compute-bound** — they're spending time **blocked** (waiting on I/O), **idle** (not enough data per fetch), or **not assigned** the work (skew / under-parallelized). Diagnosis is about finding which. ## Step 1 — Decompose lag per partition Run `kafka-consumer-groups.sh --describe --group <g>` to see per-partition lag. - **One/few partitions lag, rest fine** -> **key skew / hot partition**. A single partition is processed serially by one consumer; if most traffic hashes to one key/partition, that partition bottlenecks no matter how many consumers you add. Fix: better partitioning key or repartition. No fetch config solves this. - **All partitions lag evenly** -> systemic slowness; continue. ## Step 2 — Is the poll loop blocking? Profile or add timing around processing. The classic cause: each record triggers a **synchronous downstream call** (database write, REST call). The poll thread waits on network I/O -> CPU idle but throughput capped by downstream latency. - If so, scale consumers up to partition count, or introduce **in-process parallelism** (worker pool) so multiple records process concurrently. Batch the downstream sink if possible. ## Step 3 — Are you partition-bound? Check instances vs partitions. If instances == partitions, you've hit the parallelism ceiling; adding more does nothing. Options: repartition or in-process parallelism (see Step 2). ## Step 4 — Are fetches starved? If each poll() returns little data: - **max.poll.records** too low -> many poll() calls with per-call overhead. - **fetch.max.bytes / max.partition.fetch.bytes** too small -> small responses, more round trips. - **fetch.min.bytes** high with **fetch.max.wait.ms** high -> consumer waits, throughput on bursty topics can suffer. For catch-up throughput, raise max.poll.records and the fetch byte limits so each round trip moves more data. ## Step 5 — Is it rebalancing? Frequent rebalances stop consumption fleet-wide each time. Causes: - Processing a batch exceeds **max.poll.interval.ms** -> member evicted -> rebalance -> CommitFailedException -> reprocessing -> more lag (a death spiral). Fix: lower max.poll.records or raise the interval; use CooperativeStickyAssignor and static membership to reduce disruption. - Check the **rebalance rate** metric and consumer logs for 'Attempt to heartbeat failed' / 'rebalance'. ## Key metrics to pull - **records-lag-max** (per consumer): max lag across owned partitions. - **fetch-latency-avg/max**: time the broker takes to answer fetches (high = broker-side or network). - **poll-idle-ratio-avg**: fraction of time the consumer is idle inside poll() waiting for data (high = starved fetches or low traffic; low = processing-bound). - **rebalance-rate-per-hour** / **last-rebalance-seconds-ago**: rebalance churn. - **commit-latency**: slow commits can stall the loop. ## Putting it together Low CPU + even lag + high poll-idle-ratio -> fetch starvation or downstream blocking; raise fetch sizes or parallelize. Low CPU + skewed per-partition lag -> repartition/key fix. Frequent rebalances -> poll-interval tuning. The anti-pattern is blindly adding consumer instances when you're already partition-bound or downstream-bound.

  • Which JMX/client metric best distinguishes 'consumer is starved waiting for data' from 'consumer is busy processing'?
    poll-idle-ratio-avg: it measures the fraction of time the consumer spends idle inside poll() waiting on the broker. A high ratio means the consumer is waiting for data (fetch starvation or low traffic); a low ratio with rising lag means it's processing-bound and you should parallelize or speed up processing.
  • You find one partition has 10x the lag of others. Why won't adding consumers help, and what will?
    A partition is consumed serially by exactly one member, so its throughput is fixed regardless of consumer count — adding consumers leaves them idle. The fix is addressing key skew: choose a partitioning key with better cardinality/distribution, or repartition so the hot key's traffic spreads across partitions.

saying these in an interview costs you the question

  • Reflexively adding consumer instances when already partition-bound or downstream-bound.
  • Ignoring per-partition lag breakdown and missing key skew.
  • Assuming low CPU means the consumer is fine (it usually means it's blocked or starved).
  • Overlooking a rebalance loop driven by max.poll.interval.ms as the lag source.

context