A consumer group has growing lag but CPU on the instances is low and the network isn't saturated. Walk through how you'd diagnose and which fetch/parallelism configs you'd suspect.
answer
- lag up + CPU idle = blocked or starved, not compute-bound
- decompose per-partition lag first (skew?)
- poll-idle-ratio high -> starved fetch or blocking I/O
- check instances vs partitions (ceiling)
- rebalance death spiral via max.poll.interval.ms
basics
~20 sLow CPU + growing lag usually means the consumers are blocked waiting, not working: a slow downstream call, partition skew (one hot partition), too few partitions to parallelize, or fetches that are too small/infrequent. Check per-partition lag, then look at max.poll.records, fetch sizes, and partition assignment.
solid answer
~50 sGrowing lag with idle CPU points to blocking or under-utilization, not raw compute limits. First, break lag down per partition (kafka-consumer-groups.sh --describe): if one partition lags and the rest are fine, it's key skew or a hot partition — no consumer config fixes that, you need a better partitioning key. If lag is even but consumers are slow, profile where the poll loop spends time: a synchronous downstream call (DB/HTTP) blocks the thread, so adding consumers up to partition count or parallelizing in-process helps. Check whether you're partition-bound (instances == partitions already). On the fetch side, tiny fetch.max.bytes / max.partition.fetch.bytes or low max.poll.records can throttle how much each poll pulls, adding round-trip overhead. Also confirm you're not rebalancing repeatedly (each rebalance pauses consumption) due to max.poll.interval.ms being exceeded. Metrics: records-lag-max, fetch-latency, poll-idle-ratio, and rebalance rate.
go deeper
Know to check consumer lag with kafka-consumer-groups.sh and that low CPU doesn't mean healthy.
Decompose lag per partition and connect symptoms to max.poll.records / fetch sizes / rebalances.
Run a structured diagnosis using metrics (poll-idle-ratio, fetch-latency, rebalance-rate) and pick the right lever among skew, parallelism, fetch sizing.
Establish lag SLOs, dashboards, and runbooks; teach teams to avoid the add-more-consumers anti-pattern and design partitioning for skew resilience.
## What 'low CPU + rising lag' tells you Lag = (log end offset) - (committed consumer offset). If lag grows while the consumer CPU is idle, the consumers are **not compute-bound** — they're spending time **blocked** (waiting on I/O), **idle** (not enough data per fetch), or **not assigned** the work (skew / under-parallelized). Diagnosis is about finding which. ## Step 1 — Decompose lag per partition Run `kafka-consumer-groups.sh --describe --group <g>` to see per-partition lag. - **One/few partitions lag, rest fine** -> **key skew / hot partition**. A single partition is processed serially by one consumer; if most traffic hashes to one key/partition, that partition bottlenecks no matter how many consumers you add. Fix: better partitioning key or repartition. No fetch config solves this. - **All partitions lag evenly** -> systemic slowness; continue. ## Step 2 — Is the poll loop blocking? Profile or add timing around processing. The classic cause: each record triggers a **synchronous downstream call** (database write, REST call). The poll thread waits on network I/O -> CPU idle but throughput capped by downstream latency. - If so, scale consumers up to partition count, or introduce **in-process parallelism** (worker pool) so multiple records process concurrently. Batch the downstream sink if possible. ## Step 3 — Are you partition-bound? Check instances vs partitions. If instances == partitions, you've hit the parallelism ceiling; adding more does nothing. Options: repartition or in-process parallelism (see Step 2). ## Step 4 — Are fetches starved? If each poll() returns little data: - **max.poll.records** too low -> many poll() calls with per-call overhead. - **fetch.max.bytes / max.partition.fetch.bytes** too small -> small responses, more round trips. - **fetch.min.bytes** high with **fetch.max.wait.ms** high -> consumer waits, throughput on bursty topics can suffer. For catch-up throughput, raise max.poll.records and the fetch byte limits so each round trip moves more data. ## Step 5 — Is it rebalancing? Frequent rebalances stop consumption fleet-wide each time. Causes: - Processing a batch exceeds **max.poll.interval.ms** -> member evicted -> rebalance -> CommitFailedException -> reprocessing -> more lag (a death spiral). Fix: lower max.poll.records or raise the interval; use CooperativeStickyAssignor and static membership to reduce disruption. - Check the **rebalance rate** metric and consumer logs for 'Attempt to heartbeat failed' / 'rebalance'. ## Key metrics to pull - **records-lag-max** (per consumer): max lag across owned partitions. - **fetch-latency-avg/max**: time the broker takes to answer fetches (high = broker-side or network). - **poll-idle-ratio-avg**: fraction of time the consumer is idle inside poll() waiting for data (high = starved fetches or low traffic; low = processing-bound). - **rebalance-rate-per-hour** / **last-rebalance-seconds-ago**: rebalance churn. - **commit-latency**: slow commits can stall the loop. ## Putting it together Low CPU + even lag + high poll-idle-ratio -> fetch starvation or downstream blocking; raise fetch sizes or parallelize. Low CPU + skewed per-partition lag -> repartition/key fix. Frequent rebalances -> poll-interval tuning. The anti-pattern is blindly adding consumer instances when you're already partition-bound or downstream-bound.
- Which JMX/client metric best distinguishes 'consumer is starved waiting for data' from 'consumer is busy processing'?poll-idle-ratio-avg: it measures the fraction of time the consumer spends idle inside poll() waiting on the broker. A high ratio means the consumer is waiting for data (fetch starvation or low traffic); a low ratio with rising lag means it's processing-bound and you should parallelize or speed up processing.
- You find one partition has 10x the lag of others. Why won't adding consumers help, and what will?A partition is consumed serially by exactly one member, so its throughput is fixed regardless of consumer count — adding consumers leaves them idle. The fix is addressing key skew: choose a partitioning key with better cardinality/distribution, or repartition so the hot key's traffic spreads across partitions.
saying these in an interview costs you the question
- Reflexively adding consumer instances when already partition-bound or downstream-bound.
- Ignoring per-partition lag breakdown and missing key skew.
- Assuming low CPU means the consumer is fine (it usually means it's blocked or starved).
- Overlooking a rebalance loop driven by max.poll.interval.ms as the lag source.