How do you observe the backlog and throughput of a ThreadPoolExecutor using its queue and getCompletedTaskCount()/getTaskCount()?
answer
- getQueue().size() = backlog (watch the trend)
- completedTaskCount delta / time = throughput
- taskCount = scheduled total (done+active+queued)
- invariant: task = completed + active + queued
- getQueue() is live: read-only, don't mutate
basics
~10 sCall getQueue().size() to see how many tasks are waiting (the backlog), getCompletedTaskCount() to see how many have finished, and getTaskCount() for the total ever scheduled. A growing queue means the pool can't keep up.
solid answer
~40 sBacklog is read from the work queue via getQueue().size(): a steadily growing queue means arrival rate exceeds the pool's service rate, the classic saturation symptom. Throughput is derived from getCompletedTaskCount(), an approximate running total of finished tasks; sampling it over an interval (delta divided by time) gives tasks/second. getTaskCount() is the approximate total ever scheduled (completed + active + queued). The useful invariant is taskCount approximately equals completedTaskCount + activeCount + queue.size(). All these counters are approximations because the pool isn't fully locked while you read them, and getQueue() returns the live queue, so its size can change as you look. Treat getQueue() as read-only for monitoring; mutating it (e.g. draining) bypasses executor bookkeeping. Together with activeCount/poolSize, the queue trend and the completed-task rate are the core throughput and backlog signals.
code
java · 10 linesThreadPoolExecutor pool = ...;
long c1 = pool.getCompletedTaskCount();
Thread.sleep(1000);
long c2 = pool.getCompletedTaskCount();
double throughputPerSec = (c2 - c1) / 1.0; // sampled over 1s
int backlog = pool.getQueue().size(); // tasks waiting
long scheduled = pool.getTaskCount(); // ever scheduled
// sanity (approximate): scheduled ~= completed + active + queued
System.out.printf("throughput=%.1f/s backlog=%d scheduled=%d%n",
throughputPerSec, backlog, scheduled);go deeper
Knows getQueue().size() shows waiting tasks and getCompletedTaskCount() shows how many finished.
Computes throughput as a sampled delta of completedTaskCount over time and reads the queue trend as the saturation signal; states the task = completed + active + queued invariant.
Explains why the counters are approximate, why getQueue() must be treated as read-only, and how an unbounded queue converts saturation into memory growth that only queue monitoring catches.
Connects backlog/throughput signals to capacity planning and Little's Law style reasoning (queue length ~ arrival rate x wait time), and designs alerting on queue-growth slope and throughput collapse rather than single thresholds.
## Setup: queue and counters A **ThreadPoolExecutor** holds tasks it cannot run immediately in an internal **work queue** (a `BlockingQueue<Runnable>` you pass in, e.g. `LinkedBlockingQueue` or `ArrayBlockingQueue`). When all worker threads are busy and the pool is at its max, incoming tasks wait in this queue. Watching the queue and the lifetime task counters is how you measure **backlog** (work waiting) and **throughput** (work completing). ## Backlog: the queue - **`getQueue()`** returns the *live* queue instance. `getQueue().size()` is the number of tasks currently waiting (not yet started). - A **single reading** is just a snapshot; the signal that matters is the **trend**. If `queue.size()` keeps climbing, the **arrival rate** (tasks submitted per second) exceeds the **service rate** (tasks the pool can finish per second). This is the textbook saturation symptom and usually precedes rejections (when a bounded queue fills, the RejectedExecutionHandler kicks in). - For an **unbounded** queue (default `LinkedBlockingQueue` with no capacity) the queue size can grow without limit, masking saturation as memory growth instead of rejection, which is exactly why monitoring queue size matters. ## Throughput and totals: the task counters - **`getCompletedTaskCount()`** is an *approximate* running total of tasks that have **finished executing** since the pool started. It only increases. - **`getTaskCount()`** is the *approximate* total number of tasks that have **ever been scheduled** (completed + currently executing + currently queued). - **Throughput** is a *rate*, so you must **sample over time**: record `getCompletedTaskCount()` at t1 and t2, then `(c2 - c1) / (t2 - t1)` is tasks per second. A single absolute number tells you cumulative work, not current speed. ## The bookkeeping invariant At any instant, approximately: ``` getTaskCount() approx= getCompletedTaskCount() + getActiveCount() + getQueue().size() ``` That is: everything ever scheduled is either done, running, or waiting. It is approximate because the pool is not globally frozen while you read the parts, so the four reads are not a single atomic snapshot. ## Why everything is 'approximate' These accessors acquire the executor's internal lock only briefly (or not at all for the queue) and tasks transition states concurrently. The values are designed for **monitoring**, where small inconsistencies are acceptable, not for correctness decisions. ## Don't mutate the live queue `getQueue()` exposes the *actual* queue. Reading `size()`, iterating, or peeking is fine, but **removing/adding** elements directly bypasses the executor's accounting and can corrupt its view of pending work. To cancel queued work, use `remove(Runnable)` on the executor or `purge()`, not raw queue mutation. ## Putting it together The practical health read combines: queue size trend (backlog), completed-task rate (throughput), and active/pool size (utilization). Rising queue + flat completion rate = the pool is the bottleneck; flat queue + rising completion rate = healthy keep-up.
- Why measure throughput from getCompletedTaskCount() over an interval instead of reading it once?getCompletedTaskCount() is a cumulative lifetime total, so a single read tells you only how much total work has finished. Throughput is a rate; you need the delta between two timestamps divided by the elapsed time to get tasks per second.
- What's the danger of an unbounded queue when monitoring?With an unbounded LinkedBlockingQueue, saturation never triggers rejection; instead the queue (and heap) grows unboundedly. Monitoring queue size is the only early warning, otherwise the symptom appears as memory pressure or OutOfMemoryError.
saying these in an interview costs you the question
- Reading getCompletedTaskCount() once and calling that 'throughput' (it's a cumulative total, not a rate)
- Mutating the queue returned by getQueue() (bypasses executor bookkeeping; use remove/purge)
- Treating taskCount = completed + active + queued as exact rather than approximate
- Assuming an unbounded queue protects you, it just turns saturation into unbounded memory growth