skip to content

Requests handled by a worker pool have stopped completing and throughput has collapsed. How do you determine whether the pool is merely saturated by slow work or genuinely deadlocked, and why does the answer change what you do about it?

level: seniorimportance: should knowfreq 44%

answer

  1. is the completed counter moving?
  2. two stack samples, diff them
  3. idle CPU = deadlock, hot CPU = livelock
  4. drop load: saturation drains, deadlock does not
  5. oldest in-flight task age as the alert

basics

~20 s

Check whether anything is completing. Saturation still finishes tasks - slowly - and drains when load falls. Deadlock finishes nothing: sample worker stacks twice; identical stacks on the same tasks, with a flat completed-task counter, means stuck. Saturation needs load control; deadlock needs a structural fix plus restart.

solid answer

~50 s

Ask one question first: is the completed-task counter still advancing? Saturation is a rate problem - work completes, just too slowly, latency and queue depth are high but finite, and the pool recovers when arrivals drop or the dependency heals. Deadlock is a liveness problem - completions are exactly zero and stay zero regardless of load. The cheap discriminator is two stack samples of the worker threads a few seconds apart. Saturated workers move: different tasks, different frames. Deadlocked workers are frozen in the same frame on the same task, and the frame is telling - waiting on a result versus reading a socket versus acquiring a lock. A queue that is deep and never shrinks with zero active progress confirms it. The response differs completely: saturation is handled operationally - shed load, add timeouts, add capacity, cut service time. Deadlock cannot be tuned away; the process must be restarted to recover and the dependency structure changed so the cycle cannot form.

code

text · 5 lines
text
completions  queue     CPU    recovers when load drops
saturation       low, >0      growing   var    yes
deadlock         0            frozen    idle   no
livelock/thrash  ~0           growing   high   sometimes
slow dependency  low, >0      growing   idle   yes

go deeper

for a junior

Know the basic tell: if nothing at all is completing, it is stuck; if things complete slowly, it is overloaded.

for a middle

Add the method - two stack samples, completion counter, queue depth - and what each frame type implies.

for a senior

Own the full triage: distinguish deadlock, livelock, saturation and slow-dependency signatures, capture evidence before restarting, and pick the matching remedy.

for a principal

Push it upstream into design and observability: per-pool progress metrics and oldest-in-flight age, timeouts as a containment default, and an architectural invariant that makes the stuck case impossible rather than detectable.

## Three failure shapes that look the same from outside From a dashboard, a stalled pool, a saturated pool, and a livelocked pool all read as 'requests time out'. They need different responses, so distinguishing them is the whole job. - **Saturation:** arrivals exceed service capacity, or service time grew. Workers are busy doing real work or blocked on a real dependency. Progress continues at a reduced rate. - **Deadlock:** a set of workers waits on conditions only they could satisfy - most often parents blocked on results from their own pool, or two tasks acquiring two locks in opposite orders. Progress is exactly zero and permanent. - **Livelock or thrash:** threads are running (CPU is busy) but useful completions are near zero - retry loops, contention collapse, spinning on a contended structure. Distinguished from deadlock by CPU being high rather than idle. ## The diagnostic ladder **1. Completion counter over a window.** The single most informative signal: tasks completed per second. Non-zero and steady, however small, means the system is alive - saturated or dependency-slow. Exactly zero across a minute while work is queued means stuck. **2. Two stack samples.** Capture the state of all worker threads twice, several seconds apart, and diff them. Same task identity plus same frame equals no progress on that worker. All workers frozen equals deadlock. Some frozen, some rotating equals partial starvation - often a subset of workers is stuck while the rest still serve. **3. Read the frames.** Where a worker is parked names the cause: waiting for a result of a submitted task means the nested-submission pattern; waiting on a lock means a lock-ordering cycle; waiting in a socket read means a slow dependency with no read timeout, which is saturation dressed up as a hang. **4. Queue depth trend.** Growing with zero completions equals deadlock. Growing with non-zero completions equals overload. Flat and non-empty with zero completions and idle CPU is the strongest deadlock signature. **5. CPU.** Deadlocked pools burn no CPU; livelocked ones burn plenty. This one number separates 'everyone is waiting' from 'everyone is spinning'. **6. Load experiment, if you can afford it.** Cut arrivals to near zero. Saturation drains and recovers; deadlock stays exactly where it was. This is the definitive test, because the defining property of a deadlock is that it is not load-dependent once entered. ## Why the distinction dictates the response For **saturation** the levers are load and capacity: shed or throttle arrivals, enforce deadlines so abandoned work stops occupying workers, cap concurrency to the slow dependency so it cannot consume the whole pool, add workers if the bottleneck is genuinely wait-bound, or reduce service time. The system heals itself once the imbalance is corrected, and no restart is needed. For **deadlock** none of those help. Adding workers postpones recurrence; shedding load does not release the held threads; the process must be restarted to recover, and until the structure changes it will happen again at the next burst. The fix is structural: remove same-pool blocking, isolate pools by dependency layer, impose a global lock order, or make waiting non-blocking. Timeouts on waits are worth adding as a safety net - they convert a permanent hang into a bounded failure and let the pool recover on its own - but a timeout is containment, not a cure. ## Making it observable in advance Instrument pools with active worker count, queue depth, completed count, and task age of the oldest in-flight task. A rising 'oldest in-flight task age' with a flat completed count is the earliest machine-readable signature of a stuck pool, and it is the alert worth having, because by the time latency alerts fire the pool is already unrecoverable.

  • You find every worker parked waiting for the result of another task. What do you check next to confirm the diagnosis?
    Confirm the awaited tasks are queued on the same pool rather than executing elsewhere - if the pending work sits in that pool's own queue with no free worker, the cycle is closed. Then check whether the number of blocked parents equals the pool size, which is the condition for the permanent form. Finally check the submission site to see whether the parent-child relationship is intrinsic or accidental, since that decides between layered pools and asynchronous composition.
  • Are timeouts on every wait a sufficient answer to this class of bug?
    They are a valuable safety net but not a fix. A timeout bounds how long a worker is held, so the pool eventually recovers on its own instead of needing a restart, and the failure becomes a visible error rather than a silent hang. But the work still fails, the pool still stalls for the timeout duration under load, and if callers retry you regain the same state immediately. Use timeouts for containment and change the dependency structure for the cure.

saying these in an interview costs you the question

  • Restarting without capturing stacks, destroying the only evidence
  • Treating a permanent stall as overload and adding workers
  • Assuming high CPU means the pool is working, when it may be a retry livelock
  • Reading one stack sample and calling a slow task 'stuck'
  • Concluding deadlock without checking whether completions are truly zero

context