A running service stops serving requests while its processor usage sits near zero. How do you use thread dumps — snapshots of every thread's stack and state — to work out what it is stuck on?
answer
- several dumps, seconds apart — compare for progress
- zero processor use = everyone waiting
- group by stack, count the crowd
- blocked → find the lock owner → what is the owner waiting on
- capture before restart; name your threads
basics
~20 sTake several dumps a few seconds apart and compare. Near-zero processor use means nobody is running, so look for threads blocked on a lock (and who owns it), waiting on a condition or a remote response, and for whole worker pools stuck in the same stack — that shared frame is the culprit.
solid answer
~60 sZero processor usage plus no progress means no thread is runnable, so the answer is in what they are all waiting for. I take three to five dumps ten seconds apart — one dump cannot distinguish stuck from merely slow, whereas identical stacks across dumps prove no progress. Then I group threads by stack rather than reading them one by one: in a stuck service the interesting signal is usually a hundred worker threads sharing one frame. From there I read the states. Threads blocked acquiring a lock normally name the owning thread, so I follow the chain to the owner and ask what *it* is waiting on. Threads parked on a condition, a latch or a queue tell me a producer never signalled. Threads waiting on a socket read point at a downstream dependency and, usually, a missing timeout. If every pool thread is in the same downstream call, the pool is exhausted and the real fault is elsewhere. Before restarting I capture the dumps — restarting destroys the evidence.
code
text · 13 linesthread "http-worker-17" state: BLOCKED on lock 0x7f2 owned by "http-worker-3"
at InventoryCache.read(...)
at OrderHandler.handle(...)
... 149 more threads with the identical stack ...
thread "http-worker-3" state: RUNNABLE (in socket read) holds lock 0x7f2
at SocketInput.read(...)
at PricingClient.call(...) <-- remote call inside a critical section
at InventoryCache.refresh(...)
Reading: 150 workers serialised behind one lock; the owner is
waiting on a remote service with no timeout. Fault is the
remote call inside the lock, not the cache logic.go deeper
Know what a thread dump is, that several spaced dumps are needed to prove nothing is progressing, and that near-zero processor use means threads are waiting rather than working.
Read the states confidently, follow blocked-thread to lock-owner chains, group by stack, and recognise pool exhaustion behind a slow dependency.
Correlate dumps with pool and dependency metrics, recognise starvation deadlock from nested submissions, and translate findings into remedies such as timeouts, bulkheads and shorter critical sections.
Insist on the preconditions that make dumps useful at all — named threads, automated capture on health-check failure, durable storage, a rule against restarting before capture — and on isolation design so one dependency cannot consume all capacity.
## What a thread dump gives you A thread dump is a point-in-time snapshot of every thread in the process: its identity, its scheduling state, its call stack, and — in most managed runtimes — which locks it holds and which one it is waiting to acquire. It is the cheapest and most universally available diagnostic for a hung service, and it can normally be requested from outside the process without a restart. The states you care about, in language-neutral terms: - **Runnable** — eligible to run. Note that a thread blocked in a socket read is often reported as runnable, because the wait happens in the kernel; the stack, not the state label, tells you the truth. - **Blocked on a monitor/lock** — waiting to enter a critical section someone else holds. This entry usually names the owner. - **Waiting / timed waiting** — parked on a condition, a latch, a queue, a future, or a sleep, until someone signals it. ## Method **1. Take several dumps, spaced.** Three to five, roughly five to ten seconds apart. A single dump cannot distinguish "stuck forever" from "busy right now": only the comparison can. If a thread has the identical stack in every dump, it made no progress; if stacks churn while throughput is zero, you are looking at a livelock or a spin, not a block. **2. Aggregate, do not read linearly.** Group threads by stack signature and count. A dump of 400 threads is unreadable one by one, but "312 threads share these top ten frames" is a diagnosis in one line. **3. Interpret near-zero processor use.** Nothing is executing, so the system is waiting. That narrows the world to three families: (a) a lock nobody will release, (b) a signal that never comes, (c) an external call that never returns. **4. Follow the ownership chain.** For blocked threads, find the lock owner. Then examine that owner: it may itself be waiting on a slow remote call while holding a lock the whole pool needs — the classic "one slow dependency freezes everything" shape. If the chain loops back on itself, you have a deadlock. **5. Look at the waiters' meaning.** Threads parked on a work queue with an empty queue are simply idle and healthy — do not misread an idle pool as a hang. Threads parked on a latch or a future indicate a producer/completer that died or never ran; check whether the thread that was supposed to complete it still exists (an uncaught exception may have killed it). **6. Check for pool exhaustion.** If every request-handling thread is inside the same downstream client call, the service is not broken — it is saturated by a dependency, and the fix is timeouts, bulkheads and circuit breaking rather than anything local. A related trap is a pool whose tasks block on results computed by the *same* pool; with all threads waiting, no thread remains to produce the result, which is a starvation deadlock even though no lock is involved. **7. Correlate with the rest.** Pool queue depth, in-flight request counts, downstream latency and connection-pool metrics turn a stack snapshot into a story. A thread dump tells you where threads are; metrics tell you when it started and what changed. ## Common shapes and what they mean - **All workers blocked on one lock; the owner in a socket read** → one slow dependency serialised behind a shared lock. Shorten the critical section, never call remote services while holding a lock, add a timeout. - **All workers in the same downstream client frame; no locks** → dependency-induced pool exhaustion; needs timeouts and isolation. - **A cycle of blocked threads** → deadlock; the runtime may even report it explicitly. - **Threads waiting on futures completed by their own pool** → starvation deadlock from nested submission. - **Everything idle on an empty queue while requests pile up elsewhere** → the problem is upstream: the acceptor, the connection pool, or a rate limiter. ## Practical cautions - **Capture before you restart.** A restart is the most common way to destroy the only evidence you had. - **Dumping is not free but is usually acceptable** — it typically pauses the process briefly. In an already-hung service that cost is irrelevant. - **Name your threads.** Meaningful pool and thread names turn an anonymous dump into a readable map; this is a cheap investment made long before the incident. - **Automate capture.** A watchdog that dumps when health checks fail, storing the files durably, means the next hang is diagnosed from evidence rather than from memory.
- The dump shows every worker thread inside the same downstream client call, with no locks involved. What has actually gone wrong and what is the fix?The service is suffering pool exhaustion caused by a slow or hung dependency: each request occupies a worker for the duration of the remote call, so once the dependency slows past the pool's capacity every thread is consumed and new requests queue forever. The local code is not the defect. Fixes are timeouts on every remote call, bulkheads or separate pools so one dependency cannot consume all capacity, circuit breaking to fail fast, and load shedding at the edge.
- You see many threads parked and waiting on a work queue that is empty. Is that a problem?By itself, no — an idle pool waiting for work is the healthy resting state, and mistaking it for a hang sends you in the wrong direction. It becomes a symptom only if work should be arriving and is not, which points upstream: the acceptor thread, a saturated connection pool, a rate limiter, or a producer that died. Check the queue arrival rate and the upstream component rather than the idle workers.
- Why take several dumps instead of one, and what does it mean if the stacks change but throughput stays at zero?One dump cannot distinguish a thread that is stuck from a thread that is simply executing that method at that instant; identical stacks across several dumps are what prove absence of progress. If stacks keep changing while no work completes, threads are running but achieving nothing, which is the signature of a livelock or a retry/spin loop rather than a block — and it usually comes with high processor usage instead of near-zero.
A traffic jam photographed once tells you cars are stopped; photographed three times a minute apart it tells you they never moved — and following the queue forward finds the one blocked junction.
saying these in an interview costs you the question
- Drawing conclusions from a single dump, so 'currently in this method' is mistaken for 'stuck in this method'.
- Restarting the process first and then trying to investigate, destroying all evidence.
- Reading a dump thread by thread instead of grouping by stack and counting.
- Interpreting an idle pool parked on an empty queue as the hang.
- Assuming near-zero processor usage rules out a concurrency fault, when it is precisely the signature of blocking.