A service stops serving requests, the process is alive, and CPU is near zero. Working only from stack snapshots of every thread, how do you confirm it is a deadlock and identify which threads form the cycle?
answer
- three snapshots, seconds apart, before restarting
- identical stacks + zero CPU = blocked, not slow
- waiting-on plus held-by = one graph edge
- detectors miss pools, permits, queue slots, remote locks
- no cycle, just external wait = hang, needs a timeout
basics
~20 sTake several snapshots seconds apart. Deadlocked threads are blocked, not running, and their stacks are identical across snapshots. For each, read what it waits on and what it already holds, build the wait-for graph, and look for a cycle.
solid answer
~60 sCapture **three snapshots a few seconds apart** before restarting anything - a single snapshot cannot distinguish stuck from slow. Then: 1. **Filter to blocked threads.** Near-zero CPU plus identical stacks across all snapshots means truly parked, not looping (that would be livelock) and not progressing slowly. 2. **Extract edges.** For each blocked thread, note the resource it is waiting to acquire and the resources it already holds. Most runtimes record lock ownership, so waiting-on and held-by are both visible; some print a ready-made deadlock section. 3. **Build the wait-for graph and search for a cycle.** A cycle among threads that hold what the others want is the proof. Threads queued behind the cycle are victims, not members - do not chase them. 4. **Rule out lookalikes.** Frozen stacks parked on a socket read or an external call with no timeout form no cycle: that is a hang, not a deadlock. Application-managed resources (pool permits, queue slots) have no ownership metadata, so the detector stays silent and you must reconstruct edges from what each thread was doing.
code
text · 10 linesthread "worker-3" BLOCKED
waiting to acquire <lock 0x7f...A> (owner: worker-7)
holds <lock 0x7f...B>
thread "worker-7" BLOCKED
waiting to acquire <lock 0x7f...B> (owner: worker-3)
holds <lock 0x7f...A>
=> worker-3 -> worker-7 -> worker-3 : cycle on A and B,
acquired in opposite ordersgo deeper
Know that you take a thread dump, that blocked threads show what they wait on, and that two threads waiting on each other is the pattern to look for.
Explain multiple snapshots as the way to distinguish stuck from slow, and reconstruct the wait-for graph from waiting-on and held-by entries.
Own the whole triage: classify by CPU and stack movement, rule out unbounded external waits, and know that pools, permits and queue slots are the blind spot no automatic detector covers.
Push it upstream to observability - lock-wait and pool-saturation metrics, a standing dump-capture runbook, and bounded waits everywhere so a hang surfaces as a timeout with a trace instead of a silent freeze.
## Start by classifying the hang "Requests stopped, process alive" has several possible causes, and the diagnosis path branches immediately on two cheap observations: **CPU usage** and **whether stacks change over time**. - CPU near zero, stacks frozen: threads are blocked. Candidates are deadlock or an unbounded external wait. - CPU high, stacks moving, no work completing: livelock or a retry storm. - CPU high, stacks in one hot region: a spin or an infinite loop. - Everything stopped periodically and then resumes: a pause of the runtime, not a lock problem. That is why the first action is to take **multiple snapshots** of all threads, spaced a few seconds apart. One snapshot cannot distinguish a thread that is stuck from a thread that is merely slow at that instant; three snapshots showing byte-identical frames for the same threads is strong evidence of a permanent block. Capture them *before* restarting, because a restart destroys the only evidence and the bug is timing dependent, so it may not recur for weeks. ## Extract the edges A snapshot of a blocked thread usually tells you two things: the call it is parked in and, if the runtime tracks ownership, which lock it is trying to acquire plus which locks it already owns. Those two facts are exactly the endpoints of a wait-for edge: *this thread* -> *the owner of the thing it wants*. Write one edge per blocked thread, then search for a cycle. A cycle whose members each hold what the next one is waiting for is a proof of deadlock, not a hypothesis. Many runtimes will do this for you and print an explicit "found a deadlock" section covering their built-in lock types. Treat that as a fast path, not as the definition: it only sees locks the runtime itself manages. ## The blind spot: resources the runtime does not know about Detectors are silent for anything the application implements itself or borrows from a library that keeps its own bookkeeping: - **Connection or object pool permits.** Every thread sits in "borrow from pool" and no owner is recorded. The dump looks like contention, not deadlock. - **Semaphore permits and bounded-queue slots.** A blocked put or take names no owner. - **Worker threads in a bounded pool** occupied by tasks that are waiting for other tasks that are still queued. - **Locks held remotely** - a database row lock, a distributed lease. For these you reconstruct the graph by reasoning about roles, not by reading ownership fields: group blocked threads by what they were doing, count how many hold each scarce resource, and ask whether the holders can ever finish. A useful tell is *N threads blocked acquiring a resource type whose total capacity is N and whose holders are themselves blocked* - a saturated pool whose holders are waiting is the classic resource deadlock, and no lock detector will ever report it. ## Rule out the lookalikes deliberately Before concluding deadlock, check three things. **Is there a cycle at all?** Frozen stacks parked on a socket read, an unbounded blocking receive, or a remote call with no timeout mean threads are waiting on the outside world. There is no cycle, so it is a hang: the fix is a timeout and a bounded wait, not lock ordering. **Do stacks really not change?** If frames advance slowly, you have a slow dependency or a convoy, not a deadlock. **Is one thread blocked while holding a lock across I/O?** This is not itself a deadlock but is the most common precursor: a thread parked in a network call while owning a lock stalls everyone behind it for as long as the call takes, and if the remote side is in turn waiting on this service, you have a genuine distributed cycle. ## Turn the finding into a fix A confirmed cycle names the exact resources and the exact acquisition orders involved, and that is the actionable output. Record the two orders observed - which thread took X then Y, and which took Y then X - because the remedy is chosen from that evidence. Also record whether the resources are runtime locks or application resources, since only the former are visible to future automatic detection; the latter needs metrics such as pool wait time, saturation, and borrow depth to be visible at all next time.
- The dump shows 50 threads all blocked borrowing from a 50-connection pool, and no lock cycle is reported. Is this a deadlock?It is if every holder is itself blocked waiting for a connection it can never get - for example each task grabs one connection and then needs a second to finish. Capacity is exhausted by holders that cannot progress, which is a circular wait over a multi-instance resource, and no lock detector reports it because the pool is application state. If instead the holders are simply running slow queries, it is saturation, not deadlock, and the pool will drain.
- Why capture snapshots rather than attaching a debugger and stepping?Attaching and stepping perturbs timing, and a deadlock is already a state you cannot reproduce on demand, so anything that changes scheduling may hide the evidence. Snapshots are cheap, non-invasive, and can be taken repeatedly, giving you the change-over-time signal that distinguishes blocked from slow. They also produce an artifact you can attach to the ticket and reason about after the process has been restarted.
saying these in an interview costs you the question
- Concluding deadlock from a single snapshot, with no evidence that the stacks are unchanging.
- Restarting the process before capturing evidence, then being unable to explain the outage.
- Treating every frozen stack as a deadlock, including threads parked on an external call where no cycle exists.
- Trusting the runtime's automatic deadlock report as complete, so pool, semaphore and queue deadlocks go unnoticed.
- Blaming the threads merely queued behind the cycle instead of the cycle members that actually hold the resources.