Why do you take multiple thread dumps over time, and how do you read them to diagnose stuck threads and contention?
answer
- One dump = one instant → duration is invisible
- 3–5 dumps ~5s apart → crude time series
- Same frame across all dumps = stuck; changing frames = healthy
- Growing BLOCKED count on one monitor = contention; inspect the owner
- Script jstack in a loop with sleep; analyze/diff the set
basics
~20 sA single dump is just one instant, so you can't tell a momentarily-busy thread from a truly stuck one. Take several dumps a few seconds apart: a thread on the same stack frame in every dump is stuck, and many threads piling up on one lock across dumps reveals contention.
solid answer
~50 sOne thread dump is a single snapshot — it can't distinguish a thread that's stuck forever from one that just happened to be in a method at that instant. So you capture several dumps a few seconds apart (e.g. 3–5 dumps, ~5s spacing) and compare. A thread whose stack trace is identical across all dumps is genuinely stuck; one whose frames change is making progress. For contention, look at the distribution over time: if the count of threads BLOCKED on a particular monitor stays high or grows across dumps, that lock is a bottleneck, and the thread that consistently *holds* it (and is slow to release) is the cause. You also watch which threads accumulate in WAITING on a pool — steady growth signals resource exhaustion. The technique turns an instantaneous snapshot into a crude time series, which is what lets you separate transient blips from persistent stalls and trends. Tools like fastthread.io or `jstack` diffing help when dumps are large.
code
java · 12 lines// Capture a short time series of dumps, then compare:
//
// PID=12345
// for i in 1 2 3 4 5; do
// jstack $PID > dump_$i.txt # or: jcmd $PID Thread.print
// sleep 5
// done
//
// Reading the set:
// - grep a suspect thread name across files; identical stack in all 5 => STUCK
// - count 'waiting to lock <0x...abc>' per dump; persistently high => CONTENTION
// - find the matching 'locked <0x...abc>' line => the OWNER causing itgo deeper
Understands that one dump is a snapshot and that taking a few helps, without the detailed contention reading.
Takes several spaced dumps and compares stacks to tell stuck from progressing threads.
Reads contention as a distribution over dumps, identifies the lock owner driving it, and scripts/analyzes the capture; applies the RUNNABLE/native caveat.
Automates triggered dump capture on alarms, sets organization-wide triage runbooks, and reasons about overhead and when to escalate to continuous profiling.
## The core limitation: a dump is one instant A thread dump is a **snapshot** — it freezes the JVM's threads at a single microsecond. That instant cannot tell you about *duration*. A thread shown inside `processOrder()` might be: - **stuck there forever** (the bug), or - **passing through** it normally (perfectly healthy). The snapshot looks identical in both cases. **Duration is invisible in one dump.** This is the fundamental reason single dumps mislead. ## The fix: sample over time Taking **multiple dumps spaced a few seconds apart** turns the snapshot into a crude **time series**. A common recipe: 3–5 dumps, ~5 seconds between each (sometimes scripted in a loop). Now duration becomes observable: | Observation across dumps | Interpretation | |---|---| | Same thread, **same stack frame** in every dump | **Stuck** — it has not moved; investigate that frame | | Same thread, **different frames** each dump | Healthy — it is doing work and progressing | | A monitor with a **growing count of BLOCKED** threads | **Contention** building on that lock | | Threads steadily **accumulating in WAITING** on a pool | **Resource exhaustion** (connections/threads running out) | The key insight: you're inferring *time-in-state* from *repeated samples*, the same way a sampling profiler infers hot methods from periodic stack samples. ## Reading for a stuck thread Pick a suspect thread (often one whose stack sits in your code or an external call) and trace its stack across the dumps. If frames `a()->b()->c()` appear unchanged in dump 1, 2, and 3, that thread has been parked on `c()` the whole window. Common culprits: an external call with no timeout (HTTP/DB), an infinite loop, a wait on a condition that never signals. If the top frame is a blocking I/O call (e.g. `socketRead0`), the thread is waiting on a slow/dead downstream — note that such a thread still shows **RUNNABLE** (the native-I/O caveat). ## Reading for lock contention Contention is a **distribution** phenomenon, best seen across dumps: 1. In each dump, group threads by the monitor they are `waiting to lock`. 2. If the same monitor has many waiters **persistently** across dumps, that critical section is a bottleneck. 3. Identify the **owner** — the one thread that `locked` that monitor — and look at *its* stack: it's doing the slow work (e.g. a long computation or I/O while holding the lock) that serializes everyone else. 4. The fix follows from the owner's stack: shrink the critical section, move I/O out of the lock, use a finer-grained or read/write lock, or a lock-free structure. ## Practical capture Script it so the dumps are timestamped and aligned: ``` for i in 1 2 3 4 5; do jstack $PID > dump_$i.txt # or: jcmd $PID Thread.print sleep 5 done ``` Then diff or load the set into an analyzer (e.g. fastthread.io, IBM TDA) which clusters identical stacks and highlights threads that never moved and monitors with many waiters. For an intermittent production hang you may also trigger capture automatically on a latency/CPU alarm. ## Why not just one big profiler run? A sampling profiler gives smoother data, but jstack needs **nothing installed**, has negligible overhead per dump, and works headless over SSH — ideal for a live incident. The multi-dump technique is the cheapest way to add the missing *time* dimension to a snapshot tool.
- How can you tell, from a set of dumps, that a thread is genuinely stuck rather than busy?Its stack trace is identical across all the dumps; a busy-but-progressing thread shows different frames from one dump to the next.
- You see many threads BLOCKED on one monitor across all dumps. What single thread do you examine next, and why?The thread that holds (locked) that monitor — its stack shows the slow work being done inside the critical section that serializes everyone else, which is what you must fix.
- Why prefer multiple jstack dumps over a single full profiler session during a live incident?jstack needs nothing installed, has negligible per-dump overhead, and works headless over SSH; the multi-dump trick cheaply adds the time dimension a single snapshot lacks.
saying these in an interview costs you the question
- Concluding a thread is stuck from a single dump — you cannot see duration in one snapshot
- Spacing dumps too close or taking only one, so trends and stuck-vs-busy can't be distinguished
- Focusing only on the BLOCKED waiters and ignoring the lock OWNER whose slow work causes the contention
- Assuming a thread in your method is broken when its frames actually change across dumps (it's just working)