How do you capture a thread dump with jcmd, and what would you look for in the output when diagnosing a hang or high CPU?
answer
- Thread.print (jstack replacement); add -l for lock info
- States: RUNNABLE / BLOCKED / WAITING / TIMED_WAITING
- Hang: look for deadlock block or many BLOCKED on same monitor
- High CPU: several dumps, find RUNNABLE in same frames
- Match top -H TID (hex) to nid= in the dump
basics
~10 sRun 'jcmd <pid> Thread.print' to print every thread's stack. To diagnose a hang look for BLOCKED threads and deadlocks; for high CPU look for RUNNABLE threads stuck in the same code across several dumps.
solid answer
~50 sA thread dump is a snapshot of every thread's call stack and state at one instant. With jcmd you take it via 'jcmd <pid> Thread.print' (the jstack replacement); add '-l' for lock/ownable-synchronizer details. In the output each thread shows a name, state (RUNNABLE, BLOCKED, WAITING, TIMED_WAITING), and its stack. For a hang, scan for a 'Found one Java-level deadlock' section, or many threads BLOCKED waiting to enter the same monitor while one holds it — that points to lock contention. For high CPU, a single snapshot is rarely enough: take three or four dumps a few seconds apart and look for threads that stay RUNNABLE in the same frames (a hot loop, a regex, GC pressure). Correlate native thread ids with OS-level CPU (e.g. top -H) to find which thread burns CPU, then match it back to the dump.
code
java · 12 lines// Shell session for a hang / high-CPU triage with jcmd (pid 4242):
//
// jcmd // 4242 com.acme.OrderService
// jcmd 4242 Thread.print -l > d1.txt // dump 1
// sleep 3; jcmd 4242 Thread.print -l > d2.txt // dump 2 (compare across)
// top -H -p 4242 // find the busy native TID, e.g. 18001
// printf '%x\n' 18001 // -> 4651 ; grep nid=0x4651 d1.txt
//
// In the dump you scan for either:
// "Found one Java-level deadlock:" (deadlock)
// or many lines like:
// "- waiting to lock <0x000000070f...> (a com.acme.Ledger)" (contention)go deeper
Knows the command 'jcmd <pid> Thread.print' produces a thread dump and that stacks show what each thread is doing.
Reads thread states, recognizes a deadlock block and lock contention, and knows to take multiple dumps for CPU issues.
Correlates native TIDs with OS CPU, interprets monitor ownership/contention patterns, and distinguishes CPU-bound from resource-starved hangs.
Builds an incident playbook (automated multi-dump capture, naming conventions, JFR for continuous data) and reasons about safepoint cost at scale.
## What a thread dump is A running JVM has many **threads**, each with its own **call stack** (the chain of method calls it is currently inside) and a **state**. A **thread dump** freezes the JVM for a moment and prints, for every thread, its name, its state, and its full stack. It is the single most useful artifact for diagnosing *liveness* problems: hangs, deadlocks, and CPU-bound loops. ## Capturing it with jcmd ``` jcmd <pid> Thread.print # standard dump jcmd <pid> Thread.print -l # also print lock/ownable-synchronizer info jcmd <pid> Thread.print -e # include extended thread info (newer JDKs) ``` `Thread.print` is jcmd's replacement for the legacy `jstack <pid>`. Redirect to a file and timestamp it if you are taking several: ``` jcmd <pid> Thread.print > dump-$(date +%s).txt ``` ## Reading thread states Each thread reports a `java.lang.Thread.State`: - **RUNNABLE** — executing (or able to execute) JVM bytecode; may also be blocked in a native/OS call. CPU burners live here. - **BLOCKED** — waiting to acquire a **monitor** (the lock behind a `synchronized` block) that another thread holds. - **WAITING** — parked indefinitely via `Object.wait()`, `LockSupport.park()`, `Thread.join()`, etc., until signalled. - **TIMED_WAITING** — same, but with a timeout (`sleep`, `wait(ms)`, `park(ns)`). ## Diagnosing a hang 1. **Deadlock first.** With `-l`, the JVM does deadlock detection and prints a `Found one Java-level deadlock:` block naming the threads and the locks forming the cycle. That is a definitive answer. 2. **Lock contention.** No deadlock but the app is stuck: look for *many* threads in **BLOCKED** state all `waiting to lock <0x...>` the **same** monitor address, while one thread *holds* that monitor (`- locked <0x...>`) and is itself stuck deep in slow work. That one holder is your bottleneck. 3. **External wait.** If almost everything is **WAITING/TIMED_WAITING** in socket reads or a connection pool's `await`, the app isn't burning CPU — it's blocked on a downstream resource (DB, remote service, exhausted pool). ## Diagnosing high CPU A single dump cannot tell a *busy* thread from one that just happened to be RUNNABLE. The technique: 1. Take **several dumps** (e.g. 4) a few seconds apart. 2. Find threads that are **RUNNABLE in the same or similar frames** across all of them — that's a hot path (tight loop, pathological regex, JSON churn, or constant GC). 3. To pin the *exact* hot thread, correlate with the OS. On Linux, `top -H -p <pid>` shows per-thread CPU and a native thread id (TID). Convert that TID to hex and match the `nid=0x...` field in the dump's thread header. Now you know which Java thread (by name and stack) is eating the core. ## Practical tips - Thread **names** are gold — name your pools (`http-nio-8080-exec-7`, `order-worker-3`) so a dump immediately tells you *which subsystem* is stuck. - Capturing a dump is cheap and safe; do it liberally in incidents. (It does briefly reach a safepoint, but the pause is tiny.) - For a continuous picture rather than snapshots, JFR (`JFR.start`) records thread/lock/CPU events over time with low overhead.
- Your dump shows hundreds of threads all WAITING in a connection-pool await. What does that tell you?The app is starved on a downstream resource, not CPU-bound. The pool is exhausted because borrowed connections aren't being returned fast enough — look at who holds connections and the downstream latency, not at the JVM's CPU.
- How do you map an OS thread eating CPU back to a Java thread in the dump?Get the native TID from top -H -p <pid>, convert it to hexadecimal, and find the matching nid=0x... in the thread dump header; the thread's name and stack identify it.
saying these in an interview costs you the question
- Trying to diagnose high CPU from a single dump
- Assuming RUNNABLE always means burning CPU (can be blocked in native I/O)
- Forgetting -l, then missing deadlock/lock-ownership detail
- Confusing WAITING (parked) with BLOCKED (waiting for a monitor)