skip to content

You're told a service is 'hung.' How would you tell whether it's suffering from deadlock, livelock, or starvation?

level: seniorimportance: should knowfreq 45%

answer

  1. Two instruments: CPU usage + repeated thread dumps (jstack/jcmd a few seconds apart)
  2. Deadlock = low CPU, identical dumps, JVM prints 'Found one Java-level deadlock'
  3. Livelock = high CPU, dumps change, throughput still zero
  4. Starvation = system progresses but one thread perpetually behind
  5. Decision tree: CPU zero? → deadlock; CPU high + no progress? → livelock; partial progress + laggard? → starvation

basics

~20 s

Check CPU usage and thread dumps. Near-zero CPU with threads BLOCKED in a cycle on each other's locks = deadlock. High CPU but no progress = livelock. The system works but one specific thread never advances = starvation.

solid answer

~50 s

Combine a CPU-usage observation with repeated thread dumps (jstack / jcmd Thread.print) taken a few seconds apart. Deadlock: CPU near zero; threads are BLOCKED waiting on monitors, and successive dumps are identical showing a stable cycle — the JVM's dump even prints a 'Found one Java-level deadlock' section detecting the cycle automatically. Livelock: CPU is high but the application makes no forward progress; dumps differ each time (threads keep changing state, often spinning in tryLock/retry or yield loops) yet throughput stays zero. Starvation: the overall system is making progress, but one specific thread (or a low-priority group) is repeatedly stuck waiting on the same lock or never appears RUNNABLE across dumps, while others keep advancing. The decision tree is: low CPU + stable BLOCKED cycle → deadlock; high CPU + zero progress → livelock; partial progress + one thread perpetually behind → starvation.

go deeper

for a junior

Knows to look at CPU and capture a thread dump, and that deadlocked threads are BLOCKED.

for a middle

Uses jstack/jcmd, reads thread states, and recognizes the JVM's automatic deadlock section.

for a senior

Applies the CPU + repeated-dump decision tree to separate all three failures, knows the JVM only auto-detects lock cycles, and reaches for JFR/profilers for contention.

for a principal

Builds observability so these are caught proactively (lock-contention metrics, per-thread CPU, progress/heartbeat signals, alerts), and standardizes runbooks/diagnostics across services.

## The problem 'The service is hung' is a symptom, not a diagnosis. Three different liveness failures can all look like a hang from the outside, but they have **distinct fingerprints**. Telling them apart quickly saves a lot of time, because the fixes differ. ## Your two main instruments 1. **CPU usage** (top / OS metrics / per-thread CPU). This single number splits the cases fast: *is anything actually running?* 2. **Thread dumps.** A **thread dump** is a snapshot of every JVM thread's state and stack. Capture it with `jstack <pid>`, `jcmd <pid> Thread.print`, or by sending the process a SIGQUIT. Each thread shows a **state** — notably `RUNNABLE` (running or ready), `BLOCKED` (waiting to enter a `synchronized` monitor), `WAITING`/`TIMED_WAITING` (parked, e.g. on a `Lock` or `Object.wait`). Take **two or three dumps several seconds apart** so you can see whether threads are *moving* between dumps. ## The fingerprints ### Deadlock — low CPU, stable, detectable - **CPU:** near zero. The threads are blocked, doing nothing. - **Dumps:** essentially **identical** across snapshots — the situation is frozen. Threads are `BLOCKED` on monitors, each waiting for a lock another holds. - **Bonus:** the JVM **detects monitor deadlocks automatically** and prints a section like `Found one Java-level deadlock:` naming the threads and the cycle. (Note: it detects `synchronized`/`ReentrantLock` cycles, but not every kind of hang, e.g. a logical wait/notify lost-signal.) ### Livelock — high CPU, churning, zero progress - **CPU:** **high** — the threads are busy. - **Dumps:** **different each time** — threads keep changing state, often seen spinning in `tryLock`/retry loops, `Thread.yield()`, or condition rechecks. They're `RUNNABLE`, not `BLOCKED`. - **Key tell:** business metrics (queue depth, processed count) **don't move** despite the CPU burn. Busy + no progress = livelock. ### Starvation — partial progress, one laggard - **CPU:** the *system* may be fine/busy; the **victim** thread shows little CPU. - **Progress:** the application **is** doing work overall — it's not fully hung — but **one specific thread or group never advances**. Across successive dumps that thread is perpetually `WAITING`/`BLOCKED` on the same lock, or rarely appears `RUNNABLE`, while peers keep moving. - **Common settings:** unfair lock with constant barging, or a low-priority thread on a busy box. ## The decision tree 1. **Is CPU near zero?** - Yes → look for `Found one Java-level deadlock` / a stable BLOCKED cycle → **deadlock** (or a lost wait/notify if no cycle is reported). - No → go to 2. 2. **Is CPU high but throughput zero, with dumps changing each time?** → **livelock**. 3. **Is the system making *some* progress while one thread/group is perpetually behind?** → **starvation**. ## Tooling beyond jstack - **VisualVM / JFR (Java Flight Recorder)** to watch per-thread CPU and lock contention over time. - **Async-profiler / lock profilers** to see which monitor is hot and who's barging. - For deadlock specifically, `jcmd <pid> Thread.print` includes the same deadlock-detection section as jstack. ## How to derive the answer Anchor on two questions: *Is anything running (CPU)?* and *Is the whole system stuck or just one thread (progress)?* Low CPU + frozen → deadlock; high CPU + stuck → livelock; partial progress + one laggard → starvation. Two dumps a few seconds apart resolve almost every case.

  • Why take multiple thread dumps a few seconds apart instead of one?
    A single snapshot can't tell you whether threads are stuck or merely passing through a state. Comparing dumps shows movement: identical dumps imply a frozen deadlock; constantly changing states with no progress imply livelock; one thread stuck across dumps while others move implies starvation.
  • Does the JVM's automatic deadlock detection catch a lost wait/notify hang?
    Not necessarily. Its 'Found one Java-level deadlock' detection finds cycles of threads waiting on monitors/locks. A thread blocked in Object.wait() because a notify was missed has no lock cycle, so it won't be reported as a deadlock — you must diagnose it from the stacks.

saying these in an interview costs you the question

  • Assuming the JVM auto-detects all three — it only reports monitor/lock deadlock cycles
  • Diagnosing from a single thread dump (you need two+ to see whether threads are moving)
  • Equating high CPU with 'healthy' — high CPU plus zero progress is livelock
  • Forgetting to check whether the system is making partial progress, which distinguishes starvation

context