skip to content

How do you use jcmd to control Java Flight Recorder (JFR) on a running JVM, and why is JFR often preferred over ad-hoc heap/thread dumps for performance investigations?

level: seniorimportance: should knowfreq 52%

answer

  1. JFR.start / JFR.check / JFR.dump / JFR.stop
  2. settings=default (~1%) vs settings=profile (more detail)
  3. duration + maxsize/maxage bound the recording
  4. Analyze .jfr in JDK Mission Control (flame graphs, allocation, locks, GC)
  5. Continuous + low overhead → beats one-shot dumps for perf

basics

~20 s

Use jcmd to start, check, dump, and stop a flight recording: 'JFR.start', 'JFR.check', 'JFR.dump filename=...', 'JFR.stop'. JFR records CPU, allocations, locks, and GC over time at very low overhead, so it captures continuous behavior that one-off dumps miss.

solid answer

~50 s

Java Flight Recorder (JFR) is a built-in, low-overhead event recorder in the JVM. jcmd is the command-line way to drive it on an already-running process: 'JFR.start name=rec settings=profile duration=120s' begins a recording, 'JFR.check' lists active recordings and their status, 'JFR.dump name=rec filename=/tmp/rec.jfr' writes what's been captured so far, and 'JFR.stop name=rec filename=...' ends it and flushes to disk. You then open the .jfr file in JDK Mission Control to see flame graphs of CPU hot methods, allocation profiles, lock contention, GC pauses, and I/O. JFR is preferred over ad-hoc thread/heap dumps because it samples continuously at roughly 1-2% overhead, so it captures behavior over a window rather than a single instant — you see where time and allocations actually go, not just one frozen snapshot. It's also production-safe enough to leave running, and you can start it from the command line on a live JVM without restart.

code

java · 10 lines
java
// jcmd-driven JFR lifecycle on a live JVM (pid 4242):
//
//   jcmd 4242 JFR.start name=diag settings=profile duration=120s \
//        maxsize=200m filename=/tmp/diag.jfr
//   jcmd 4242 JFR.check                         // -> recording 'diag' running
//   jcmd 4242 JFR.dump  name=diag filename=/tmp/diag-now.jfr   // partial snapshot
//   jcmd 4242 JFR.stop  name=diag filename=/tmp/diag-final.jfr // end + flush
//
// Then analyze: open diag-final.jfr in JDK Mission Control, or:
//   jfr print --events jdk.ExecutionSample /tmp/diag-final.jfr | head

go deeper

for a junior

Knows JFR records JVM events and that jcmd can start and stop a recording into a .jfr file.

for a middle

Can run the JFR.start/dump/stop lifecycle with name/settings/duration and open the result in Mission Control.

for a senior

Chooses JFR vs dumps appropriately, uses maxsize/maxage rings, reasons about overhead, and interprets flame graphs / allocation / lock / GC views.

for a principal

Standardizes always-on low-overhead JFR with incident-triggered dumps, manages overhead/storage trade-offs fleet-wide, and integrates .jfr capture into automated diagnostics pipelines.

## What JFR is **Java Flight Recorder (JFR)** is a profiling and event-collection engine *built into the JVM itself* (open-sourced from JDK 11; the engine exists in current OpenJDK builds). Instead of taking a single snapshot, JFR continuously emits thousands of small **events** — method-sampling for CPU, object allocations, monitor (lock) contention, GC pauses, thread parks, file/socket I/O, exceptions, and JVM-internal metrics — into an in-memory ring buffer that is periodically flushed to a `.jfr` file. Its whole design goal is **very low overhead** (commonly ~1-2% with the default/profile settings), so it can run on production. ## Driving JFR with jcmd jcmd is the way to control JFR on a JVM that is *already running* (you didn't pre-configure recording at launch): ``` # Start a 2-minute recording using the richer 'profile' template jcmd <pid> JFR.start name=diag settings=profile duration=120s filename=/tmp/diag.jfr # See what recordings exist and their state/size jcmd <pid> JFR.check # Snapshot the data captured so far WITHOUT stopping jcmd <pid> JFR.dump name=diag filename=/tmp/diag-partial.jfr # Stop the recording and flush the final file jcmd <pid> JFR.stop name=diag filename=/tmp/diag-final.jfr ``` Key arguments: - **settings** — a template: `default` (low overhead, ~1%) or `profile` (more detail, more overhead). You can supply a custom `.jfc`. - **duration** — auto-stop after a time window; omit it for an open-ended recording you stop manually. - **maxsize / maxage** — cap the on-disk/ring size so a long recording doesn't grow unbounded (oldest data is dropped). - **name** — a label so you can target dump/stop/check at a specific recording. You can also start JFR at launch (`-XX:StartFlightRecording=...`) or run it always-on as a continuous ring you `JFR.dump` from when an incident occurs. ## Analyzing the file Open the `.jfr` in **JDK Mission Control (JMC)** (or process it with `jfr print`). You get: **CPU flame graphs / hot methods** (where wall/CPU time goes), **allocation profiles** (which call sites allocate the most — the leak/GC-pressure source), **lock contention** (which monitors block which threads), **GC pause timelines**, and **I/O and exception** views. ## Why JFR beats ad-hoc dumps 1. **Continuous vs. instantaneous.** A thread dump is one frozen instant; a heap dump is one memory snapshot. JFR records *over a window*, so you see the *distribution* of where time and allocations actually go — far better for "the app is slow/CPU-heavy" questions where the cause is statistical, not a single stuck thread. 2. **Low, bounded overhead.** Sampling at ~1-2% means you can leave it on in production; a giant heap dump pauses the app and writes gigabytes. 3. **Correlated, multi-dimensional.** One recording ties together CPU, allocation, locks, GC, and I/O on a common timeline, so you can see *why* a latency spike happened (e.g. a GC pause coinciding with an allocation storm), not just one facet. 4. **No restart, command-driven.** jcmd starts/stops it on a live JVM — ideal during an incident. ## When you still want the dumps JFR is the default for *performance* (where does time/allocation go?). You still take a **heap dump** when you need the full object-reference graph to chase a *retention/leak* path, and a **thread dump** for a quick, human-readable view of a *deadlock or hang*. The mature practice: JFR running continuously, with on-demand thread/heap dumps via jcmd for the specific forensic question.

  • An intermittent latency spike happens a few times a day. Would you use a thread dump or JFR, and why?
    JFR. The cause is statistical and time-distributed, so you want a continuous recording (a maxage ring) that you JFR.dump right after a spike, then read the CPU/GC/lock/allocation timeline in JMC. A single thread dump would almost certainly miss the moment.
  • How do you keep a long-running JFR recording from filling the disk?
    Set maxsize and/or maxage on JFR.start so the recording behaves as a bounded ring buffer — once the cap is reached, the oldest events are discarded rather than growing the file indefinitely.

saying these in an interview costs you the question

  • Believing JFR needs a JVM restart to start (jcmd starts it live)
  • Thinking JFR replaces heap dumps for leak reference-path analysis
  • Ignoring maxsize/maxage and letting an open recording grow unbounded
  • Assuming JFR is a paid/commercial feature — it's open-source in current OpenJDK

context