skip to content

What is the difference between a continuous (always-on) JFR recording and a profiling recording, and how does this shape an incident-investigation strategy in production?

level: seniorimportance: should knowfreq 42%

answer

  1. Continuous = unbounded + default settings + ring buffer (maxage/maxsize)
  2. Profiling = time-boxed + profile settings + richer/denser
  3. Black-box flight recorder → dump the recent window on incident
  4. Always-on everywhere; escalate to profiling when reproducible
  5. Don't run 'profile' always — overhead + data volume

basics

~20 s

A continuous recording runs all the time with low-overhead settings and a bounded ring buffer, so you can dump the recent history when something goes wrong. A profiling recording is a short, time-boxed session with richer events and slightly higher cost, used for a focused investigation.

solid answer

~50 s

A **continuous recording** is unbounded in time, uses the low-overhead 'default' settings, and is capped by a ring buffer (maxage and/or maxsize) so old data rolls off — it is meant to run always, in production, as a flight recorder you can dump after the fact. A **profiling recording** is time-boxed (a fixed duration) with the denser 'profile' settings — more event types and higher sampling rate — for a deliberate investigation where you accept slightly more overhead to get richer data. The strategy implication: keep a continuous recording running everywhere so that when an incident occurs, you can JFR.dump the last N minutes and analyze what actually happened — no reproduction needed. Reserve profiling recordings for when you can target a specific window or reproduce a problem and want maximum detail. Continuous recording answers 'what happened during the incident I didn't expect'; profiling answers 'let me deeply study this known scenario'.

go deeper

for a junior

Knows there are two modes: an always-on continuous recording and a shorter profiling recording with more detail.

for a middle

Can configure a continuous recording with maxage/maxsize and a time-boxed profiling recording, and explain why default vs profile settings differ in cost.

for a senior

Designs the black-box + targeted-probe strategy: continuous recording fleet-wide with incident-triggered dumps, escalating to profiling only for reproducible scenarios; justifies the overhead/data trade-offs.

for a principal

Operationalizes JFR as part of observability — automated dump triggers on SLO breaches/OOM, standardized templates and retention windows, and policy for when deeper profiling is warranted across many services.

## Two modes, two purposes JFR recordings come in two practical flavours, distinguished by **how long they run**, **what settings they use**, and **how the buffered data is bounded**. ### Continuous (always-on) recording - **Duration:** unbounded — it just keeps running. - **Settings:** the **`default`** template, tuned for **< ~1% overhead**, so it is safe to leave on permanently. - **Bounding:** because it never stops, it must cap how much data it retains. You set **`maxage`** (keep only the last, say, 10 minutes) and/or **`maxsize`** (keep only the last, say, 100 MB). Older events roll off — it behaves like a **ring buffer / black-box flight recorder** (the aviation analogy JFR is named after). - **Purpose:** to *always have the recent past available*. When something unexpected happens, you **`JFR.dump`** the buffered window to a `.jfr` file and analyze it. You did not have to anticipate the failure or reproduce it. Start example: ``` jcmd <pid> JFR.start name=cont settings=default maxage=15m maxsize=200m # ... later, on incident ... jcmd <pid> JFR.dump name=cont filename=incident.jfr ``` ### Profiling recording - **Duration:** time-boxed — a fixed `duration` (e.g. 2 minutes) or an explicit start/stop around the scenario. - **Settings:** the **`profile`** template (or a custom `.jfc`) — denser method sampling, more event types enabled, lower thresholds. This costs **somewhat more than 1%**, which is fine for a bounded session. - **Purpose:** a **focused investigation** where you can target the window (you can reproduce the problem, or you know exactly when load will hit) and want maximum resolution — detailed flame graphs, fine-grained allocation profiles, etc. Start example: ``` jcmd <pid> JFR.start name=prof settings=profile duration=120s filename=profile.jfr ``` ## The strategy: black box + targeted probe The two modes compose into a robust production strategy: 1. **Leave a continuous recording on across the fleet.** Cost is negligible (< 1%), and the payoff is that *every* incident comes with a recording of the lead-up — you investigate the real event instead of trying to reproduce a transient failure. 2. **Trigger a dump on signals.** Configure dump-on-exit, dump-on-OutOfMemoryError, or have your alerting/ops automation run `JFR.dump` when latency/error SLOs breach, capturing the recent window automatically. 3. **Escalate to a profiling recording** only when you have a *reproducible* or *schedulable* scenario and the default recording's resolution is not enough — e.g. you want a high-fidelity flame graph of a specific endpoint under load. ## Why not just always run 'profile'? Because the higher sampling density and extra events raise overhead beyond the always-on budget and produce far more data. Continuous = cheap and broad but coarser; profiling = rich but costly and time-boxed. Picking the wrong one means either missing detail (too coarse to diagnose) or paying overhead you did not need (and a flood of data). The discipline is: **broad and cheap by default, deep and targeted on demand.**

  • Why bound a continuous recording with maxage or maxsize instead of letting it grow?
    An unbounded recording with no cap would consume ever-growing memory/disk. maxage/maxsize turn it into a ring buffer that retains only the recent window — enough to dump the lead-up to an incident — while keeping resource use constant.
  • Give a concrete trigger for automatically dumping a continuous recording.
    Dump on OutOfMemoryError (JFR can be configured to capture a recording when the JVM hits OOM), dump on JVM exit (dumponexit), or have monitoring/ops automation invoke jcmd JFR.dump when an SLO (latency/error rate) is breached.

saying these in an interview costs you the question

  • Saying you should always run the 'profile' template in production — its overhead and data volume make it unsuitable for always-on use.
  • Thinking a continuous recording keeps everything forever — it is bounded by maxage/maxsize and old data rolls off.
  • Believing you must reproduce an incident to profile it — the whole point of continuous recording is dumping the real event's lead-up.
  • Conflating 'continuous' with 'higher overhead' — continuous uses the low-overhead default settings precisely so it can run always.

context