skip to content

Which JVM, process and system meters does Boot expose, and how do you read jvm.memory / jvm.gc / system.cpu?

level: middleimportance: should knowfreq 48%

answer

  1. jvm.memory.used tags: area(heap/nonheap) + id(pool)
  2. jvm.gc.pause Timer; allocated/promoted counters
  3. system.cpu.usage vs process.cpu.usage — 0..1 fractions
  4. jvm.threads.states, jvm.classes.loaded
  5. process.uptime / process.start.time

basics

~10 s

Boot exposes jvm.memory.used/max/committed (tagged by area and pool), jvm.gc.pause and allocation counters, jvm.threads.live/daemon/peak, plus system.cpu.usage and process.cpu.usage, process.uptime and process.start.time. They come from Micrometer JVM/System binders.

solid answer

~30 s

These are the resource (USE-style) meters from Micrometer's JVM and system binders that Boot auto-registers. `JvmMemoryMetrics`: `jvm.memory.used/committed/max` as Gauges tagged `area`=heap|nonheap and `id`=memory pool (e.g. G1 Eden Space). `JvmGcMetrics`: `jvm.gc.pause` (a Timer of collection pauses tagged by action/cause), `jvm.gc.memory.allocated`, `jvm.gc.memory.promoted`, `jvm.gc.max.data.size`. `JvmThreadMetrics`: `jvm.threads.live/daemon/peak/states`. `ClassLoaderMetrics`: `jvm.classes.loaded/unloaded`. `ProcessorMetrics`: `system.cpu.usage` (whole host, 0..1), `process.cpu.usage` (this JVM, 0..1), `system.cpu.count`, `system.load.average.1m`. `UptimeMetrics`: `process.uptime`, `process.start.time`. To read heap: `GET /actuator/metrics/jvm.memory.used?tag=area:heap`. These are the go-to signals for OOM, GC-pressure and CPU-saturation alerts.

code

java · 13 lines
java
// These binders are the ones Boot auto-registers; you rarely register them
// yourself, but you CAN register extras (e.g. a second uptime source):
@Configuration
class ExtraJvmMetrics {
    @Bean
    JvmMemoryMetrics jvmMemoryMetrics() { return new JvmMemoryMetrics(); }
    @Bean
    ProcessorMetrics processorMetrics() { return new ProcessorMetrics(); }
}
// Prometheus exposition sample:
//   jvm_memory_used_bytes{area="heap",id="G1 Old Gen"} 5.24288E7
//   jvm_gc_pause_seconds_count{action="end of minor GC",cause="G1 Evacuation Pause"} 12.0
//   process_cpu_usage 0.37   # == 37%

go deeper

for a junior

Name the categories: memory, gc, threads, cpu, process uptime.

for a middle

Know the exact meter names, their types, tag keys, and the 0–1 CPU scale.

for a senior

Map meters to USE alerts (heap ratio, GC pause rate, cpu saturation) and know the -1/NaN edge cases.

for a principal

Reason about container cgroup CPU/memory views and cross-instance saturation alerting strategy.

## Why these exist Alongside the request metrics (RED), you need **resource/saturation** metrics (USE = Utilization, Saturation, Errors). Boot auto-registers Micrometer's JVM and system MeterBinders to provide them with no code. ## The binders and their meters ### Memory — `io.micrometer.core.instrument.binder.jvm.JvmMemoryMetrics` - `jvm.memory.used`, `jvm.memory.committed`, `jvm.memory.max` — all **Gauges**. - Tags: `area` = `heap` or `nonheap`; `id` = the specific memory pool name (`G1 Eden Space`, `G1 Old Gen`, `Metaspace`, `CodeCache`, …). So a single metric name fans out into many series by pool. ### Garbage collection — `JvmGcMetrics` - `jvm.gc.pause` — a **Timer** measuring stop-the-world pause durations, tagged `action` (end of minor/major GC) and `cause`. - `jvm.gc.memory.allocated` / `jvm.gc.memory.promoted` — counters of bytes allocated in young gen / promoted to old gen (great for allocation-rate dashboards). - `jvm.gc.max.data.size`, `jvm.gc.live.data.size`, `jvm.gc.overhead`. - Note: `JvmGcMetrics` registers a NotificationListener and is `Closeable` — Boot manages its lifecycle. ### Threads / classes — `JvmThreadMetrics`, `ClassLoaderMetrics` - `jvm.threads.live`, `jvm.threads.daemon`, `jvm.threads.peak`, `jvm.threads.states` (tagged by `state`=runnable/blocked/waiting/...). - `jvm.classes.loaded`, `jvm.classes.unloaded`. ### CPU / load — `ProcessorMetrics` - `system.cpu.usage` — recent CPU for the whole machine, **0.0–1.0**. - `process.cpu.usage` — recent CPU for this JVM process, 0.0–1.0. - `system.cpu.count` — available processors; `system.load.average.1m` — OS 1-minute load (-1 on Windows where unavailable). ### Uptime / files - `UptimeMetrics`: `process.uptime` (seconds), `process.start.time` (epoch). - `FileDescriptorMetrics`: `process.files.open`, `process.files.max` (Unix). ## Reading them - List: `GET /actuator/metrics/jvm.memory.used` -> shows available tags. - Filtered: `GET /actuator/metrics/jvm.memory.used?tag=area:heap` sums heap only. - Prometheus: names are dot->underscore + unit suffix, e.g. `jvm_memory_used_bytes{area="heap",id="G1 Old Gen"}`. ## Gotchas - `system.cpu.usage`/`process.cpu.usage` are **fractions 0–1**, not percentages — multiply by 100 in dashboards. On some restricted containers they can read as NaN/-1 if the JMX bean isn't available. - `jvm.memory.max` can be **-1/unbounded** for pools without a fixed max (e.g. some non-heap pools) — guard division in `used/max` ratios. - `system.load.average.1m` is `-1` on platforms that don't report it. - Container CPU: with cgroup limits, `system.cpu.count` reflects the JVM's view (`Runtime.availableProcessors()`), which honors container CPU quotas on modern JDKs — align alert thresholds accordingly. ## When to use Base your infra alerts on these: heap used/max ratio and old-gen growth for memory leaks, `jvm.gc.pause` sum/rate for GC pressure, `process.cpu.usage` for saturation, thread-state counts for thread-pool starvation.

  • What's the difference between system.cpu.usage and process.cpu.usage?
    system.cpu.usage is the whole host's recent CPU utilization; process.cpu.usage is only this JVM's share. Both are fractions 0.0–1.0. For pod/container saturation of your own app you alert on process.cpu.usage; system.cpu.usage shows noisy-neighbor pressure.
  • Why might jvm.memory.max return -1 for some pools?
    Some memory pools (certain non-heap pools) have no fixed upper bound, so the JVM reports max as -1 (undefined). A used/max ratio must guard against this or it produces meaningless/negative values.

saying these in an interview costs you the question

  • Reporting cpu.usage as already a percentage (0–100) instead of a 0–1 fraction
  • Assuming jvm.memory.used is a single series rather than one-per-pool via the id tag
  • Thinking jvm.gc.pause is a gauge of current pause rather than a Timer of pause durations

context