skip to content

How does the /actuator/metrics endpoint work, and how do you drill into a specific metric and filter by its tags?

level: middleimportance: should knowfreq 50%

answer

  1. MetricsEndpoint over Micrometer
  2. /metrics = names[]; /metrics/{name} = measurements + availableTags
  3. ?tag=key:value, repeated = AND
  4. snapshot, NOT a scrape (that's /prometheus)
  5. http.server.requests uri = templated (cardinality)

basics

~10 s

GET /actuator/metrics lists available metric names. GET /actuator/metrics/{name} (e.g. jvm.memory.used) returns that metric's current measurements plus its availableTags. Add ?tag=key:value to filter to a specific dimension.

solid answer

~40 s

`/actuator/metrics` is backed by `MetricsEndpoint` and sits on top of **Micrometer**. `GET /actuator/metrics` returns a `names` array listing every registered meter name (e.g. `http.server.requests`, `jvm.memory.used`, `system.cpu.usage`). You then GET `/actuator/metrics/{requiredMetricName}` to see that meter's current snapshot: a `measurements` array (statistic + value pairs such as `COUNT`, `TOTAL_TIME`, `MAX`) and an `availableTags` array listing each dimension (tag key) and its observed values. To narrow down, append one or more `?tag=key:value` selectors — e.g. `/actuator/metrics/http.server.requests?tag=status:500&tag=uri:/orders`. Multiple tag selectors are ANDed and the measurements are aggregated across whatever dimensions remain. Crucially this endpoint returns *point-in-time* aggregated values for ad-hoc inspection — it is not a time-series scrape. For Prometheus scraping you use the separate `/actuator/prometheus` endpoint.

code

java · 35 lines
java
// 1) List metric names
// GET /actuator/metrics  ->  { "names": ["http.server.requests", "jvm.memory.used", ...] }

// 2) Drill into one meter
// GET /actuator/metrics/http.server.requests
// {
//   "name": "http.server.requests",
//   "measurements": [
//     { "statistic": "COUNT",      "value": 1420 },
//     { "statistic": "TOTAL_TIME", "value": 88.7 },
//     { "statistic": "MAX",        "value": 0.42 }
//   ],
//   "availableTags": [
//     { "tag": "status", "values": ["200", "404", "500"] },
//     { "tag": "uri",    "values": ["/orders/{id}", "/orders"] },
//     { "tag": "method", "values": ["GET", "POST"] }
//   ]
// }

// 3) Filter to failed calls on one route (tags AND together)
// GET /actuator/metrics/http.server.requests?tag=uri:/orders/{id}&tag=status:500
// -> COUNT/TOTAL_TIME recomputed for just that slice

// Custom meter that will then appear under /actuator/metrics:
import io.micrometer.core.instrument.MeterRegistry;
import org.springframework.stereotype.Service;

@Service
public class OrderService {
    private final io.micrometer.core.instrument.Counter placed;
    public OrderService(MeterRegistry registry) {
        this.placed = registry.counter("orders.placed", "channel", "web");
    }
    public void place() { placed.increment(); /* orders.placed{channel=web} */ }
}

go deeper

for a junior

Know /metrics lists names and /metrics/{name} shows a value; know you can filter by tag.

for a middle

Explain measurements vs availableTags, the ?tag=key:value AND semantics, and that it's a snapshot not a scrape.

for a senior

Discuss Micrometer registry backing, http.server.requests tag design, and cardinality control.

for a principal

Reason about observability architecture: snapshot endpoint vs Prometheus/OTLP export, cardinality budgets, histogram config.

**Micrometer** is Spring Boot's metrics facade (a 'SLF4J for metrics'). Your app and Boot's auto-configuration register **meters** (counters, gauges, timers, distribution summaries) into a `MeterRegistry`. The `/actuator/metrics` endpoint (`MetricsEndpoint`) is a window onto that registry. **Listing.** `GET /actuator/metrics` returns: ```json { "names": [ "http.server.requests", "jvm.memory.used", "system.cpu.usage", "jvm.gc.pause", ... ] } ``` These are meter *names* (dotted convention). Names present depend on what's registered — HTTP server timings, JVM memory/GC/threads, system CPU, datasource pool, logback events, Tomcat sessions, etc. **Drilling in.** `GET /actuator/metrics/jvm.memory.used`: ```json { "name": "jvm.memory.used", "description": "The amount of used memory", "baseUnit": "bytes", "measurements": [ { "statistic": "VALUE", "value": 1.34E8 } ], "availableTags": [ { "tag": "area", "values": [ "heap", "nonheap" ] }, { "tag": "id", "values": [ "G1 Eden Space", "G1 Old Gen", "Metaspace" ] } ] } ``` - **`measurements`**: each entry is a `Statistic` (e.g. `VALUE` for a gauge; `COUNT`, `TOTAL_TIME`, `MAX` for a timer) and its current value. For a meter with tags, the value shown is the **aggregate across all tag combinations** unless you filter. - **`availableTags`**: the dimensions (tag keys) present and the observed values for each. This is your menu for drilling down. **Filtering by tag.** Append `?tag=<key>:<value>`. To see only heap memory: `GET /actuator/metrics/jvm.memory.used?tag=area:heap`. To combine dimensions, repeat the parameter — they are **ANDed**: `GET /actuator/metrics/http.server.requests?tag=uri:/api/orders&tag=status:500`. The returned `measurements` are recomputed over just the matching meters, and `availableTags` narrows to the remaining dimensions so you can drill further. **`http.server.requests`** is the metric interviewers love: a `Timer` auto-instrumented per request with tags `method`, `uri` (the *templated* path, e.g. `/orders/{id}`, to avoid cardinality explosion), `status`, `outcome`, `exception`. `COUNT` = number of requests, `TOTAL_TIME` = summed latency, `MAX` = slowest recent request. **Critical gotchas.** - **Not a scrape / not time-series.** `/actuator/metrics` gives a *current aggregated snapshot* for human/ad-hoc use. It does not return history and is not what Prometheus scrapes. The Prometheus format lives at the separate `/actuator/prometheus` endpoint (needs `micrometer-registry-prometheus`). Don't wire monitoring to `/actuator/metrics`. - **`uri` tag uses the route template, not the raw path**, precisely to keep tag cardinality bounded. High-cardinality tags (raw ids, user ids) can blow up memory — a real production concern. - **A metric only appears after it has been recorded at least once** — e.g. `http.server.requests` won't list until a request has been served. - **Percentiles** (like p95) are not in `availableTags`; they require enabling client-side percentiles/histograms on the timer, and are typically computed in the monitoring backend rather than read here. - Tag filter syntax is `tag=KEY:VALUE` (colon-separated), a common syntax slip in interviews. **When to use.** Quick manual inspection ('what's my current heap? how many 500s on /orders?') during incident triage without a dashboard. For real observability you export via a Micrometer registry (Prometheus/OTLP/etc.).

  • Why is the `uri` tag on http.server.requests a template like /orders/{id} rather than the literal path?
    To bound tag cardinality. Using literal paths (each unique id) would create a distinct time series per id, exploding memory in the registry and the backend. The templated route keeps the dimension small and aggregatable.
  • A teammate points Prometheus at /actuator/metrics and gets JSON it can't parse. What's wrong?
    /actuator/metrics is a human-oriented JSON snapshot, not the Prometheus exposition format. They should scrape /actuator/prometheus (provided by micrometer-registry-prometheus), which emits text-format time series.

saying these in an interview costs you the question

  • Believing /actuator/metrics is a time-series or the Prometheus scrape endpoint
  • Using raw path (not templated route) as a tag and ignoring cardinality
  • Writing the tag filter as tag=key=value instead of tag=key:value
  • Expecting p95/percentiles to appear without enabling histograms

context