skip to content

Timer & DistributionSummary

Timers record latency, long task timers track work still in flight, and distribution summaries handle non-time distributions, all with optional percentiles and histogram buckets. Interviewers ask about client-side percentiles versus histograms, because you cannot average percentiles across instances.

part ofSpring Frameworkoverview, primer and where to startread it →
on this pageshow

questions

5

What is a Micrometer Timer, and how do you use it to record the latency of a block of code?

level: juniorimportance: must knowfreq 70%

answer

  1. count + total time + max
  2. Timer.builder(...).register(registry)
  3. record(Runnable/Supplier) is exception-safe
  4. nanos internally, base unit per backend
  5. short events; bound tag cardinality

basics

~20 s

A Timer measures how long operations take and counts how often they run. Register one on the MeterRegistry, then wrap the code with timer.record(() -> ...) or record a Duration, and Micrometer tracks count, total time, and max.

solid answer

~40 s

A Micrometer Timer is a meter for short-duration events (latency). Every Timer tracks at least three statistics: the count of recordings, the total time accumulated, and the maximum recorded time. You build it with Timer.builder("http.server.requests").tags(...).register(registry) or the fluent MeterRegistry helpers. To record, either pass a lambda to timer.record(Supplier/Runnable) so timing is automatic even on exceptions, or record a known Duration with timer.record(duration). Time is stored in nanoseconds internally but exposed in a configurable base unit (usually seconds or milliseconds) per monitoring system. You never call it repeatedly with different tag values from a hot loop without bounding cardinality. Spring Boot auto-instruments web requests, so you mostly add Timers for your own service/business operations.

code

java · 23 lines
java
import io.micrometer.core.instrument.MeterRegistry;
import io.micrometer.core.instrument.Timer;
import org.springframework.stereotype.Service;

@Service
class OrderService {

    private final Timer processTimer;

    OrderService(MeterRegistry registry) {
        this.processTimer = Timer.builder("orders.process")
            .description("Time to process a single order")
            .tag("type", "standard")
            .register(registry);
    }

    Order process(Order order) {
        // Times the lambda; records even if it throws.
        return processTimer.record(() -> doProcess(order));
    }

    private Order doProcess(Order order) { /* ... */ return order; }
}

go deeper

for a junior

Know that Timer measures duration + frequency, and that record(lambda) is the safe way to time a block.

for a middle

Explain the count/total/max trio, base-unit handling, and cardinality dangers of tags.

for a senior

Contrast Timer vs LongTaskTimer vs DistributionSummary and know Spring Boot's built-in http.server.requests/@Timed.

for a principal

Reason about aggregation cost, series cardinality budgets, and why averages are insufficient for SLOs.

## What a Timer is Micrometer is the metrics facade Spring Boot uses (like SLF4J but for metrics). A **`Timer`** is one kind of *meter* designed to measure the **frequency and duration of short-lived events** — typically request/method latency. It is for events measured in milliseconds/seconds that complete quickly, as opposed to long-running jobs (that's `LongTaskTimer`). Every `Timer` publishes at minimum three time series: - **count** — how many times something was recorded. - **total time** — the sum of all recorded durations (from which you compute *average* = total/count). - **max** — the maximum single recording observed in a decaying interval. ## Creating one You obtain a `MeterRegistry` (Spring Boot auto-configures one, e.g. a `PrometheusMeterRegistry`) via constructor injection, then build the timer: ```java Timer timer = Timer.builder("orders.process") .description("time to process an order") .tags("region", "eu") .register(registry); ``` Meter names use dot-separated lowercase; Micrometer translates the name+tags into each backend's naming convention (dots→underscores for Prometheus, etc.). **Tags** are key/value dimensions — keep their value set *bounded* (low cardinality) because each unique tag combination is a separate time series. ## Recording Three main ways: 1. **Wrap a lambda** — timing is automatic and exception-safe: ```java timer.record(() -> orderService.process(order)); Order o = timer.record(() -> orderService.load(id)); // Supplier overload returns the value ``` 2. **Record a known Duration**: ```java timer.record(Duration.ofMillis(elapsed)); ``` 3. **Timer.Sample** — start now, stop later (covered in its own question), useful when start and stop are far apart in the code. ## Base units & storage Internally durations are kept in **nanoseconds**. Each monitoring system has a preferred **base unit** (Prometheus → seconds, others → milliseconds). You don't scale manually — you read the timer in a specific unit: `timer.mean(TimeUnit.MILLISECONDS)`, `timer.max(TimeUnit.SECONDS)`. ## Gotchas - **Averages hide tail latency.** count/total gives you a mean only; for p95/p99 you must enable percentiles or histograms (separate concern). - **Cardinality explosion**: never tag with unbounded values (user id, raw URL with path params, exception message). Each combination is a persistent series. - **Don't register a new Timer per call** — register once (registries dedupe by name+tags, but building repeatedly is wasteful and error-prone). Injected/field-held is idiomatic. - Spring Boot already times HTTP server requests (`http.server.requests`) and `@Timed`-annotated methods, so check before rolling your own. ## When to use Use a `Timer` for any operation that starts and finishes quickly and where you care about how long and how often: DB calls, downstream HTTP calls, business operations. For durations of *concurrently in-flight* long tasks, use `LongTaskTimer`. For non-time magnitudes (payload sizes, batch counts), use `DistributionSummary`.

  • What three statistics does every Timer always expose?
    count (number of recordings), total time (sum of durations, giving mean), and max (largest single recording over a decaying window). Percentiles/histograms are opt-in on top of these.
  • Why is timer.record(Runnable) preferable to manually measuring start/stop with System.nanoTime()?
    It records the duration even when the wrapped code throws (via try/finally internally), avoids off-by errors in unit conversion, and keeps timing logic out of business code.

saying these in an interview costs you the question

  • Thinking a Timer stores every individual sample so you can query any percentile later (it stores aggregates unless histograms/percentiles are enabled)
  • Registering a new Timer on every method invocation
  • Tagging with unbounded values like user id or full request URL
  • Believing average latency from count/total is enough to reason about tail behavior

context

open as a page

What is a LongTaskTimer and how does it differ from a regular Timer?

level: middleimportance: should knowfreq 50%

basics

~10 s

A LongTaskTimer measures tasks that are still running. It reports how many are in flight and their current accumulated duration, updated while they run. A regular Timer only records after an operation finishes.

open as a page

When and how do you use Timer.Sample instead of Timer.record(Runnable)?

level: middleimportance: should knowfreq 55%

basics

~10 s

Use Timer.Sample when the start and stop of an operation are in different places (e.g. async callbacks) so you can't wrap them in one lambda. Call Timer.start(registry), carry the Sample, then sample.stop(timer) later.

open as a page

What is a DistributionSummary and when would you use it instead of a Timer?

level: seniorimportance: should knowfreq 45%

basics

~10 s

A DistributionSummary tracks the distribution of non-time measurements, like request payload sizes or items per batch. It records count, total, and max like a Timer but for arbitrary magnitudes instead of durations, via summary.record(value).

open as a page

Explain the difference between publishPercentiles, publishPercentileHistogram, and serviceLevelObjectives on a Timer/DistributionSummary, and why it matters for aggregation across instances.

level: principalimportance: should knowfreq 40%

basics

~10 s

publishPercentiles computes percentiles inside each instance (not mergeable across instances). publishPercentileHistogram exports histogram buckets the backend aggregates to compute percentiles. serviceLevelObjectives adds explicit boundary buckets so you can measure the fraction meeting a target.

open as a page