Micrometer offers Counter, Gauge, Timer and DistributionSummary. Explain how you choose between them for a given measurement, and why a Micrometer Gauge must be registered against an object it observes rather than being set imperatively.
answer
- counter = monotonic, consumed as a rate
- gauge = sampled at publish, weak reference, NaN when collected
- register gauges once over a long-lived field; push via AtomicInteger, MultiGauge for changing keys
- timer = count + sum + max; LongTaskTimer for in-flight work
- DistributionSummary = non-time distribution
basics
~20 sCounter is a monotonically increasing total you consume as a rate. Gauge is a value sampled at publish time that can go up or down. Timer records durations plus a count. DistributionSummary records non-time distributions. A Gauge holds the observed object weakly and reads it on publish, so instrumentation never keeps it alive.
solid answer
~60 sAsk what the backend will do with the number. - **Counter** — a cumulative, monotonically increasing total (requests, errors, bytes). You query its rate, never the raw value; never use one for something that can decrease. - **Gauge** — an instantaneous value that may rise or fall: queue depth, pool size, cache entries. Micrometer's gauge *samples*: you hand the registry an object plus a function, and the value is read when metrics are published. That is why there is no `set()` on `Gauge`. - **Timer** — duration events; publishes count, total time and max. - **DistributionSummary** — the same machinery for non-time quantities (payload bytes, items per batch). The registry holds the gauged object **weakly**, so instrumentation cannot cause a leak — but a gauge over a lambda-captured local reports `NaN` once that object is collected. The idiom is to keep the observed object in a long-lived field and register once: `Gauge.builder("queue.size", this.queue, Collection::size).register(registry)`. When the value must be pushed, register the gauge over an `AtomicInteger`/`AtomicLong` field and mutate that number, or use `MultiGauge` for a changing key set.
code
text · 12 lines// wrong: the only reference to the observed object is captured by the lambda
gauge("cache.size", () -> buildTempCache().size())
-> object becomes unreachable -> weak ref cleared -> gauge reports NaN
// right: registry weakly observes a long-lived field, registered once
this.cache = new Cache() // field on a singleton component
gauge("cache.size", this.cache, c -> c.size())
// right, when the value must be pushed
this.depth = new AtomicInteger() // long-lived field
gauge("queue.depth", this.depth) // register once
depth.set(n) // mutate the observed numbergo deeper
Name the four meter types with one example each, and state the two rules that matter: a counter only ever goes up, and a gauge is read at publish time rather than set.
Explain the gauge's sampling and weak-reference semantics, the NaN failure mode and its structural fix (register once over a long-lived field or an AtomicInteger), and the count/sum/max triple a Timer always yields.
Add the operational consequences: transients lost between publishes, re-registration returning the existing meter so gauges must not be registered per request, MultiGauge for changing key sets, and LongTaskTimer for work you must see while it is still running.
Frame it as a query-contract decision — what the backend must be able to compute (rates, quantiles, current level) determines the meter, cardinality and cost follow from that, and per-event fidelity belongs in logs or traces rather than in metrics at all.
## The question behind the question Interviewers ask this because picking the wrong meter type produces metrics that are *unqueryable* rather than merely imprecise. A counter that goes down breaks every rate function; a gauge used to represent a total silently loses every event between scrapes. ## Counter A monotonically increasing `double`. Its absolute value is meaningless on its own — it depends on when the process started — which is why every backend consumes it as a rate (`rate()`, `increase()`, or delta temporality on export). Because it is cumulative, a process restart resets it to zero, and rate functions detect and correct for that reset. Two rules follow: never decrement (Micrometer's API only offers `increment`), and never model "current number of X" as a counter. A subtle counting rule: count the *event*, not the state. "Requests received" is a counter; "requests in flight" is a gauge. ## Gauge A value read at publish time. The design deliberately offers no `gauge.set(x)`; you register something like `Gauge.builder("queue.size", queue, Collection::size).register(registry)` and Micrometer invokes the function when the registry publishes or is scraped. Consequences worth naming: - **Weak reference.** The registry holds the observed object weakly, so instrumentation never causes a leak. If the object becomes unreachable, the gauge reports `NaN` and the meter is effectively dead. This is the number-one gauge bug: registering a gauge over a lambda that captures a local, then wondering why the metric shows `NaN` after a while. The fix is structural — hold the observed object in a long-lived field (a field of a singleton component, say) and register the gauge once against that field. - **Sampling loses transients.** A gauge only ever shows the value at publish instants. A queue that spikes to 10,000 between scrapes is invisible. If peaks matter, pair the gauge with a counter of enqueue events, or record observed depths into a `DistributionSummary`. - **Registration is idempotent, not replacing.** Re-registering the same name and tags returns the *existing* meter rather than rebinding the function, so a gauge registered inside a request handler stays bound to the first object it ever saw. Register gauges once, at construction. For a value your code genuinely computes and must push, the correct shape is still a registered gauge over a mutable number: keep an `AtomicInteger` or `AtomicLong` field, register a gauge over it once, and mutate the field. When the set of gauged keys itself changes (per-tenant depths, per-partition lag), `MultiGauge` exists to re-register a whole row set in bulk. Note also that `Metrics.gauge(...)` returns the object you passed, not the meter, which makes the wrap-at-construction idiom read naturally. ## Timer Measures duration events and always yields at least `count`, `sum` (total time) and `max`. That triple already gives throughput (rate of count), average latency (sum/count) and a decaying maximum. Timers can additionally emit client-side percentiles or histogram buckets; that choice has its own trade-offs and belongs to distribution configuration rather than meter selection. A `LongTaskTimer` is the sibling for operations that are *in progress*. An ordinary Timer records nothing until the operation completes, so a job that hangs for an hour produces no signal at all until it finishes — or never, if it dies. `LongTaskTimer` reports the active task count and current durations while tasks run. Knowing that distinction is a common differentiator. ## DistributionSummary Identical statistics to Timer but for a unitless or non-time quantity: response payload bytes, items per batch, price per order. It accepts a base unit for the backend's benefit. Use it whenever you would want percentiles of something that is not a duration. ## A decision procedure 1. Is it a duration? → Timer (LongTaskTimer if you need visibility while it runs). 2. Is it a count of events that only accumulates? → Counter. 3. Is it a current level that can go down? → Gauge, registered once over a long-lived object or a mutable number field. 4. Is it a distribution of a non-time value? → DistributionSummary. And one negative rule: if you find yourself wanting the exact value of every individual event, you want logs or traces, not metrics. Metrics are aggregates by construction; per-event fidelity is the wrong tool.
- Why does Micrometer's Gauge have no set() method?Because a gauge is defined as a value sampled at publish time, not a value pushed by the application. The registry stores a weak reference to an object plus a function and invokes it when metrics are scraped or exported, which keeps instrumentation from causing leaks and guarantees the reported value matches the object's state at the moment of publication. If you need to push, you register a gauge once over a mutable number field such as an AtomicInteger and mutate that field.
- How do you gauge a set of values whose keys change at runtime, such as per-partition lag?Registering a gauge per key as keys appear leaves dead meters behind, because re-registering the same name and tags returns the existing meter rather than rebinding it, and removed keys keep reporting their last observed object. MultiGauge exists for exactly this: it manages a row set keyed by tags and re-registers the whole set on each refresh, adding new rows and optionally removing stale ones. You call its register method periodically with the current snapshot.
- You need to know that a nightly job is stuck. Why is a Timer insufficient?A Timer records nothing until the operation completes, so a job that has been running for six hours contributes no samples and looks identical to a job that never started. A LongTaskTimer reports the number of active tasks and their current durations while they are still in flight, which is what an alert on stuck work needs. In practice you use both: the LongTaskTimer for liveness and the Timer for completed-run latency.
A counter is the odometer, a gauge is the fuel needle, and a timer is a stopwatch that also keeps a tally of how many laps it timed.
saying these in an interview costs you the question
- Using a Counter for a value that can decrease, breaking rate queries
- Expecting a Gauge to capture peaks that occur between publishes
- Registering a gauge over a lambda-captured local and then wondering why the metric reads NaN
- Believing a Timer reports anything about an operation still in progress
- Reading a counter's absolute value instead of its rate