When would you back a metric with an AtomicInteger/DoubleAdder and gauge it, versus using a Counter or a plain bound Gauge? How do you reason about it at design scale?
answer
- Counter=events (rate, sums across pods)
- bound Gauge=cheap read off a live object
- Atomic/LongAdder+Gauge=fluctuating, no natural source or O(n) size
- never decrement a Counter
- register once, strong ref, bounded tags, MeterFilter guardrails
basics
~20 sUse a Counter for cumulative, only-up event totals. Use a bound Gauge when you already have a live object to read a current value from. Back a metric with an AtomicInteger/LongAdder + Gauge when the current value goes up and down but no natural object exposes it, so you maintain the number yourself.
solid answer
~50 sPick by data semantics and by whether a readable source exists. A **Counter** is for monotonic event totals where the consumer wants a rate — never lose increments between scrapes. A **bound Gauge** is for an instantaneous value you can *read* from an existing live object (pool size, queue depth) — sampled on scrape, cheap when size() is O(1). When the quantity fluctuates but **no object naturally exposes it** (e.g. in-flight requests, current retry budget), you keep the truth in an `AtomicInteger`/`AtomicLong`/`DoubleAdder`/`LongAdder`, update it on state changes, and register a Gauge that reads it — this decouples the measurement from any collection's cost and avoids O(n) scrape reads. At scale you also weigh: cardinality/tag governance, leak-safety (weak refs), aggregation across instances (Counters sum cleanly; Gauges need avg/max), scrape-path cost, and whether a `FunctionCounter`/`TimeGauge`/`MultiGauge` fits better. Standardize these as conventions so teams don't each rediscover the NaN and re-registration traps.
code
java · 20 linesimport io.micrometer.core.instrument.MeterRegistry;
import org.springframework.stereotype.Component;
import java.util.concurrent.atomic.LongAdder;
@Component
public class InFlightTracker {
// Fluctuating value, no single collection exposes it -> we own the number.
// LongAdder scales under high write contention better than AtomicLong.
private final LongAdder inFlight = new LongAdder();
public InFlightTracker(MeterRegistry registry) {
// Gauge reads our authoritative counter on each scrape (O(1)).
// Strong reference: inFlight is a field, so no NaN-from-GC.
registry.gauge("http.server.requests.inflight", inFlight, LongAdder::doubleValue);
}
public void begin() { inFlight.increment(); } // up
public void end() { inFlight.decrement(); } // down -> must be a Gauge, not a Counter
}go deeper
Can distinguish Counter vs Gauge but not the atomic-backed pattern.
Knows atomic-backed Gauge for up/down values; may not weigh aggregation or contention.
Chooses correctly among the three and cites O(n) size, LongAdder contention, and register-once/strong-ref.
Adds cross-instance aggregation semantics, cardinality governance via MeterFilter/customizers, variant meters (FunctionCounter/TimeGauge/MultiGauge), and team-wide conventions.
This is a design-judgment question: choosing the *right meter shape* and *backing storage* for a quantity, and codifying that as a convention. **The three building blocks.** - **Counter** — monotonic cumulative total. You `increment()`. Consumers derive a **rate**. Aggregates trivially across instances (sum of totals). Lossless between scrapes (every increment is captured). Use for: events happened. - **Bound Gauge** — `Gauge.builder(name, object, fn)`. **Sampled on read**; value can go up/down. Cheap iff the read (`fn`) is O(1) and non-blocking. Weakly references the source. Use when a **live object already exposes** the current value. - **Atomic-backed Gauge** — you own an `AtomicInteger`/`AtomicLong`/`LongAdder`/`DoubleAdder`, mutate it in your business code (`incrementAndGet`/`add`), and register a Gauge reading it (`registry.gauge(name, atomic, AtomicInteger::get)`). `registry.gauge(...)` even *returns* the atomic you pass, so you can inline it into a field. Use when the value fluctuates but **no natural object** exposes it, or when reading the natural object is **expensive** (O(n) size) or **racy**. **Decision heuristic.** 1. 'How many X have occurred (ever)?' -> **Counter** (or `FunctionCounter` if a library already holds the monotonic total). 2. 'What is the current value, and can I read it cheaply from an object I hold?' -> **bound Gauge** on that object. 3. 'Current value, but it fluctuates and nothing exposes it cheaply/safely' -> **AtomicInteger/LongAdder + Gauge**. **Why not a Counter for up-and-down values?** Counters are contractually monotonic; decrementing violates the model and breaks rate math (backends interpret a drop as a counter reset). In-flight-requests must be a Gauge. **Why an atomic instead of gauging the collection?** Two reasons: (a) **cost** — some collections' `size()` is O(n) (`ConcurrentLinkedQueue`), and the Gauge reads on every scrape; an atomic read is O(1). (b) **decoupling** — the counted quantity may not map to one collection (e.g. sum across shards), so you maintain a single authoritative number. **LongAdder/DoubleAdder vs Atomic.** Under high write contention, `LongAdder`/`DoubleAdder` scale better than `AtomicLong` (striped cells), at the cost of a slightly more expensive `sum()`/reset. For hot in-flight counters, `LongAdder` is often the better backing store; gauge its `doubleValue()`/`sum()`. **Cross-instance aggregation (a scale concern).** Counters **sum** across pods to a global total/rate. Gauges do **not** sum meaningfully — you aggregate with avg/max/sum-of-current depending on semantics (e.g. total in-flight = sum of per-pod Gauges; utilization = avg). Choosing Gauge vs Counter therefore changes what dashboards/alerts can express. **Leak-safety and lifecycle (design invariants to standardize).** Bound/atomic Gauges hold the source **weakly** — you must keep a strong field. Registration is **idempotent by name+tags** (first wins), so register **once** at construction, never per request. These two facts cause most production Gauge incidents (NaN, stale, silently-ignored). Bake them into a team convention or a small helper. **Related meter variants worth knowing at this level.** - `FunctionCounter` — expose a monotonic total held elsewhere without calling increment. - `TimeGauge` — a Gauge whose value is a duration, carrying a `TimeUnit`. - `MultiGauge` — manage a dynamic *set* of gauges with varying tag values (rebuilt via `register(rows)`), useful for per-category current values without leaking removed categories. - `Meter.builder`/custom meters for exotic cases. **Governance at scale.** Enforce naming (dot names), **bounded tag values** (no user/tenant ids), register-once, strong-reference, and 'don't decrement a Counter'. A `MeterRegistryCustomizer`/`MeterFilter` can cap cardinality, add common tags, or deny bad names centrally — the principal-level answer is as much about *guardrails* as about picking a meter.
- Why can't 'current in-flight requests' be modeled as a Counter?It goes both up and down, but a Counter is contractually monotonic. Decrementing violates the model and backends read a decrease as a counter reset, corrupting rate calculations. It must be a Gauge (typically atomic-backed).
- How do Counters and Gauges differ when aggregating across many instances?Counters sum cleanly to a global total and rate across pods. Gauges don't sum by default — you aggregate by semantics: sum for a total in-flight, avg for utilization, max for peak. This affects what alerts/dashboards you can build, so it's part of the design choice.
- When would you reach for MultiGauge or FunctionCounter?MultiGauge when you need a dynamic set of gauges keyed by changing tag values (e.g. current count per active category) and must add/remove series cleanly. FunctionCounter when some object already holds a monotonic total and you just want to expose it without calling increment yourself.
saying these in an interview costs you the question
- Modeling an up-and-down value as a decrementable Counter
- Gauging an O(n)-size collection on the scrape path when an atomic would do
- Assuming Gauges sum across instances like Counters
- Ignoring cardinality/tag governance at fleet scale
- Registering meters per request instead of once