skip to content

When would you back a metric with an AtomicInteger/DoubleAdder and gauge it, versus using a Counter or a plain bound Gauge? How do you reason about it at design scale?

level: principalimportance: should knowfreq 25%

answer

  1. Counter=events (rate, sums across pods)
  2. bound Gauge=cheap read off a live object
  3. Atomic/LongAdder+Gauge=fluctuating, no natural source or O(n) size
  4. never decrement a Counter
  5. register once, strong ref, bounded tags, MeterFilter guardrails

basics

~20 s

Use a Counter for cumulative, only-up event totals. Use a bound Gauge when you already have a live object to read a current value from. Back a metric with an AtomicInteger/LongAdder + Gauge when the current value goes up and down but no natural object exposes it, so you maintain the number yourself.

solid answer

~50 s

Pick by data semantics and by whether a readable source exists. A **Counter** is for monotonic event totals where the consumer wants a rate — never lose increments between scrapes. A **bound Gauge** is for an instantaneous value you can *read* from an existing live object (pool size, queue depth) — sampled on scrape, cheap when size() is O(1). When the quantity fluctuates but **no object naturally exposes it** (e.g. in-flight requests, current retry budget), you keep the truth in an `AtomicInteger`/`AtomicLong`/`DoubleAdder`/`LongAdder`, update it on state changes, and register a Gauge that reads it — this decouples the measurement from any collection's cost and avoids O(n) scrape reads. At scale you also weigh: cardinality/tag governance, leak-safety (weak refs), aggregation across instances (Counters sum cleanly; Gauges need avg/max), scrape-path cost, and whether a `FunctionCounter`/`TimeGauge`/`MultiGauge` fits better. Standardize these as conventions so teams don't each rediscover the NaN and re-registration traps.

code

java · 20 lines
java
import io.micrometer.core.instrument.MeterRegistry;
import org.springframework.stereotype.Component;
import java.util.concurrent.atomic.LongAdder;

@Component
public class InFlightTracker {

    // Fluctuating value, no single collection exposes it -> we own the number.
    // LongAdder scales under high write contention better than AtomicLong.
    private final LongAdder inFlight = new LongAdder();

    public InFlightTracker(MeterRegistry registry) {
        // Gauge reads our authoritative counter on each scrape (O(1)).
        // Strong reference: inFlight is a field, so no NaN-from-GC.
        registry.gauge("http.server.requests.inflight", inFlight, LongAdder::doubleValue);
    }

    public void begin() { inFlight.increment(); }   // up
    public void end()   { inFlight.decrement(); }   // down -> must be a Gauge, not a Counter
}

go deeper

for a junior

Can distinguish Counter vs Gauge but not the atomic-backed pattern.

for a middle

Knows atomic-backed Gauge for up/down values; may not weigh aggregation or contention.

for a senior

Chooses correctly among the three and cites O(n) size, LongAdder contention, and register-once/strong-ref.

for a principal

Adds cross-instance aggregation semantics, cardinality governance via MeterFilter/customizers, variant meters (FunctionCounter/TimeGauge/MultiGauge), and team-wide conventions.

This is a design-judgment question: choosing the *right meter shape* and *backing storage* for a quantity, and codifying that as a convention. **The three building blocks.** - **Counter** — monotonic cumulative total. You `increment()`. Consumers derive a **rate**. Aggregates trivially across instances (sum of totals). Lossless between scrapes (every increment is captured). Use for: events happened. - **Bound Gauge** — `Gauge.builder(name, object, fn)`. **Sampled on read**; value can go up/down. Cheap iff the read (`fn`) is O(1) and non-blocking. Weakly references the source. Use when a **live object already exposes** the current value. - **Atomic-backed Gauge** — you own an `AtomicInteger`/`AtomicLong`/`LongAdder`/`DoubleAdder`, mutate it in your business code (`incrementAndGet`/`add`), and register a Gauge reading it (`registry.gauge(name, atomic, AtomicInteger::get)`). `registry.gauge(...)` even *returns* the atomic you pass, so you can inline it into a field. Use when the value fluctuates but **no natural object** exposes it, or when reading the natural object is **expensive** (O(n) size) or **racy**. **Decision heuristic.** 1. 'How many X have occurred (ever)?' -> **Counter** (or `FunctionCounter` if a library already holds the monotonic total). 2. 'What is the current value, and can I read it cheaply from an object I hold?' -> **bound Gauge** on that object. 3. 'Current value, but it fluctuates and nothing exposes it cheaply/safely' -> **AtomicInteger/LongAdder + Gauge**. **Why not a Counter for up-and-down values?** Counters are contractually monotonic; decrementing violates the model and breaks rate math (backends interpret a drop as a counter reset). In-flight-requests must be a Gauge. **Why an atomic instead of gauging the collection?** Two reasons: (a) **cost** — some collections' `size()` is O(n) (`ConcurrentLinkedQueue`), and the Gauge reads on every scrape; an atomic read is O(1). (b) **decoupling** — the counted quantity may not map to one collection (e.g. sum across shards), so you maintain a single authoritative number. **LongAdder/DoubleAdder vs Atomic.** Under high write contention, `LongAdder`/`DoubleAdder` scale better than `AtomicLong` (striped cells), at the cost of a slightly more expensive `sum()`/reset. For hot in-flight counters, `LongAdder` is often the better backing store; gauge its `doubleValue()`/`sum()`. **Cross-instance aggregation (a scale concern).** Counters **sum** across pods to a global total/rate. Gauges do **not** sum meaningfully — you aggregate with avg/max/sum-of-current depending on semantics (e.g. total in-flight = sum of per-pod Gauges; utilization = avg). Choosing Gauge vs Counter therefore changes what dashboards/alerts can express. **Leak-safety and lifecycle (design invariants to standardize).** Bound/atomic Gauges hold the source **weakly** — you must keep a strong field. Registration is **idempotent by name+tags** (first wins), so register **once** at construction, never per request. These two facts cause most production Gauge incidents (NaN, stale, silently-ignored). Bake them into a team convention or a small helper. **Related meter variants worth knowing at this level.** - `FunctionCounter` — expose a monotonic total held elsewhere without calling increment. - `TimeGauge` — a Gauge whose value is a duration, carrying a `TimeUnit`. - `MultiGauge` — manage a dynamic *set* of gauges with varying tag values (rebuilt via `register(rows)`), useful for per-category current values without leaking removed categories. - `Meter.builder`/custom meters for exotic cases. **Governance at scale.** Enforce naming (dot names), **bounded tag values** (no user/tenant ids), register-once, strong-reference, and 'don't decrement a Counter'. A `MeterRegistryCustomizer`/`MeterFilter` can cap cardinality, add common tags, or deny bad names centrally — the principal-level answer is as much about *guardrails* as about picking a meter.

  • Why can't 'current in-flight requests' be modeled as a Counter?
    It goes both up and down, but a Counter is contractually monotonic. Decrementing violates the model and backends read a decrease as a counter reset, corrupting rate calculations. It must be a Gauge (typically atomic-backed).
  • How do Counters and Gauges differ when aggregating across many instances?
    Counters sum cleanly to a global total and rate across pods. Gauges don't sum by default — you aggregate by semantics: sum for a total in-flight, avg for utilization, max for peak. This affects what alerts/dashboards you can build, so it's part of the design choice.
  • When would you reach for MultiGauge or FunctionCounter?
    MultiGauge when you need a dynamic set of gauges keyed by changing tag values (e.g. current count per active category) and must add/remove series cleanly. FunctionCounter when some object already holds a monotonic total and you just want to expose it without calling increment yourself.

saying these in an interview costs you the question

  • Modeling an up-and-down value as a decrementable Counter
  • Gauging an O(n)-size collection on the scrape path when an atomic would do
  • Assuming Gauges sum across instances like Counters
  • Ignoring cardinality/tag governance at fleet scale
  • Registering meters per request instead of once

context