skip to content

What are the pitfalls of gauging a collection's size, especially around references and re-registration?

level: seniorimportance: should knowfreq 45%

answer

  1. weak ref -> keep strong field or NaN
  2. bound to instance, not variable -> mutate in place
  3. register once, first-wins de-dup
  4. ConcurrentLinkedQueue.size() is O(n) per scrape
  5. tag values must stay bounded

basics

~20 s

Micrometer keeps only a weak reference to the collection, so if you don't hold a strong reference it gets collected and the Gauge shows NaN. Also, if you replace the collection with a new instance, the Gauge still points at the old one, and re-registering a same-named Gauge is ignored.

solid answer

~40 s

Three classic pitfalls. (1) Weak reference: Micrometer won't keep your collection alive, so a collection referenced only by the Gauge is GC-eligible and the Gauge then reports NaN — always retain a strong field. (2) Identity, not variable: the Gauge is bound to the specific object instance. If you do `this.list = new ArrayList<>()`, the Gauge keeps reading the old list; you must **mutate in place**. (3) Idempotent registration: Micrometer de-duplicates meters by name+tags and keeps the first; re-registering on every request silently no-ops and, worse, each call allocates a lambda that may leak. Additional concerns: `size()` on some concurrent collections (e.g. `ConcurrentLinkedQueue`) is O(n) and runs on every scrape; unbounded tag values create runaway cardinality; and the read must be thread-safe since scraping races with mutation.

code

java · 31 lines
java
import io.micrometer.core.instrument.MeterRegistry;
import io.micrometer.core.instrument.Tags;
import org.springframework.stereotype.Component;
import java.util.Map;
import java.util.concurrent.ConcurrentHashMap;
import java.util.concurrent.atomic.AtomicInteger;

@Component
public class SessionRegistry {

    // Strong reference kept in a field (prevents NaN from GC)
    private final Map<String, Session> sessions = new ConcurrentHashMap<>();
    private final AtomicInteger sessionCount = new AtomicInteger();

    public SessionRegistry(MeterRegistry registry) {
        // ConcurrentHashMap.size() is O(1)-ish, safe to gauge directly:
        registry.gaugeMapSize("sessions.active", Tags.empty(), sessions);

        // Alternative for O(n)-size collections: gauge an atomic you maintain
        registry.gauge("sessions.active.atomic", sessionCount, AtomicInteger::get);
    }

    public void open(String id, Session s) {
        if (sessions.putIfAbsent(id, s) == null) sessionCount.incrementAndGet();
    }

    public void close(String id) {
        if (sessions.remove(id) != null) sessionCount.decrementAndGet();
    }
    // NEVER: this.sessions = new ConcurrentHashMap<>();  // gauge would go stale
}

go deeper

for a junior

Likely unaware of weak-reference/NaN and re-registration semantics.

for a middle

Knows to keep a strong field; may miss identity-vs-variable and O(n) size costs.

for a senior

Articulates all three reference/registration pitfalls plus read cost, thread-safety, and cardinality, with concrete fixes.

for a principal

Establishes team conventions (register-once, back-with-atomic, bounded tags) and reasons about metrics-path performance and leak-safety trade-offs.

Gauging a collection's size (`registry.gaugeCollectionSize`, `gaugeMapSize`, or a hand-built `Gauge.builder(name, coll, Collection::size)`) is common — active sessions, cache entries, pending work — and it's where most Gauge bugs live. **Pitfall 1 — the weak reference / NaN trap.** Micrometer stores the state object with a **weak reference** by design (a metric must never keep an object alive and cause a leak). If the *only* thing pointing at your collection is the Gauge registration, the garbage collector is free to reclaim it. After that, the Gauge's value function has nothing to read and the meter reports **NaN**, which exporters usually drop — so your metric silently disappears. Fix: hold a **strong reference**, normally an instance field on a long-lived Spring bean. **Pitfall 2 — bound to object identity, not to the variable.** The Gauge captures the *specific instance* you passed. Consider: ``` private List<Job> jobs = new ArrayList<>(); // register gauge on jobs ... public void reset() { this.jobs = new ArrayList<>(); } // BUG ``` After `reset()`, the field points at a fresh list, but the Gauge still samples the **original** list (until it's GC'd), reporting a stale/frozen size. The rule: **mutate the collection in place** (`jobs.clear()`), never reassign the field. If you must swap instances, you have to re-bind — which runs into pitfall 3. **Pitfall 3 — registration is idempotent (first-wins).** Micrometer identifies a meter by **name + tags**. Registering a second Gauge with the same name+tags returns the existing meter and **ignores** your new state object/function. So a Gauge registered inside a controller method or a loop is created once and every later call is a no-op — plus each call constructs a lambda/builder that can accumulate. Register **once** at construction. **Pitfall 4 — cost and thread-safety of the read.** The value function runs on **every scrape**, potentially concurrently with writers. `size()` is O(1) on `ArrayList`/`HashMap`/`ConcurrentHashMap`, but **O(n)** on `ConcurrentLinkedQueue`/`ConcurrentLinkedDeque` (it walks the list) — an expensive, blocking read on the metrics path. Prefer a collection with O(1) size, or track the count in an `AtomicInteger` and gauge that. The read must also be safe under concurrent mutation (weakly-consistent iterators/size are fine; a plain `LinkedList` shared without synchronization is not). **Pitfall 5 — cardinality.** If you tag the collection Gauge with high-cardinality tag *values* (user id, tenant id from an unbounded set), you multiply time series. Keep tag values bounded. **Practical guidance:** - Use `registry.gaugeCollectionSize("sessions.active", Tags.empty(), sessions)` and **store `sessions` in a field**. - Never reassign the gauged collection — clear/mutate it. - Register once (constructor / `@PostConstruct`), never per request. - For O(n)-size or expensive-to-read state, back it with an `AtomicInteger`/`LongAdder` and gauge the atomic (`registry.gauge("...", counterField, AtomicInteger::get)`), updating the atomic on add/remove.

  • A teammate registers a Gauge on a list inside a @Scheduled method that runs every minute. Why is that wrong?
    Registration is idempotent by name+tags, so only the first run actually registers; every later run is a silent no-op that still allocates a builder/lambda. It also risks binding to whatever list instance existed on the first run. Register once at construction instead.
  • Why can gauging a ConcurrentLinkedQueue's size be a performance problem?
    ConcurrentLinkedQueue.size() is O(n) — it traverses the whole queue. Since the Gauge function runs on every scrape, a large queue makes each scrape walk the list, adding latency and CPU on the metrics path. Back it with an AtomicInteger/LongAdder and gauge that instead.
  • How do you fix a Gauge that intermittently reports NaN?
    NaN means the weakly-referenced source object was garbage collected. Ensure a strong reference is held — store the collection/object in a field on a singleton bean — and make sure you're not reassigning that field to a new instance.

saying these in an interview costs you the question

  • Relying on Micrometer to keep the collection alive
  • Reassigning the gauged field to a new instance
  • Registering the Gauge on every request/schedule tick
  • Gauging an O(n)-size concurrent collection on the hot scrape path
  • High-cardinality tag values on the collection Gauge

context