In Micrometer, a meter is identified by its name plus a set of tags (key/value dimensions). What makes a tag safe or unsafe to add, and what exactly breaks — in the application process, at export/scrape time, and in the metrics backend — when tag cardinality is unbounded?
answer
- meter identity = name + full tag set
- cardinality multiplicative, not additive
- meters never evicted → slow leak
- scrape timeout loses ALL series from that instance
- route template; unmatched → one constant; MeterFilter caps
basics
~20 sA tag is safe when its value set is bounded and decided by your code, not by callers. Each distinct tag combination is a permanent time series: registry memory that is never reclaimed, a larger export payload, backend index growth, and slower queries.
solid answer
~60 sIn Micrometer a meter's identity is name + the full set of tags, so **every distinct tag-value combination is a separate meter and a separate time series**. A tag is safe when its value set is bounded and owned by your code — HTTP method, status class, outcome, route template, queue name. Unsafe values are caller-controlled or unbounded: user id, order id, raw request path, full exception message, session id, timestamp. Cardinality is **multiplicative**: five tags with ten values each is 100,000 series from one meter name. Four things break, in order: 1. **Process** — the registry keeps a map of `Meter.Id` → meter; meters are not evicted on their own, so an unbounded tag is a slow leak (worse for Timers with histogram buckets). 2. **Export** — a pull-based scrape must serialise every series; slow enough and the scrape times out, losing *all* metrics from that instance. 3. **Backend** — index entries and active-series memory, degrading a shared cluster. 4. **Queries** — aggregations touch every matching series. Guard with `MeterFilter`, not with discipline alone.
code
text · 10 linesone meter name, four tags:
method (5) x status (12) x uri-template (40) x outcome (4)
= 9,600 series <- bounded, fine
swap uri-template for the RAW path:
method (5) x status (12) x raw-path (unbounded) x outcome (4)
= grows with traffic, forever
unmatched paths must collapse:
/admin.php, /.env, /wp-login -> uri="NOT_FOUND" (1 value, not 1 per probe)go deeper
Know that a meter's identity is its name plus all its tags, that each distinct combination is its own time series, and that ids and raw URLs must never be tags. Being able to state the safe/unsafe test is enough.
Explain the multiplicative math, that meters are not evicted so unbounded tags leak memory, and the route-template fix including collapsing unmatched requests to one value. Name MeterFilter as the enforcement point.
Walk the full failure chain in order — process memory, scrape payload and timeout losing all series from the instance, backend index and active-series memory, query and alert-evaluation cost — and describe how you would locate the offending metric and cap it in production before fixing instrumentation.
Frame cardinality as a shared-resource governance problem: per-service series budgets enforced centrally with filters rather than by convention, ownership of tag keys, and a deliberate split of high-cardinality identifiers into traces and logs with exemplars linking back to the aggregate.
## What a tag actually creates Micrometer identifies a meter by its `Meter.Id`: a name plus an unordered set of tags (string key/value pairs), plus meter type and base unit. Two registrations with the same name but different tag values are **two different meters**, each with its own state, exported as its own time series. (Aside: the name you register is not necessarily the name you query — each registry rewrites it into the backend's idiom on export; that translation is a separate topic from cost.) So the question "is this tag safe?" is really "how many distinct values will this key ever take, over the lifetime of the process?" ## The safe/unsafe test A tag is safe when its value set is **bounded, small, and chosen by your code**. HTTP method, status or status class, success/failure outcome, cache name, queue name, matched route template, a fixed enum of error kinds — all finite and enumerable at design time. A tag is unsafe when its values are **produced by data or by callers**: user id, tenant id in a large multi-tenant system, order id, the raw request path, a full exception message (which often embeds ids), session id, correlation id, a timestamp, a free-text search term. The practical heuristic: *if an external caller can influence the value, treat it as unbounded until proven otherwise.* And remember cardinality is **multiplicative across tags on the same meter**, not additive. Four tags with 10, 20, 5 and 8 values is 8,000 series — before anyone adds a fifth. ## Stage 1 — the application process The registry holds a concurrent map from `Meter.Id` to the meter instance. Each new tag combination allocates a meter object and its accumulators. For a `Timer` or `DistributionSummary` configured with percentile histograms, that is an array of bucket counters — hundreds of bytes to kilobytes each. Crucially, **meters are not evicted automatically**. There is no TTL; a meter created once for a one-off user id stays registered for the life of the process (you can call the registry's remove operation, but nothing does it for you). An unbounded tag is therefore a memory leak that looks like ordinary heap growth, and it also lengthens every export cycle because the registry iterates all meters. ## Stage 2 — export With a pull-based registry, every scrape serialises the entire meter set into the response body. Ten thousand series is unremarkable; a million makes the endpoint slow and the payload large. When the scrape exceeds the collector's timeout you do not lose the noisy series — **you lose every series from that instance**, so the blast radius of one bad tag is total blindness for that process, exactly when you need it. Push-based registries fail differently: larger batches, more egress, and dropped or throttled publishes at the vendor. ## Stage 3 — the backend Time-series databases size memory largely by **active series count**, and each series carries an index entry plus its own chunk stream. One careless service can push a shared cluster into memory pressure, ingestion rejection or compaction backlog — degrading observability for teams that did nothing wrong. This is why cardinality is a platform-level control, not a per-team preference. ## Stage 4 — queries An aggregation must read every series matching the selector, even when the result is a single line. A dashboard panel over a million-series metric is slow, and alert rule evaluation over it eats evaluation budget, so alerts start lagging or timing out. ## The route-template fix The canonical offender is tagging with the raw request path. `/orders/8412` and `/orders/8413` are different values, so the series count grows with your data. Tag the **route template your router matched** (`/orders/{id}`) so the value set equals the number of endpoints you wrote. The subtle half: requests that match *no* route. Unmatched paths are attacker- or scanner-controlled, so a vulnerability scanner probing thousands of paths mints a series per probe — a cheap denial-of-service against your own monitoring. Fold every unmatched request into a **single constant value** (a `NOT_FOUND`-style placeholder), and do the same for redirects. Any per-tenant or per-entity dimension deserves the same question: is the set finite and owned by us? ## Guard rails Discipline does not survive contact with a fleet, so enforce it centrally with `MeterFilter` on the registry configuration: - `MeterFilter.maximumAllowableTags(prefix, tagKey, maxValues, onMaxReached)` — tracks distinct values for a tag key and applies the supplied filter (commonly `MeterFilter.deny()`) once the cap is exceeded, so the damage stops at a known ceiling. - `MeterFilter.replaceTagValues(tagKey, mapping, exceptions...)` — normalises values, e.g. mapping anything unrecognised to a single `other` bucket. Use this when you want an overflow bucket rather than outright denial; `maximumAllowableTags` alone does not create one. - `MeterFilter.maximumAllowableMetrics(max)` — a global ceiling on meters in the registry. - `MeterFilter.denyNameStartsWith(prefix)` — drop whole families you never query. Treat these as production controls with the same seriousness as a connection-pool limit. ## When you genuinely need per-entity detail If the question is "which request was slow" rather than "what fraction is slow", metrics are the wrong store. Logs and traces are built for high-cardinality identifiers, and exemplars can attach a trace id to a metric bucket so you can jump from the aggregate to one concrete request — without minting a series per entity.
- Why is tagging with the raw request path dangerous even though your service only exposes a few dozen endpoints?The raw path includes path parameters and, more importantly, paths that matched nothing at all. The value set is therefore controlled by callers rather than by your code, so scanners and stray 404 traffic mint a new series per probed path. Because meters are not evicted, that memory is never reclaimed and the export payload keeps growing until the scrape times out. Tag the matched route template instead, and collapse every unmatched request into one constant value.
- How would you find out which metric is responsible before it takes the backend down?From the backend, count active series grouped by metric name and look at the top few — the offender is usually one name with orders of magnitude more series than its neighbours, and the tag key at fault becomes obvious once you group by each key. From inside the process, enumerate the registry's meters and group by name to get the same picture without touching the backend. Then apply a `MeterFilter` cap to that name immediately and fix the instrumentation afterwards.
- Does `MeterFilter.maximumAllowableTags` merge the excess values into an 'other' bucket?No. It tracks how many distinct values a tag key has taken and, once the cap is exceeded, applies the filter you supply — most often `MeterFilter.deny()`, which simply stops registering new meters beyond the ceiling. If you want an overflow bucket you compose it yourself, typically with `MeterFilter.replaceTagValues` mapping unrecognised values to a single placeholder. The distinction matters because deny loses the excess traffic entirely, while an overflow bucket keeps the totals correct.
Tags are the filter dropdowns on a report. A dropdown with ten options is useful; a dropdown with one entry per customer is not a filter, it is a separate report per customer that you now have to store forever.
saying these in an interview costs you the question
- Saying high cardinality is 'just a cost thing' — the first failure is in your own process memory, and the second is a scrape timeout that loses every metric from the instance.
- Assuming unused meters are cleaned up or expire; Micrometer does not evict them, so a one-off tag value is permanent for the process lifetime.
- Treating cardinality as additive across tags rather than multiplicative — 'it's only five tags' hides a six-figure series count.
- Claiming a raw URI tag is fine 'because we only have 40 endpoints', forgetting path parameters and unmatched/404 traffic that callers control.
- Believing a tag cap folds excess values into an 'other' bucket automatically; the standard cap denies new meters unless you add a value-replacement filter yourself.
- Reaching for a user id or request id tag to answer 'which request was slow' — that is a tracing or logging question, not a metrics one.