skip to content

How do you bound metric label cardinality across a fleet, and where between instrumentation and storage should each control sit?

level: principalimportance: should knowfreq 46%

answer

  1. Four controls, not one
  2. A catch-all keeps totals honest
  3. Placement is the real decision
  4. Central rules hide the cost from producers
  5. Alert on creation rate, not the total

basics

~20 s

Allowlist a label's values with a catch-all for the rest, rewrite or drop it before storage, normalise continuous values into categories, or move the dimension onto traces. Emit-time fixes are correct but slow; central rules hide the cost.

solid answer

~50 s

Four controls bound cardinality — meaning the number of distinct series a metric produces, not one label's value count. **Allowlist** the permitted values and map the rest to a catch-all such as `other`, which fixes the series count while keeping totals honest. **Rewrite or drop** the label in the tier that collects telemetry, before it reaches storage. **Normalise** a continuous or templated value into the category you would actually group by — a route template, a small set of ranges. **Move** the per-request dimension onto a signal that carries it as an attribute rather than as identity. Placement is the real decision: emit-time is correct but costs a release per service, so it is never the incident tool; the collection tier is immediate and fleet-wide but leaves the application still paying to compute a label that silently vanishes. A store-side series limit is a backstop, not a design.

code

pseudocode · 9 lines
pseudocode
ALLOWED_TIERS = { "day_pass", "monthly", "annual", "staff" }

function tier_label(raw):
    if raw in ALLOWED_TIERS: return raw
    return "other"

counter("gym_checkin_requests_total")
    .with_labels(site = site, tier = tier_label(raw_tier))
    .increment()

go deeper

for a junior

Recall the four moves: allowlist the permitted values, drop or rewrite the label before storage, normalise a continuous value into a few categories, or move the dimension onto another signal.

for a middle

Explain what each control costs in lost resolution, and why folding the long tail into a catch-all value keeps totals such as request counts and error rates correct.

for a senior

Argue placement: at emit time the fix is correct but needs a release per service; in the collection tier it is instant and fleet-wide but invisible to the team still paying to compute the label.

for a principal

Own the standing policy — a bright-line rule on what may be a label, per-team series budgets, detection on series-creation rate rather than instantaneous totals, and who carries the cost when a control is applied centrally.

## Four controls, and what each one actually removes **Cardinality** here means the number of distinct time series a metric produces, which is the product of the distinct values of its labels — not the distinct value count of any single label. Four controls bound it, and they are not interchangeable: | Control | What it does | What you lose | |---|---|---| | **Allowlist the values** | Emit a label only for a known enumeration, mapping everything else to one catch-all value such as `other` | The identity of the long tail; totals are preserved | | **Drop or rewrite before storage** | Remove the label, or replace its value with a normalised one, in the tier that collects telemetry rather than in the application | Nothing at query time, but the app still computes and ships it | | **Normalise a continuous or templated value** | Replace the concrete value with the category it belongs to: a route template for a path, a small set of ranges for a size or duration | Per-value resolution, deliberately | | **Move the dimension to another signal** | Keep the per-request field as an attribute on a trace or a structured event, where the store does not make it part of a series' identity | A different query model, different retention, and sampling | The first and third are the ones that keep the metric useful. Allowlisting keeps the total honest — a request against an unknown value is still counted, just under `other` — so error rates and throughput stay correct while the series count is fixed by the size of the enumeration. Normalising is the right answer whenever the unbounded value is a *rendering* of a bounded one, which paths and error messages almost always are. ## Where each control can live, and why that is the real decision The same rewrite has very different properties depending on which stage applies it: 1. **In the instrumentation, at emit time.** The correct place. The label is never created, nothing downstream needs configuration, and the constraint is visible in the code review that introduced it. The cost is a release cycle per fix, per service — which is why it is never the tool you reach for during an incident. 2. **In the collection tier, before storage.** Central, immediate, applies across services without touching any of them. Three costs: the application still spends CPU and network shipping a label that is then discarded; the rule lives somewhere the emitting team does not read, so the next engineer sees a label in the code that never appears in queries and has no idea why; and central rules accumulate into a layer nobody dares delete. 3. **At the store, as a limit.** A cap on series per metric or per source. This is a backstop that turns a catastrophe into lost data for one metric. It is not a bounding strategy, and treating it as one means your design is "fail somewhere else". The honest senior answer is that you use the collection tier to stop the bleeding within minutes and the instrumentation to fix it within a sprint, and you track the central rules as debt with owners. A rewrite applied centrally and never removed is the mechanism by which one incident becomes permanent invisible complexity. ## Making it policy rather than incident response At fleet scale the interesting question is not which control but how you stop needing them. What a lead actually owns: - **A rule with a bright line.** A label's value set must be enumerable in code, or come from a small reference set that changes only at deploy time. Anything else requires an explicit, named exception with an owner. Bright lines survive review; "be thoughtful about cardinality" does not. - **A budget per team, expressed in series.** With a telemetry budget halved, a per-team series allocation turns a platform-vs-everyone argument into each team's own prioritisation. It also makes the cost visible to the people who create it, which central dropping specifically destroys. - **Detection on the derivative, not the level.** Alert on the rate at which new series appear and on the series count attributable to each metric name and each emitting service. A total that has already doubled tells you too late; a creation rate that steps up tells you on the deploy that did it. - **Attribution before mitigation.** Rank metric names by series count and by creation rate, then attribute by the labels the collection tier attaches, so you can name the service and the change. Dropping a label you cannot attribute just moves the problem. - **A default review question, not a specialist's job.** "What is this label's ceiling, and what is it multiplied by across the fleet?" belongs in the same reflex as "is this input validated". ## What you accept Every control here trades resolution for boundedness, and pretending otherwise is the failure mode. The judgement is deciding which questions the metrics layer will simply refuse to answer — per-member, per-request, per-identifier breakdowns — and making sure another signal answers them, rather than letting each team rediscover the tradeoff by melting the shared store.

  • A team proposes raising the store's series limit instead of fixing the label. When is that the right call?
    Only when the dimension is genuinely queried, its growth is bounded and understood, and the budget exists for the resulting footprint over the full retention window. A limit is a backstop that converts a catastrophe into lost data for one metric, so raising it deliberately for a known, bounded increase is legitimate. Raising it because something hit the ceiling is deferring the diagnosis, not making one.
  • With a telemetry budget halved, would you cut collection frequency or cut labels first?
    Labels, almost always. Series count drives the fixed per-series overhead that dominates the bill, while frequency drives sample volume, which compresses well and scales linearly. Halving frequency also degrades every incident investigation across the estate; dropping one unbounded label degrades one breakdown nobody could safely query anyway. Frequency cuts are the second lever, applied selectively to metrics nobody alerts on.
  • How do you stop a central rewrite rule from becoming permanent invisible complexity?
    Give it an owner, an expiry and a link back to the change that made it necessary, and treat it as debt on the emitting team's backlog rather than the platform's. Report the labels being discarded back to that team so the cost stays visible. The failure mode is a growing central layer nobody dares delete because no one remembers what depends on it.

saying these in an interview costs you the question

  • Reaches for a bigger store instead of bounding the label
  • Drops labels centrally and never tells the emitting team
  • Keeps a label nobody ever filters or groups by
  • Treats a store-side series limit as a design, not a backstop
  • Assumes moving a dimension to another signal is free
  • Alerts only on the total series count, never on its growth