skip to content

Which metric label values count as unbounded, and why is 'there are only a few thousand today' the wrong test?

level: middleimportance: should knowfreq 58%

answer

  1. Unbounded is not the same as large
  2. Ask who chooses the value
  3. A ceiling, not a current count
  4. Rotating values accumulate as churn
  5. Validation does not make a set finite

basics

~20 s

A metric label is unbounded when nothing caps its value set — identifiers, raw URL paths, error message text, timestamps. Counting today's distinct values measures current traffic, not the ceiling, and misses values that rotate and accumulate over the retention window.

solid answer

~50 s

Boundedness is a property of **where the value comes from**, not of how many values exist right now. If the code chooses it — an enum, a branch name, a short reference list that changes only at deploy — it is bounded. If a request, a clock, a generated identifier or an exception's text chooses it, it is unbounded however small it looks today. The canonical offenders are entity identifiers, a raw request path with an id in it, an error message string, a timestamp or duration, and a user-agent string. 'Only a few thousand today' fails for two reasons: it samples a distribution with no ceiling, and it reads an instant while the store bills a window. Values that rotate per deploy or per process look tiny at any moment and accumulate as churn.

code

text · 5 lines
text
UNBOUNDED  path="/members/8814/checkins"       -> one series per member
BOUNDED    route="/members/{id}/checkins"      -> one series per route

UNBOUNDED  error="upstream timed out after 4.812s" -> one series per message
BOUNDED    error_class="upstream_timeout"      -> one series per class

go deeper

for a junior

Recall the classic unsafe label values: entity identifiers, raw URL paths that contain them, error message text, timestamps and user-agent strings. Knowing the list is enough at this level.

for a middle

Explain that boundedness comes from where the value originates rather than how many values exist today, and give the bounded rewrite for each unsafe example.

for a senior

Bring up churn: values that rotate per deploy or per process look small at an instant and accumulate across the retention window, which is exactly what a current-count check cannot see.

for a principal

Set the review rule the organisation uses: a label's value set must be enumerable in code, or come from a small reference set that changes only at deploy time. Anything else needs a named exception with an owner.

## Bounded and unbounded, precisely A **metric label** is a key/value dimension that forms part of a time series' identity, so the set of values a label can take is the set of series that label will create. A label is **bounded** when that value set is finite, small, and fixed by something you control — an enumeration in code, a short reference list, a set of environment names. It is **unbounded** when the value set is decided by something outside your control: user input, a generated identifier, a clock, a free-text message. "Unbounded" is not a synonym for "large". It means *nothing in the system caps it*. That distinction is the whole question, because the numbers are deceptive in both directions. A bounded label with 400 values is fine. An unbounded label with 30 values today is not. ## The canonical sources | Unsafe label value | Why it is unbounded | Bounded form | |---|---|---| | A member, order or session identifier | One value per entity, growing forever | The entity's *type* or *tier* | | A raw request path | One value per identifier embedded in the URL | The route template | | An error or exception message | Includes timings, ids, remote text | A stable error class or code | | A timestamp, date or duration | Every value is new by construction | A bucketed range, or nothing | | A user agent string | Effectively unlimited variants in the wild | A parsed family name | | A generated instance or build identifier | New value per process or per release | Kept, but counted as churn | Every row has the same shape: something that varies *per event* was made part of an identity that is supposed to vary *per category*. A metric label answers "which kind of thing is this?"; per-event detail belongs on a signal that carries it as an attribute rather than as identity. ## Why "there are only a few thousand today" is the wrong test Three separate reasons, and a strong answer gives at least two: 1. **It measures the symptom, not the mechanism.** The number of distinct values today is an observation about current traffic. Boundedness is a property of where the values come from. If the source is unconstrained, today's count is a sample of a distribution with no ceiling, and it tells you nothing about tomorrow's. 2. **It measures an instant, and the store bills a window.** Values that *rotate* look tiny at any moment and accumulate relentlessly. A label carrying a container or process identifier on a fleet of 47 hosts running 12 containers each shows 564 values at any instant; redeploy eleven times a week and it has created several thousand distinct series across a retention window, none of which disappear when the process does. This is **churn**, and it is invisible to any check that reads the current value count. 3. **It is measured on the wrong traffic.** Counts are usually taken in a low-traffic environment or soon after a rollout, before the long tail of real-world values has appeared. The trap is that the label works. Dashboards render, queries return, nothing warns. The count grows monotonically with traffic and time, and the failure arrives as a step change when the store crosses a memory or index threshold — typically weeks after the change that caused it, and usually attached to the wrong deploy in everyone's mental model. ## The test that actually works Ask **who chooses the value**: - **The code chooses it** — a constant, an enum member, a branch label. Bounded. Its ceiling is visible by reading the source, and it only changes at deploy time. - **A small reference set chooses it** — a list of regions, sites, tiers. Bounded, with a caveat: confirm it is a reference set and not a growing table. A `site` label on nine climbing gyms is fine; a `member` label sourced from the same database is not, and both are "a column in a table". - **Anything else chooses it** — a request, a clock, a remote system, a generated id, an exception's text. Treat it as unbounded, whatever the current count says. Validation does not rescue an unbounded label. A validator that rejects malformed values still admits an unlimited number of well-formed ones; only a validator whose accepted set is an enumeration makes the label bounded. ## What to do with the dimension instead Keeping the dimension usually means changing its *shape* rather than dropping the question. Replace the identifier with the category you would actually group by, replace the concrete path with the route template, replace the message with a class. When the per-event detail genuinely matters, carry it on a signal whose store does not fold it into identity, and keep the metric as the thing that answers "how many, how often, how slow" across categories.

  • A colleague argues the label is safe because a validator rejects unknown values. Does that settle it?
    Only if the validator's accepted set is an enumeration. A validator that rejects malformed values still admits an unlimited number of well-formed ones — a well-formed identifier, a well-formed path, a well-formed date. Rejecting bad input constrains the shape of the value, not the size of its value set, and it is the size that becomes series.
  • Is a label whose values come from a database column bounded?
    It depends entirely on which column. A reference table of nine sites or four membership tiers is bounded and changes at deploy speed. A column of member or booking identifiers from the same database is unbounded and grows with the business. 'It comes from our own data' is not the test; whether the set is small, enumerable and slow-changing is.
  • How would you spot an unbounded label that is already in production but has not yet caused a problem?
    Look at the rate at which new series appear for that metric rather than its current total. A bounded label reaches a plateau within a day of full traffic; an unbounded one keeps creating series at a rate roughly proportional to traffic and never flattens. Comparing series count on a fresh deploy against one that has been running a week shows the same thing.

saying these in an interview costs you the question

  • Judges a label safe from its current distinct-value count
  • Calls a rotating value safe because few exist at once
  • Puts a raw URL path containing identifiers into a label
  • Uses a full error or exception message as a label value
  • Assumes a validated input is automatically a finite set
  • Measures the value count only in a development environment