Which metric label values count as unbounded, and why is 'there are only a few thousand today' the wrong test?
answer
- Unbounded is not the same as large
- Ask who chooses the value
- A ceiling, not a current count
- Rotating values accumulate as churn
- Validation does not make a set finite
basics
~20 sA metric label is unbounded when nothing caps its value set — identifiers, raw URL paths, error message text, timestamps. Counting today's distinct values measures current traffic, not the ceiling, and misses values that rotate and accumulate over the retention window.
solid answer
~50 sBoundedness is a property of **where the value comes from**, not of how many values exist right now. If the code chooses it — an enum, a branch name, a short reference list that changes only at deploy — it is bounded. If a request, a clock, a generated identifier or an exception's text chooses it, it is unbounded however small it looks today. The canonical offenders are entity identifiers, a raw request path with an id in it, an error message string, a timestamp or duration, and a user-agent string. 'Only a few thousand today' fails for two reasons: it samples a distribution with no ceiling, and it reads an instant while the store bills a window. Values that rotate per deploy or per process look tiny at any moment and accumulate as churn.
code
text · 5 linesUNBOUNDED path="/members/8814/checkins" -> one series per member
BOUNDED route="/members/{id}/checkins" -> one series per route
UNBOUNDED error="upstream timed out after 4.812s" -> one series per message
BOUNDED error_class="upstream_timeout" -> one series per classgo deeper
Recall the classic unsafe label values: entity identifiers, raw URL paths that contain them, error message text, timestamps and user-agent strings. Knowing the list is enough at this level.
Explain that boundedness comes from where the value originates rather than how many values exist today, and give the bounded rewrite for each unsafe example.
Bring up churn: values that rotate per deploy or per process look small at an instant and accumulate across the retention window, which is exactly what a current-count check cannot see.
Set the review rule the organisation uses: a label's value set must be enumerable in code, or come from a small reference set that changes only at deploy time. Anything else needs a named exception with an owner.
## Bounded and unbounded, precisely A **metric label** is a key/value dimension that forms part of a time series' identity, so the set of values a label can take is the set of series that label will create. A label is **bounded** when that value set is finite, small, and fixed by something you control — an enumeration in code, a short reference list, a set of environment names. It is **unbounded** when the value set is decided by something outside your control: user input, a generated identifier, a clock, a free-text message. "Unbounded" is not a synonym for "large". It means *nothing in the system caps it*. That distinction is the whole question, because the numbers are deceptive in both directions. A bounded label with 400 values is fine. An unbounded label with 30 values today is not. ## The canonical sources | Unsafe label value | Why it is unbounded | Bounded form | |---|---|---| | A member, order or session identifier | One value per entity, growing forever | The entity's *type* or *tier* | | A raw request path | One value per identifier embedded in the URL | The route template | | An error or exception message | Includes timings, ids, remote text | A stable error class or code | | A timestamp, date or duration | Every value is new by construction | A bucketed range, or nothing | | A user agent string | Effectively unlimited variants in the wild | A parsed family name | | A generated instance or build identifier | New value per process or per release | Kept, but counted as churn | Every row has the same shape: something that varies *per event* was made part of an identity that is supposed to vary *per category*. A metric label answers "which kind of thing is this?"; per-event detail belongs on a signal that carries it as an attribute rather than as identity. ## Why "there are only a few thousand today" is the wrong test Three separate reasons, and a strong answer gives at least two: 1. **It measures the symptom, not the mechanism.** The number of distinct values today is an observation about current traffic. Boundedness is a property of where the values come from. If the source is unconstrained, today's count is a sample of a distribution with no ceiling, and it tells you nothing about tomorrow's. 2. **It measures an instant, and the store bills a window.** Values that *rotate* look tiny at any moment and accumulate relentlessly. A label carrying a container or process identifier on a fleet of 47 hosts running 12 containers each shows 564 values at any instant; redeploy eleven times a week and it has created several thousand distinct series across a retention window, none of which disappear when the process does. This is **churn**, and it is invisible to any check that reads the current value count. 3. **It is measured on the wrong traffic.** Counts are usually taken in a low-traffic environment or soon after a rollout, before the long tail of real-world values has appeared. The trap is that the label works. Dashboards render, queries return, nothing warns. The count grows monotonically with traffic and time, and the failure arrives as a step change when the store crosses a memory or index threshold — typically weeks after the change that caused it, and usually attached to the wrong deploy in everyone's mental model. ## The test that actually works Ask **who chooses the value**: - **The code chooses it** — a constant, an enum member, a branch label. Bounded. Its ceiling is visible by reading the source, and it only changes at deploy time. - **A small reference set chooses it** — a list of regions, sites, tiers. Bounded, with a caveat: confirm it is a reference set and not a growing table. A `site` label on nine climbing gyms is fine; a `member` label sourced from the same database is not, and both are "a column in a table". - **Anything else chooses it** — a request, a clock, a remote system, a generated id, an exception's text. Treat it as unbounded, whatever the current count says. Validation does not rescue an unbounded label. A validator that rejects malformed values still admits an unlimited number of well-formed ones; only a validator whose accepted set is an enumeration makes the label bounded. ## What to do with the dimension instead Keeping the dimension usually means changing its *shape* rather than dropping the question. Replace the identifier with the category you would actually group by, replace the concrete path with the route template, replace the message with a class. When the per-event detail genuinely matters, carry it on a signal whose store does not fold it into identity, and keep the metric as the thing that answers "how many, how often, how slow" across categories.
- A colleague argues the label is safe because a validator rejects unknown values. Does that settle it?Only if the validator's accepted set is an enumeration. A validator that rejects malformed values still admits an unlimited number of well-formed ones — a well-formed identifier, a well-formed path, a well-formed date. Rejecting bad input constrains the shape of the value, not the size of its value set, and it is the size that becomes series.
- Is a label whose values come from a database column bounded?It depends entirely on which column. A reference table of nine sites or four membership tiers is bounded and changes at deploy speed. A column of member or booking identifiers from the same database is unbounded and grows with the business. 'It comes from our own data' is not the test; whether the set is small, enumerable and slow-changing is.
- How would you spot an unbounded label that is already in production but has not yet caused a problem?Look at the rate at which new series appear for that metric rather than its current total. A bounded label reaches a plateau within a day of full traffic; an unbounded one keeps creating series at a rate roughly proportional to traffic and never flattens. Comparing series count on a fresh deploy against one that has been running a week shows the same thing.
saying these in an interview costs you the question
- Judges a label safe from its current distinct-value count
- Calls a rotating value safe because few exist at once
- Puts a raw URL path containing identifiers into a label
- Uses a full error or exception message as a label value
- Assumes a validated input is automatically a finite set
- Measures the value count only in a development environment