A team adds a label for user_id to a Prometheus-style metric so they can slice request latency per user, and their metrics backend falls over within a day. Using the concept of cardinality, explain what went wrong, and contrast how structured logging would have handled the same per-user breakdown safely.
answer
- cardinality = distinct label-value combinations
- each combo = its own stored time series
- Prometheus keeps full series index in memory
- logs = per-event documents, indexed not pre-aggregated
- keep IDs out of metric labels, use templated routes instead
basics
~20 sMetrics systems create a separate stored time series for every unique combination of label values. Adding something like user_id, which has millions of possible values, multiplies the number of series by millions, which overwhelms the system. Logs don't have this problem because each log line is just recorded as its own event, with the user_id as a plain field you search later, not as a new stored series.
solid answer
~60 sIn a metrics system like Prometheus, every unique combination of a metric name and its label values creates a distinct time series that must be stored and indexed independently. A metric like request_latency with labels for method and status_code might have a few dozen combinations; adding a high-cardinality label like user_id, which can take millions of distinct values, multiplies the series count by that many, which is what's meant by a cardinality explosion. This blows up memory and index size in the metrics backend, often to the point of it becoming unresponsive or crashing outright, because time-series databases are optimized for a bounded, relatively stable set of series, not an ever-growing one. Structured logging avoids this because each log line is an independent, self-contained event with fields like user_id attached as plain data, not as a series key; a log backend like Elasticsearch indexes and searches those fields without needing to pre-aggregate every possible value combination into its own persistent time series, so per-user slicing is a query-time operation, not a storage-time explosion.
go deeper
Should grasp that putting something like a user ID into a metric caused too much data to be created, without needing the precise mechanism.
Should be able to define cardinality as distinct label-value combinations and know that each combination becomes its own stored series.
Should explain the write-time-aggregation versus query-time-aggregation trade-off between metrics and logs, and name concrete safe alternatives such as route templating, bucketing, or pushing high-cardinality fields to logs or traces.
Should be establishing org-wide guardrails, such as label-cardinality limits enforced in CI or at the metrics-ingestion layer, so no individual team can accidentally take down a shared metrics backend.
## What cardinality is Cardinality, in the context of metrics, is the number of distinct combinations of label values that occur for a given metric name. A metric like `http_request_duration_seconds` with labels method, roughly 5 values, and status_code, roughly 10 to 20 common values, has a cardinality in the dozens: each combination like method equals post, status_code equals 500 is a distinct time series that the metrics backend, such as Prometheus, stores, indexes, and continuously appends new data points to. This is a deliberate design point of time-series databases like Prometheus: they are built assuming the total number of series is bounded and reasonably stable over time, and they pre-allocate memory structures, in Prometheus's case an in-memory index of every active series kept for fast lookup, proportional to that number. ## The explosion The failure mode is what happens when a label takes on a large, effectively unbounded number of distinct values, most commonly - a user ID, - an email address, - a session token, - a request ID, - or a raw IP address. Adding `user_id` as a label on a metric with, say, 2 million distinct users active over a retention window doesn't add 2 million data points, it adds 2 million entirely new, independently indexed time series, each with its own memory overhead for its label set and its own chunk of storage for its samples over time, even for users who only ever generated one data point. This is the cardinality explosion: the total number of series a metrics backend must track can go from a few thousand to tens of millions almost overnight, which - exhausts memory, since Prometheus keeps its full series index in RAM, - slows down every query across the whole system, not just queries touching the offending metric, since the index itself becomes the bottleneck, - and in the worst case causes out-of-memory crashes or ingestion backpressure that drops data for every metric, not just the misconfigured one. ## Why metrics are built to break this way The reason metrics systems are built this way, rather than being able to handle arbitrary-cardinality labels cheaply, is a direct trade-off for what makes metrics fast: pre-aggregated, indexed time series support extremely fast range queries and dashboards, sub-second queries over months of data, precisely because the storage engine can assume a bounded, well-indexed set of series and doesn't need to scan raw events at query time. That assumption is what breaks when cardinality explodes. ## How structured logging handles the same need Structured logging solves the same underlying need, per-user latency breakdown, through a fundamentally different storage model, which is why it doesn't share this failure mode. A structured log line is a self-contained JSON, or similarly structured, event: fields like `timestamp`, `service`, `user_id`, `latency_ms`, `status_code`. A log backend like Elasticsearch, Loki, or Splunk stores each event as its own document and builds inverted indexes over the field values, but it does not need to pre-create a persistent aggregated series for every distinct `user_id` value the way a metrics time-series database does; it can efficiently search, filter, and aggregate over fields at query time instead of pre-aggregating at write time. Asking what the latency distribution is for a specific user over the last day is answered by filtering and aggregating matching log events on demand, which log-oriented storage is built to handle even when the field being filtered on has millions of distinct values, at the cost of that per-query aggregation being slower and more resource-intensive at query time than a pre-computed metric would be. ## Where the work happens | | Metrics | Logs | |---|---|---| | Aggregation | at write time | at query time | | Query | cheap, but only along axes decided in advance and only safely for low-cardinality dimensions | flexible for any field including high-cardinality ones, but comparatively more expensive at scale | The trade-off, then, is not that logs are simply better at this, but that logs and metrics sit at different points on the write-cost versus query-cost spectrum: metrics do the aggregation work at write time, cheap to query later but only along axes decided in advance and only safely for low-cardinality dimensions, while logs defer aggregation to query time, flexible for any field including high-cardinality ones, but comparatively more expensive to query at scale and not well suited to the kind of always-on alerting dashboards metrics are built for. Distributed tracing spans occupy a related middle ground, carrying high-cardinality attributes per span without exploding a series count, because a trace store also indexes rather than pre-aggregates. ## What it looks like in production Common failure signals in production: - a Prometheus instance whose memory usage grows unboundedly and correlates with new users or new request IDs appearing, rather than with traffic volume; - queries that used to return in milliseconds suddenly taking seconds because the label index has bloated; - and, notoriously, any label sourced from something inherently unbounded, like a raw URL path with embedded IDs instead of a templated route used as a label, which is one of the most common real-world causes of this exact incident. The fix is almost always the same: keep genuinely high-cardinality identifiers, like `user_id`, out of metric labels entirely, and push them into structured logs or trace span attributes instead, where they belong, reserving metric labels for a small, bounded, known set of dimensions.
- What's a safe way to still get per-user latency insight without adding user_id as a metric label?Emit user_id as a field on structured log lines or as a span attribute in a trace, and derive per-user views by querying the log or trace backend, which indexes rather than pre-aggregates. If you need a small number of meaningful buckets for dashboards, you can bucket users into a bounded label, such as customer_tier, free, pro, enterprise, instead of the raw user_id.
- Would using a raw URL path like /orders/48291/items as a metric label cause the same problem as user_id?Yes, and it's actually one of the most common real-world causes of this exact incident; the fix is to label with the route template, such as /orders/:id/items, which normalizes away the embedded ID and keeps cardinality bounded to the number of distinct route shapes.
- Does distributed tracing have the same cardinality problem as metrics?Not in the same way, because trace and span storage indexes attributes per event rather than pre-creating a persistent aggregated series per unique value, so a span can safely carry a high-cardinality attribute like user_id. The cost trade-off shows up elsewhere in tracing, mainly in overall storage volume and query latency at scale, not in an unbounded series-count explosion.
A metrics backend is like a spreadsheet where you pre-create one row for every possible combination of filters you might ever want, updated live. Add a column with millions of unique values and you'd need millions of new rows, which crashes the spreadsheet. A log backend is more like a searchable filing cabinet: you can file a document with any field you want on it and search for a specific value later, without needing to have pre-built a folder for every possible value in advance.
saying these in an interview costs you the question
- Doesn't know what cardinality means in a metrics context
- Suggests fixing a cardinality explosion by just adding more memory to the metrics backend instead of removing the label
- Thinks structured logs would have the same series-explosion problem as metrics
- Can't name a real example of a high-cardinality label, such as user_id, a raw URL path, or an IP address
- Believes metrics and logs are interchangeable for every observability question