Token-cost metrics labelled by model, prompt_version and tenant exploded to 180k Prometheus series — how do you fix it?
answer
- series count is a product, not a sum
- some label domains never stop growing
- histogram buckets multiply again
- two stores, two jobs
- alerting needs speed, chargeback needs detail
basics
~20 sSplit the pipeline. Keep time-series metrics to bounded dimensions you alert on — model, effort, cache-hit bucket, feature, status — and move unbounded ones like tenant id and prompt version into per-call usage events in a warehouse or log store where you group at query time.
solid answer
~50 sSeries count is the product of distinct label-value combinations, so tenant (thousands, growing) times prompt_version (grows forever, and every rollout creates a fresh fan of series) times model times each counter is exactly the shape that melts a TSDB — and histogram buckets multiply it again. The fix is two stores with two jobs. Metrics carry only low-cardinality, operationally-actionable labels and exist to alert fast; per-call usage events carry the full attribution — tenant, prompt version, session, request id — into a columnar store or log pipeline, where cost per tenant is a nightly rollup rather than a live gauge. Where you need tenant visibility in metrics, label by tenant *tier* or keep an explicit top-N allowlist and bucket the rest as "other". Enforce the label allowlist in the metrics wrapper so a well-meaning code change cannot add a label, and drop or aggregate stray labels at the collector as a backstop.
code
json · 22 lines{
"event": "llm.usage",
"ts": "2026-05-14T09:12:44Z",
"model": "reasoning-large",
"effort": "high",
"status": "ok",
"tokens": {
"input": 2140,
"cache_read": 30512,
"cache_write": 0,
"reasoning": 7880,
"output": 1610
},
"rate_card": "2026-04-01",
"attribution": {
"tenant": "t_8391",
"plan": "scale",
"feature": "variance_explainer",
"prompt_version": "ve-2026-05-13b",
"invocation_id": "inv_01HXZ"
}
}go deeper
Know that each unique combination of label values creates its own time series, so labels with many possible values — user ids, request ids — must not go on metrics.
Compute the worst-case series count from the label domains, and explain the split between low-cardinality metrics for alerting and per-call events for attribution.
Show the operational fix end to end: label allowlist in the emit path, bucketing or top-N for tenants, counters over histograms, collector-side drops, and rollups for chargeback.
Own the standard — a documented cardinality budget per metric, new labels reviewed like schema changes, and a clear statement of which questions the metrics pipeline is allowed to answer versus the analytics store.
## Where 180,000 series comes from In a Prometheus-style time-series database, one series exists per unique combination of metric name and label values. Multiply the label domains and you have your answer: 2,000 tenants x 30 live prompt versions x 3 models is 180,000 series per counter. Add a second counter (output tokens, reasoning tokens, cache-read tokens) and you multiply again; make any of them a histogram and you multiply by the bucket count. Two properties make this worse than a big number: - **Unboundedness.** Tenants and prompt versions are not a fixed set. Every rollout mints new label values, and old ones keep their series alive through the retention window — this is churn, and it hurts the index harder than steady-state size. - **Cost is per series, not per sample.** Memory, index size, compaction and query fan-out all scale with series count, so a high-cardinality metric degrades queries and alert evaluation for everything else on the same instance. ## The two-store split The durable fix is to stop asking one system to do two jobs. **Metrics — for alerting and operations.** Low cardinality, seconds-fresh, small. Reasonable labels: model, effort or thinking setting, cache-hit bucket, feature or endpoint, call status, environment, region. Each has a handful of values, all of them things you would page on or slice a graph by. **Usage events — for attribution and finance.** One structured record per model call: token counters, rate-card version, latency, and the full attribution set (tenant, plan, prompt version, feature, session, request and trace ids). Land these in a columnar warehouse or a log/analytics store where cardinality is nearly free and any grouping is a query. Per-tenant cost, per-prompt-version cost and chargeback become scheduled rollups. They do not need to be live: an invoice-grade number that is an hour stale is fine; a paging signal that is an hour stale is not. ## Techniques for keeping metrics narrow - **Label allowlist in the emit path.** Metrics go through one helper that accepts only the approved label set. This is the control that actually holds, because it stops the next well-intentioned commit adding a user id. - **Bucket instead of enumerate.** Replace tenant id with tenant tier or plan; replace prompt version with a changed-recently flag or a small current/previous pair; replace exact cost with cost buckets. - **Top-N allowlist.** If a few whales genuinely need their own lines, maintain an explicit list of them and label everything else "other", refreshed on a slow cadence. - **Counters over histograms** on any dimension you cannot fully bound; get distributions from the event store instead. - **Collector-side filtering.** Drop or aggregate labels in the collection layer as a safety net, so a bad deploy degrades a dashboard rather than the TSDB. - **Recording rules and pre-aggregation** for the handful of series dashboards actually query. - **Exemplars or trace links** from a metric to representative calls, which gives you the drill-down without materialising the cardinality. ## Choosing which side a dimension belongs on Ask two questions. *Would I ever alert on it?* If not, it does not need to be a metric label. *Is its value set bounded and slow-changing?* If not, it belongs in events. Model passes both. Tenant fails the second badly. Prompt version fails the second and usually the first — you want it for regression analysis and cost-per-version comparison, both of which are analytical queries, not alerts. ## What you give up, and why it is acceptable You lose sub-minute per-tenant cost graphs. In exchange you get a TSDB that stays fast, alerting that evaluates on time, and per-tenant attribution that is *more* flexible than labels ever were, because you can group by plan, region, prompt version and feature in the same query without pre-declaring the combination. If a single tenant needs real-time enforcement — a hard spend cap — implement that as a counter in the gateway keyed by that tenant's quota bucket, not as a global metric dimension. ## Reviewing it Give each metric a documented cardinality budget, compute the worst case from the label domains before shipping, and treat a new label in a diff as a review item with the same weight as a schema change. Track series count per metric as an operational signal so a rollout that multiplies it is visible immediately.
- Marketing wants a live per-tenant spend graph. What do you offer instead?Offer a near-real-time view from the event store — a few minutes stale, grouped by tenant, plan and feature — plus metric-level alerting on aggregate burn and on gateway-enforced per-tenant quotas. Live per-tenant series in the TSDB buy seconds of freshness for a permanent cardinality cost, and the business question is almost always answered by an hourly rollup.
- Why is prompt version an especially damaging label rather than just a large one?Because it is churny as well as unbounded: every rollout creates a whole new fan of series across every other label, while the previous fan lingers for the retention window. Churn stresses the index and compaction more than a stable set of the same size, so cardinality spikes with deploy frequency rather than with traffic.
- Where does tenant-level cost enforcement live if not in the metrics pipeline?In the request path — a gateway or middleware that increments a per-tenant spend counter in a fast store and refuses, queues or degrades calls once the cap is hit. Observability tells you what happened; enforcement has to be synchronous with the call. Keeping them separate also means a metrics outage cannot disable your spend guardrail.
saying these in an interview costs you the question
- Adding tenant id or request id as a metric label for convenience
- Believing cardinality cost scales with sample rate rather than series count
- Using histograms on an unbounded dimension to get cost percentiles
- Deleting old prompt versions from dashboards and assuming series disappear at once
- Treating a warehouse rollup as too slow for finance-grade attribution