skip to content

What does one extra time series cost a metrics store, and why does the cost persist after you stop emitting it?

level: seniorimportance: nice to knowfreq 34%

answer

  1. Two bills, not one
  2. Fixed cost per identity, not per value
  3. Samples compress, identities do not
  4. Memory frees first, the index last
  5. Recovery is bounded by retention

basics

~20 s

Each series carries fixed overhead — index entries per label pair, identity held while active, a write buffer — regardless of sample count. Removing the label stops new series; existing ones stay indexed until retention passes.

solid answer

~40 s

A time-series store bills twice. The **per-series** bill covers index entries for every label key/value pair, the series' identity held in memory while it is active, an in-progress write buffer per actively written series, and the query-time work of resolving matchers over the index. The **per-sample** bill covers each appended value — and it is the cheap half, because timestamps and values compress extremely well against their predecessors. So two metrics of 50,000 series each, one collected every 15 seconds and one every 5 minutes, differ by 20x in samples and far less in cost. Removing the offending label only stops *new* series. Memory is reclaimed once series go stale, but the index entries and the query cost persist for the whole retention window.

go deeper

for a junior

Know that the number of distinct series, not the number of samples written, is what makes a metrics store expensive to run.

for a middle

Describe the per-series costs — index entries per label pair, identity held while the series is active, an in-progress write buffer — and contrast them with how cheaply steady samples compress.

for a senior

Explain the non-recovery: a series that has existed stays in the index for the whole retention window, so queries covering that period keep paying long after emission stopped.

for a principal

Frame the bill for a capacity plan: per-series overhead is a fixed cost bought for a retention window, sample volume is the marginal one. A plan that models only ingest rate is modelling the wrong driver.

## Two different bills A time-series store charges you twice, and the two lines behave nothing alike: - A **per-series** cost, paid once for every distinct *(metric name, label set)* that exists — and paid for the whole time that series is within the retention window, whether or not anything is still being written to it. - A **per-sample** cost, paid for each timestamped value appended to a series. Almost every surprise in this area comes from reasoning about the second bill when the first one is what moved. Sample volume scales with the collection interval, which is a number people tune deliberately and watch closely. Series count scales with a label nobody thought about. ## What one extra series actually costs | Cost | What it is | Scales with | |---|---|---| | Index entries | The store maps each label key/value pair to the series carrying it, so a query can find series by matcher rather than by scanning. Every new series adds itself to one posting list per label pair. | Series, and the number of labels each carries | | Resident series metadata | While a series is active the store keeps its identity — the full label set, in string form — reachable in memory. | Active series | | An open write buffer per series | Samples are accumulated per series before being written out compressed, so every actively-written series holds a small in-progress buffer. | Active series | | Query-time matching | A matcher must be resolved against the index and the matching set intersected, before any sample is read. | Series matching the query | | Retention-window footprint | The identity and index entries persist for as long as the data does. | Series ever created x retention | The per-series line items are *fixed* — a series costs roughly the same whether it receives one sample an hour or one every five seconds, because what it costs is identity, not data. ## Why the sample side is the cheap half Samples on a time series are highly regular: a timestamp that advances by a near-constant step and a value that usually changes very little. Stores exploit exactly that, encoding timestamps as deltas and values against the previous value, so a steady series costs remarkably little per sample — frequently a small fraction of the bytes the raw pair would take. The practical consequence is the one interviewers are listening for. Two metrics, each with 50,000 series, one collected every 15 seconds and one every 5 minutes, differ by 20x in samples written and are much closer than that in what they cost the store, because the dominant term is identical. Halving a collection interval is a modest saving. Halving a series count is a large one. ## Why the cost does not stop when you stop emitting This is the part that catches people out, and it has three layers: 1. **The series still exists.** Removing the offending label from the application stops *new* series being created. Every series already created still holds its index entries and its samples until the retention window rolls past them. Emission stops; the footprint does not. 2. **Queries still pay.** Any query whose time range overlaps the period when those series were live still has to resolve its matchers against an index containing them. A dashboard looking at "yesterday" is slow for as long as yesterday includes the incident. 3. **Only the memory recovers quickly.** Once a series receives no further samples it goes stale, the store stops treating it as active, and the resident metadata and write buffer for it are reclaimed. That is genuine and usually the fastest relief you get. It does not touch the on-disk index or the query cost. So "we deployed the fix an hour ago and it is still slow" is the expected outcome, not evidence the fix failed. Recovery is bounded below by the retention window, unless the store offers targeted deletion of the affected series and you are willing to use it. ## What this means operationally - **Capacity plans that model only ingest rate model the wrong thing.** Series count is the input that predicts memory and index size. - **Churn is a first-class metric.** The rate at which *new* series appear predicts a problem the instantaneous total will only show once it has arrived. - **Cutting a telemetry budget means cutting series, not resolution.** Faced with a halved budget, dropping one high-cardinality label usually beats doubling every collection interval, and costs the humans far less. - **A series limit is a backstop, not a design.** Refusing writes past a threshold protects the store and loses data; it is what you configure so a mistake degrades instead of destroying, not what you rely on to stay within budget.

  • Two metrics each have 50,000 series; one is collected every 15 seconds and the other every 5 minutes. Which is the bigger problem?
    They are much closer than the 20x difference in samples suggests, because the dominant term — index entries and per-series identity — is identical for both. The faster one costs more in written and stored samples, but samples compress well and scale linearly. If you have to save money, halving the series count of either beats halving the collection frequency of both.
  • You removed the bad label an hour ago and the store is still struggling. What can you actually do to speed recovery?
    Memory relief arrives on its own as the dead series go stale and stop counting as active, which is usually the first improvement you see. Beyond that, restrict dashboards and alert rules to time ranges that exclude the affected period so queries stop expanding matchers over those series, and if the store supports targeted deletion of series matching a label matcher, use it rather than waiting out the retention window.
  • Which signal warns you earliest that a metric's series count is about to become a problem?
    The rate at which new series appear, broken down by metric name and by emitting service. An instantaneous total tells you when the problem has already arrived; a creation rate that steps up at a deploy tells you which change caused it while the absolute numbers are still small enough to fix calmly.

A series is a subscription, not a message. Cancelling the traffic stops the messages; the subscription keeps billing until its term runs out.

saying these in an interview costs you the question

  • Assumes cost is driven by sample volume rather than series count
  • Expects instant recovery once the label is removed
  • Confuses collection interval with cardinality as the cost driver
  • Thinks dead series are free because nothing is written to them
  • Believes shortening a query's range reduces stored series