skip to content

Metrics Concepts

How metrics work regardless of vendor: types, labels, histogram math, and collection models. This is where interviewers separate people who read dashboards from people who can instrument a service correctly.

on this pageshow

questions

19

What identifies one time series in a dimensional metrics system, and how does adding a metric label change the total count?

level: juniorimportance: must knowfreq 72%

answer

  1. Identity, not decoration
  2. Name plus the whole label set
  3. Labels multiply, they never add
  4. Per-instance labels multiply too
  5. The product is an upper bound

basics

~20 s

A time series is the metric name plus its complete label set; change any label value and it is a different series. Adding a label multiplies the series count by that label's distinct values rather than adding to it.

solid answer

~50 s

In a dimensional metrics system a series' identity is the pair *(metric name, complete label set)*. Every distinct combination of label key/value pairs is its own stored series with its own history of samples, so there is no such thing as one series carrying two values of a label. That makes the arithmetic multiplicative: a counter with 6 door values and 4 result values is 24 series, and adding a third label with 5 values makes it 120, not 29. The product is an upper bound — only combinations that actually occur become series — but it is the number to reason with. The multiplication also covers the labels you did not choose: on 47 hosts running 12 containers each, the collection tier's per-instance labels already multiply everything by 564 before an application label is counted.

code

text · 4 lines
text
gym_checkin_requests_total{site="bristol",door="turnstile-a",result="ok"} 41273
gym_checkin_requests_total{site="bristol",door="turnstile-a",result="denied"} 118
gym_checkin_requests_total{site="bristol",door="turnstile-b",result="ok"} 29604
gym_checkin_requests_total{site="leeds",door="turnstile-a",result="ok"} 18842

go deeper

for a junior

Be ready to state that a time series is the metric name plus its full label set, and that changing any label value produces a different series rather than a variant of the same one.

for a middle

Expect to do the multiplication aloud: multiply the distinct value counts of every label, then explain why the result is an upper bound rather than the exact count.

for a senior

Show that you count the labels the collection tier adds — host, container, environment — because those multiply an application's series across the whole estate before anyone opens a dashboard.

for a principal

Own the rule teams apply before adding a label: what is its ceiling, what is it multiplied by, and will anyone ever filter or group by it. Detail nobody slices by does not earn a dimension.

## The unit a metrics store actually keeps A dimensional metrics system does not store "a metric". It stores **series**, and one series is the pair *(metric name, complete label set)* together with the run of timestamped samples written against that pair. A **metric label** here means a key/value dimension attached to the measurement — `site="bristol"`, `result="denied"` — not a log-stream label and not an orchestrator object label that happens to share the word. The word doing the work is *complete*. The label set is the series' identity, the way a compound primary key is a row's identity. `door_unlocks_total{site="bristol"}` and `door_unlocks_total{site="leeds"}` are not one metric seen two ways; they are two independent series that happen to share a name. Nothing in the store links them. A dashboard adds them together only because a query told it to. Some stores model the metric name as just another reserved label; the identity rule is unchanged either way. Two consequences fall straight out of that: - **There is no "the same series with a different value".** Changing a label value does not update a series, it starts a different one. - **Adding or renaming a label forks history.** The old label set stops receiving samples and a new one begins, so a range query spanning the change sees one series end and another start rather than a continuous line. ## The arithmetic: labels multiply, they do not add Because identity is the whole combination, the number of series a metric can produce is the **product** of the sizes of its label value sets: series <= |values(label_1)| x |values(label_2)| x ... x |values(label_n)| Take a climbing-gym membership platform's turnstile counter, adding one label at a time: | Labels on the metric | Distinct values of the new label | Upper-bound series | |---|---|---| | `site` | 9 | 9 | | `+ door` | 6 | 54 | | `+ result` | 4 | 216 | | `+ membership_tier` | 5 | 1,080 | Each new label multiplies; nothing adds. The instinct that "one more label is one more thing to store" is the most common wrong model in this whole area, and by the fourth label it is wrong by orders of magnitude. Two refinements separate a good answer from a recited one: 1. **The product is an upper bound, not a count.** Only combinations that actually occur in traffic become series. If each `door` exists at exactly one `site`, the real total is far below 1,080 — real label sets are sparse. That is a reason not to panic at a large product, never a reason to skip computing it. 2. **Sparsity only saves you while every factor is finite.** It reduces a product of bounded factors. It does nothing about a factor that grows with traffic, because every new value there is a combination that has now genuinely occurred. ## The labels you did not choose The application picks some labels. The collection tier attaches more, to record where a sample came from: a host or node identifier, a container or replica identifier, a namespace, an environment. Those are multiplied against the application's product across the entire estate. On a fleet of 47 hosts each running 12 containers, every series an application defines exists 564 times before a single application label is counted. The 216-series turnstile counter above is 121,824 series estate-wide. That estate number — not the per-process one the developer sees locally — is what reaches the store, and it is the number that matters when a telemetry budget is halved and someone has to say which metrics get cut. The same multiplication is why a label that looks harmless in one service is not harmless as a convention. A team-wide "add `build_id` to everything" decision multiplies every metric in the estate at once. ## Where the model gets misapplied - **Counting label *keys* instead of label *values*.** Four labels is not four times the cost; it is the product of their value counts. - **Assuming aggregation makes the count irrelevant.** A query that sums away a label still has to match and read every series it sums, so a high series count is a query cost even when the result has one line. - **Measuring in a development environment.** One process, one host, one environment collapses three of the multipliers to 1 and hides the real figure. ## What to actually do with this Before adding a label, ask three questions in order: 1. **What is this label's ceiling** — not how many values exist today, but how many it can take? 2. **What is it multiplied by** — the other labels already on the metric, and the per-instance labels the collection tier will attach? 3. **Will anyone ever filter or group by it?** A label nobody queries pays the full multiplication and returns nothing for it. The third question prunes more labels in practice than the other two combined. Teams add labels because the value happened to be in scope at the call site, not because a query needs it — and the store charges the product either way.

  • The product of a metric's label values is 12,000 but the store reports 900 series. What explains the gap, and can you rely on it?
    Real label sets are sparse: only combinations that actually occur become series, and most label pairs are correlated — a door exists at one site, a tier only appears on certain routes. You can note the gap, but you cannot rely on it. Sparsity is a property of current traffic, so it shrinks a product of bounded factors and does nothing about a factor that grows with traffic.
  • What happens to existing dashboards when you add a new label to a metric that is already in production?
    Every existing series stops receiving samples and a new set of series begins, because identity changed. A range query spanning the deploy sees one set end and another start rather than a continuous line. Queries that aggregate the label away recover once the range is entirely on one side; anything pinned to an exact label set, including recorded aggregates and alert rules, has to be updated.

The label set is a compound primary key: the store does not keep one row per metric, it keeps one row per distinct combination of key columns.

saying these in an interview costs you the question

  • Says adding a label adds one series rather than multiplying
  • Thinks the metric name alone identifies what is stored
  • Treats two label values as variants of a single series
  • Ignores the per-instance labels the collection tier attaches
  • Counts label keys instead of the product of their values
open as a page

What makes a metric a counter rather than a gauge, and why is a counter read as a rate over a window?

level: juniorimportance: must knowfreq 78%

basics

~20 s

A counter only ever increases and returns to zero when its publishing process restarts; a gauge is a level that moves either way. Because a counter's absolute value depends on process uptime, you read its rate over a window.

open as a page

Why can't you average the p99 latencies of ten replicas to get the fleet's p99?

level: middleimportance: must knowfreq 68%

basics

~20 s

A percentile is an order statistic over a set of observations, not a quantity that adds. Averaging per-replica p99s ignores how much traffic each served and describes no real request. Merge the underlying bucket counts first, then estimate the quantile once.

open as a page

What does a pull-based metrics collector get for free, and what does it demand of every target?

level: middleimportance: must knowfreq 74%

basics

~20 s

A pull-based collector already holds the list of what should exist, so every failed poll is itself a liveness signal, and one setting fixes the whole fleet's sampling resolution. In exchange each target must stay routable and answerable at any instant.

open as a page

Under the RED checklist, how do you instrument a request-serving service, and what counts as an error?

level: middleimportance: must knowfreq 64%

basics

~20 s

RED means emitting three things for a request-serving component: request rate, failed requests, and a duration distribution. Rate and errors come from one counter carrying an outcome dimension, and "error" must be defined in writing before the count means anything.

open as a page

What does push-based metrics collection buy you, and what does the backend lose by never asking?

level: juniorimportance: should knowfreq 56%

basics

~20 s

Push needs no inbound route and no discovery - a sender exists the moment it sends, which is what makes it work from behind address translation and from very short-lived processes. The cost: silence is ambiguous, so died and idle look identical.

open as a page

How does a metrics backend estimate a p99 latency from histogram bucket counts, and how large is the error?

level: middleimportance: should knowfreq 58%

basics

~20 s

A backend stores only per-bucket counts, never individual measurements. It locates the bucket holding the 99th-percentile rank and interpolates linearly across it, so the reported value can be wrong by the entire width of that bucket.

open as a page

Which metric label values count as unbounded, and why is 'there are only a few thousand today' the wrong test?

level: middleimportance: should knowfreq 58%

basics

~20 s

A metric label is unbounded when nothing caps its value set — identifiers, raw URL paths, error message text, timestamps. Counting today's distinct values measures current traffic, not the ceiling, and misses values that rotate and accumulate over the retention window.

open as a page

Which aggregations across a fleet of per-process gauges are meaningful, and why is an average of averages wrong?

level: middleimportance: should knowfreq 48%

basics

~20 s

It depends on the quantity. Gauges of things that add up — queued items, bytes in use — can be summed across a fleet; ratios and percentages cannot. A mean of per-process ratios ignores that each process handled a different volume.

open as a page

How do you choose histogram bucket boundaries for latency, and what breaks when nearly all requests land in one bucket?

level: seniorimportance: should knowfreq 44%

basics

~20 s

Boundaries must bracket the decision range — below the fastest realistic response, above the timeout, densest at the thresholds you act on. If one bucket holds nearly all traffic, every percentile is interpolated inside it and moves as one line.

open as a page

How do you pick a metric's type at instrumentation time, and what can you never recover afterwards?

level: seniorimportance: should knowfreq 52%

basics

~20 s

Pick from the question you will ask later: counters answer how often, gauges answer how much now, distributions answer what the slow end looks like. Nothing unrecorded can be recovered, so fixing a wrong type helps only going forward.

open as a page

In the USE method, what is saturation for a thread pool or a work queue, and why is utilization alone a weak overload signal?

level: seniorimportance: should knowfreq 45%

basics

~20 s

Saturation is work a resource has accepted but cannot serve yet: queued tasks, queue depth, wait time before starting, rejections. Utilization is capped at one hundred percent, so it stops resolving exactly when pressure starts growing.

open as a page

How do you bound metric label cardinality across a fleet, and where between instrumentation and storage should each control sit?

level: principalimportance: should knowfreq 46%

basics

~20 s

Allowlist a label's values with a catch-all for the rest, rewrite or drop it before storage, normalise continuous values into categories, or move the dimension onto traces. Emit-time fixes are correct but slow; central rules hide the cost.

open as a page

How do you scale pull-based and push-based metrics collection in one estate, and where is the line between them?

level: principalimportance: should knowfreq 46%

basics

~20 s

Polling scales by budgeting the collection interval against target count and work per target, then sharding the collection tier per network domain. Sending scales by batching, bounded queues and an explicit drop policy. Draw the line by reachability and lifetime.

open as a page

Why does recording a latency distribution create a family of time series while a counter creates one?

level: middleimportance: nice to knowfreq 30%

basics

~20 s

A counter or gauge is one time series. A distribution is stored as a family: one cumulative-count series per configured upper boundary, plus one for the observation count and one for their sum — so twelve boundaries means fourteen series.

open as a page

Why export histogram bucket counts when a service can compute and export its own p99?

level: seniorimportance: nice to knowfreq 27%

basics

~20 s

Counts are additive: any set of processes and any time range can be pooled and the quantile computed once at query time. An exported quantile is a finished number for one process and one window that nothing downstream can recombine.

open as a page

What does one extra time series cost a metrics store, and why does the cost persist after you stop emitting it?

level: seniorimportance: nice to knowfreq 34%

basics

~20 s

Each series carries fixed overhead — index entries per label pair, identity held while active, a write buffer — regardless of sample count. Removing the label stops new series; existing ones stay indexed until retention passes.

open as a page

How do you get metrics out of a batch job that exits before any collector could poll it?

level: seniorimportance: nice to knowfreq 42%

basics

~20 s

Three options: push to an intermediary that holds the last value for later collection, push straight to a receive-capable backend before exiting, or have a long-lived process report on the job's behalf. Each trades away either freshness or correct attribution.

open as a page

Neither RED nor USE fits a batch job or an event consumer. What do you instrument on those components instead?

level: seniorimportance: nice to knowfreq 32%

basics

~20 s

Instrument what the component owes someone: for a batch job, the time of its last successful completion; for an event consumer, the age of the oldest unprocessed item; for both, items processed, failed and skipped.

open as a page