Why does recording a latency distribution create a family of time series while a counter creates one?
answer
- Count stored series, not instruments
- A shape needs more than one number
- One series per boundary, plus two
- Index entry and chunks per series
basics
~20 sA counter or gauge is one time series. A distribution is stored as a family: one cumulative-count series per configured upper boundary, plus one for the observation count and one for their sum — so twelve boundaries means fourteen series.
solid answer
~40 sA counter or a gauge is one metric name and, before any dimensional label, one time series. A distribution is not: to make the recorded shape combinable across processes, the client publishes a **family** — one cumulative-count series for each configured upper boundary, plus one series for the total number of observations and one for the running sum of the values observed. Twelve boundaries therefore means fourteen stored series from a single instrument: a fourteen-fold multiplier applied before any dimension is attached, and every dimension added later multiplies the whole family rather than one series. That matters because a time-series store scales with the number of series, not the number of metric names — each series carries its own index entry, its own in-memory write buffer and its own retained chunks on disk.
code
pseudocode · 12 linesinstrument: checkout_orders (counter)
-> 1 time series
instrument: checkout_latency (distribution, 12 upper boundaries)
-> 14 time series:
12 x observations so far at or below one boundary
1 x total number of observations
1 x running sum of all observed values
# 40 publishing processes, one payment-method dimension with 6 values:
# counter -> 1 * 40 = 240 series
# distribution -> 14 * 6 * 40 = 3,360 seriesgo deeper
Recall that recording a distribution of values is far more expensive to store than counting events, and that the cost lives in how many separate time series it creates rather than in how many lines of instrumentation you wrote.
Explain the family: one cumulative-count series per upper boundary plus a count and a sum, so twelve boundaries becomes fourteen series. Be able to do the arithmetic across publishing processes and dimensions out loud.
Show that you size instrumentation before shipping it. Talk about series count as the unit a store scales on — index entries, write buffers, retained chunks — and about reserving distributions for the operations whose shape someone will genuinely read.
Frame it as a budget owned across teams: a per-service series allowance, defaults that count rather than distribute, and a review step where a new distribution states its expected series count. Without that, cost growth is discovered by an outage rather than by a plan.
## One instrument is not one time series An **instrument** is the thing you add to the code: one call site that records something. A **time series** is what the store keeps: one stream of timestamped values under one identity. Engineers estimating the cost of instrumentation almost always count instruments, and the store bills them for series. For a counter or a gauge the two numbers coincide — one instrument, one series, before any dimensional label is attached — so the habit survives right up to the moment somebody records a distribution, at which point it is off by an order of magnitude. ## Where the extra series come from A distribution has to answer questions about the *shape* of many observations, and a single number per timestamp cannot carry a shape. To make the recorded shape combinable across processes, the client publishes a **family** of series rather than one: - one **cumulative count series per configured upper boundary**, each carrying how many observations so far fell at or below that boundary; - one series carrying the **total number of observations**; - one series carrying the **running sum of all observed values**. So an instrument configured with twelve boundaries is stored as fourteen series. The count and the sum are not overhead — the count gives you throughput and the pair gives you a mean — but the boundary series are where the multiplier lives. A representation that publishes pre-computed quantiles instead behaves the same way structurally: one series per published quantile, plus a count and a sum. | Instrument as written in code | Time series stored (before any label) | |---|---| | Counter | 1 | | Gauge | 1 | | Distribution with B upper boundaries | B + 2 | | Pre-computed quantile set of size Q | Q + 2 | The last column is the multiplier that applies **before** dimensions are involved. Every dimension you attach afterwards multiplies the whole family rather than a single series, so a distribution and a counter do not merely differ in cost, they diverge as the instrumentation grows. ## What series count costs a store Series count, not metric-name count and not sample volume alone, is the number a time-series store scales with: 1. **Index.** Every distinct series needs an entry in whatever structure maps an identity to its data, and that structure is typically held in memory for the recent window so that queries can resolve fast. 2. **Write path.** Each series has its own append buffer or head chunk. Fourteen series means fourteen small in-memory structures churning instead of one, and small chunks compress worse than full ones. 3. **Read path.** A query that spans the family touches every boundary series it needs, so the cost of reading a distribution scales with the family too, not only the cost of writing it. 4. **Retention.** Every series is retained independently, so the multiplier is paid again for every day the data is kept. ## A worked example A seed-catalogue ordering service instrumented forty processes. The checkout path had one counter for orders placed and one distribution for checkout latency, the latter configured with twelve boundaries. Counting instruments, that reads as two things. Counting series, it is `(1 + 14) x 40 = 600` — the identity of the publishing process is itself a dimension, so each process contributes its own complete family. When the team then split checkout latency by payment method with six values, they expected to "add one label". They added `14 x 6 x 40 = 3,360` series where the counter beside it grew to 240, and the ingest tier's memory footprint moved accordingly. Nothing had gone wrong, and nobody had done anything foolish; they had simply budgeted using the wrong unit, and during the incident that followed a 5,400-request-per-second peak the store was already running close to its memory ceiling. ## When the family earns its keep None of this is an argument against distributions — it is an argument for knowing what you are buying. - If the only question you will ever ask is *how often*, a counter answers it at one-fourteenth of the storage, and a distribution adds nothing. - If you need a mean, a distribution's count and sum give it to you, but so do two plain counters, and two counters are far cheaper than a full family. - If you need to know what the slow end of the distribution looks like, there is no cheaper option that is honest: the shape has to be recorded at the time, and one number per timestamp cannot carry it. The practical discipline is to estimate before instrumenting. Multiply the family size by the number of publishing processes, then by every dimension's value count, and see whether the resulting series figure is one you are willing to pay for over the retention window. That arithmetic takes a minute and routinely changes the design — most often by reserving distributions for the handful of operations whose shape anyone will actually look at, and counting everything else.
- If a distribution costs that much, why not publish a mean instead and be done with it?Because a mean cannot show the slow end. A mean is two counters — a running sum and a count — and it is genuinely cheap, so it is the right answer when nobody will look past the typical case. But a mean is unchanged by a small fraction of very slow requests, which is usually the fraction users notice, and no post-hoc query can recover a shape that was never recorded.
- How does the series count change when the same instrument is published by forty processes?It multiplies. The identity of the publishing process is itself a dimension on the series, so each process contributes its own complete family: forty processes publishing a fourteen-series family means 560 series for that one instrument. The counter beside it grows to forty. The gap between the two widens with every process added.
- Are the count and sum series of a distribution redundant, given the boundary series?No, and they are the cheapest part of the family. The count is a plain event tally you would otherwise have instrumented separately, and count with sum yields a mean without touching the boundary series at all. Many dashboards read only those two, which is a good reason to keep them even where the boundary series are trimmed back.
A counter is a single tally on a wall; a distribution is a whole tally sheet — a row for each bracket plus a total and a sum — and every process on the fleet keeps its own sheet.
saying these in an interview costs you the question
- Says one instrument always maps to exactly one time series
- Counts metric names instead of series when estimating storage
- Assumes a distribution costs about the same as a counter
- Calls the count and sum series pure overhead
- Thinks adding a dimension multiplies only one series of the family
- Budgets for ingest volume while ignoring index and retention per series