Why does a metrics bill scale with the number of distinct time series while logs and traces scale with request volume?
answer
- Two different cost axes
- Count series, not requests
- One sample per series per interval
- Logs and traces bill per event
- Tenfold traffic, flat metrics bill
basics
~20 sA metric emits one value per series per interval, so its cost tracks how many distinct series exist rather than how many requests arrive. Logs and traces are produced per request, so their cost tracks traffic directly.
solid answer
~50 sA **time series** is one metric name plus one combination of dimension values. Collection writes one sample per series per interval whether that series saw four requests or four thousand, so the metrics bill is a function of *how many series exist* and *how often they are sampled* — traffic does not appear at all. Logs and traces are emitted per event: one record per log statement, one tree per kept request, so ten times the traffic is ten times the bytes, the ingest and the index. The practical consequence is that the two families respond to different dials. Metrics shrink by reducing distinct dimension combinations or lengthening the interval; logs and traces shrink by emitting less per event or keeping a smaller fraction of events. Keeping a fraction of traces does nothing for a metrics bill, and pruning dimensions does nothing for a log bill.
code
pseudocode · 8 linesseries(metric) = product of distinct values over all its dimensions
renewals{office: 14, outcome: 5, channel: 3} = 210 series
22 metrics x 210 = 4620 series
samples/day = 4620 * (86400 / 15s) = 26,611,200
# note what is absent from every line above:
# the number of renewals servedgo deeper
Recall that a metric costs the same whether it counts four requests or four thousand, because it is emitted on a schedule. Logs and traces are written per event, so they grow with traffic.
Explain the arithmetic: series count is the product of distinct dimension values, samples come from the interval, and traffic appears in neither. Then show that log and trace volume are linear in requests, and name the dial that shrinks each.
Show you have forecast both curves for a real service and pulled the right lever under pressure. Explain why a keep-a-fraction dial cannot rescue a metrics bill, and how you found a dimension that multiplied the series count.
Own telemetry as a cost model with two independent axes, and argue for where each axis is allowed to grow. Be ready to say which growth is a defect to fix and which is a legitimate consequence of the product succeeding.
## Two meters, wired to different things The question that decides a telemetry budget is not "how much data do we produce" but "what is the quantity that, when it doubles, doubles the bill". For the three signals that quantity is not the same. - **Metrics are billed by breadth and by clock.** Storage and ingest are proportional to `number of distinct series x samples per series`, and samples per series is set by the collection interval. Requests do not appear in that expression. A counter incremented once a minute and a counter incremented eighty thousand times a minute produce exactly the same number of stored samples. - **Logs and traces are billed by events.** One record per emitted log statement, one span per instrumented operation, one tree per kept request. Every one of those is proportional to traffic, and so is everything downstream of it: network egress, ingest, full-text indexing, query scan volume. ## The arithmetic, on one service A municipal parking-permit service handles 4,180 renewals an hour and is about to grow tenfold after a policy change forces every resident to re-register. **Metrics today.** Say 22 metrics, each carrying three dimensions: 14 offices, 5 outcomes and 3 payment channels. That is `14 x 5 x 3 = 210` series per metric, or about 4,620 series in total. At a 15-second interval each series produces `86400 / 15 = 5760` samples a day, so the service writes roughly 26.6 million samples a day. **Metrics after tenfold growth.** Still 4,620 series. Still 26.6 million samples a day. Nothing in the expression moved, because no dimension gained a value. The bill is flat. **Logs today.** Three records per renewal at roughly 1.4 KB each: `100,320 renewals x 3 x 1.4 KB` is about 421 MB a day. **Logs after growth.** About 4.2 GB a day — plus the index built over it and the scan cost of every query that reads it. **Traces today.** Nine operations per renewal with one renewal in a hundred kept: roughly 9,029 spans a day. **After growth:** roughly 90,288. | Signal | Cost proportional to | Tenfold traffic, same shape | The dial that reduces it | |---|---|---|---| | Metrics | distinct series x samples per series | roughly unchanged | fewer dimension combinations; longer interval | | Logs | events emitted x bytes per event | roughly tenfold | fewer or thinner records; emit less at healthy paths | | Traces | requests kept x operations per request | roughly tenfold | keep a smaller fraction; instrument fewer operations | ## The two failure modes look nothing alike Because the cost axes differ, the way each signal blows up differs too, and confusing the two leads teams to pull the wrong lever. 1. **Metrics fail as a step change, driven by breadth.** Add one dimension whose value set is large or unbounded — a permit identifier, a request path with an id embedded in it, a version string that changes every deploy — and the series count is *multiplied*, not incremented. Adding a fourth payment channel above takes the service from 210 to 280 series per metric, which is fine; adding a dimension with 40,000 distinct values takes it to eight million, which is not. Note the vocabulary trap here: how many values a single dimension takes and how many series a metric produces in total are different numbers, and they differ by orders of magnitude. 2. **Logs and traces fail as a smooth slope, driven by traffic.** Nothing is wrong with any one record; there are simply more of them every quarter, and the slope is set by product growth rather than by an engineering mistake. This failure is easy to forecast and hard to notice, because no single change causes it. ## What follows for a growing service - **Forecast the two curves separately.** A capacity plan that projects "telemetry cost" as one line will be wrong in both directions at once. - **Do not reach for a keep-a-fraction dial to fix a metrics bill.** Aggregates are already the summary of everything; keeping a fraction of them just adds error to the numbers you alert on. - **Do not reach for dimension pruning to fix a log bill.** Fields on a record are nearly free relative to the record itself; the record count is what costs. - **Watch the derivative, not the total.** A metrics bill that rises while traffic is flat means the shape changed — someone added a dimension or a dimension gained values. A log bill that rises in step with traffic is working as designed, and the conversation is about resolution, not about a defect. - **Remember that ingest and query are separate costs.** Doubling the records doubles what you pay to accept them, to index them, and to read them back in every query that touches the window.
- Traffic has been flat for a month but the metrics volume doubled overnight. What most likely happened?The shape changed rather than the load. Almost always a new dimension was added to a metric, or an existing dimension started taking values it had not taken before — a new deploy label, a pod or container identity, an error string used as a dimension value. Because series count is a product, one new dimension multiplies the total. The fix is to bound or remove that dimension, not to lengthen the interval.
- Does doubling the collection interval halve the metrics bill?It halves samples per series, which helps where you pay per sample ingested or stored, and it costs you resolution — short spikes disappear between samples. It does not touch the series count, which in many stores drives the dominant cost through per-series indexing and memory. If breadth is the problem, the interval is the wrong dial.
- Which signals respond to keeping only a fraction of events, and which does not?Traces respond directly, and logs respond if you thin what is emitted on healthy paths. Metrics do not: an aggregate already summarises every event, so keeping a fraction of the aggregates does not reduce the series count and simply makes the numbers you alert on noisier.
Metrics cost like meters bolted to a wall: you pay per meter installed, not per person walking past. Logs and traces cost like a ticket for every person who walks past.
saying these in an interview costs you the question
- Assumes ten times the traffic means ten times the metrics bill
- Counts metrics cost by number of metric names alone
- Suggests keeping a fraction of metrics to cut the bill
- Treats log cost as storage only, ignoring ingest and query
- Believes adding one dimension adds one series