skip to content

What do metrics, logs and traces each capture, and what does each one's data shape make it unable to answer?

level: juniorimportance: must knowfreq 84%

answer

  1. Three signals, three data shapes
  2. Aggregate, discrete record, causal tree
  3. Aggregation discards the individual requests
  4. Only one shape crosses process boundaries
  5. Shape decides the answerable question

basics

~20 s

Metrics are numbers aggregated over an interval, so they show trends but never an individual request. Logs are discrete timestamped records of single events. Traces are causally linked trees showing one request's path across services.

solid answer

~50 s

Each signal has a different data shape, and the shape decides what it can answer. - A **metric** is a value aggregated over an interval and a set of dimensions. It answers *how many, how often, how slow* at a fixed cost, but the individual requests are gone, so it can never say which one was slow. - A **log** is a discrete timestamped record of one thing that happened, carrying whatever fields the author chose. It answers *what exactly happened here*, with unbounded detail, but it has no built-in link to the other processes that handled the same request. - A **trace** is a tree of timed operations for one request, stitched together by a propagated identifier, so it answers *where did this request's time go across services*. It is the most expensive per request, so usually only a fraction is kept.

code

text · 12 lines
text
# aggregate: one value per series per interval, no per-request detail
renewals_total{office="north",outcome="ok"} 41822

# record: one discrete event, with whatever fields were included
{"ts":"2026-03-11T09:14:22Z","event":"renewal","plate":"KX21-ZTM","ms":587,"outcome":"error"}

# causal tree: one request, linked operations across processes
renew 587ms
  |- validate-plate    41ms
  |- charge-card      411ms
  |    |- rate-lookup 309ms
  |- issue-permit      52ms

go deeper

for a junior

Be ready to say in one sentence each what a metric, a log record and a trace hold, and to give one question each is good at. Knowing that an aggregate cannot name an individual request already puts you ahead at this level.

for a middle

Explain the mechanics behind the shapes: a series is a name plus a dimension combination sampled on an interval, a record is written per event, and a trace only exists because context is propagated across hops. Say what each shape discards.

for a senior

Show you have used all three under pressure. Talk about the aggregate covering all traffic while trace data covers only what was kept, about correlating records across services in practice, and about the questions you could not answer because a dimension was never emitted.

for a principal

Own the argument that the three shapes are complements rather than alternatives, and be able to say what a team loses concretely when it under-invests in one of them. Frame it as which classes of question the estate can afford to answer.

## Three shapes, not three products Every monitoring stack ships some version of the same three data shapes. The shape, not the tool, decides which questions are cheap, which are expensive and which are structurally impossible. - **A metric is an aggregate.** It is a number attached to a name and a set of dimensions, emitted or scraped on a fixed interval. One name plus one combination of dimension values is one **time series**: a single stream of timestamped numbers. Whether four requests or four thousand hit that combination during the interval, what gets stored is identical in shape — a timestamp and a value. - **A log is a discrete record.** One entry for one thing that happened, written when it happened, carrying whatever fields the author put in it. Nothing is combined with anything else, so the detail is unbounded and per-event. - **A trace is a causal tree.** One request produces a set of timed operations across however many processes handled it, joined by an identifier that is propagated along with the request. The tree records not only how long each step took but which step was waiting on which. ## What each shape can and cannot answer | Shape | Unit stored | Answers cheaply | Structurally cannot answer | |---|---|---|---| | Aggregate time series | one value per series per interval | rates, trends and distributions over long windows | which request, what it contained, why that one differed | | Discrete record | one event with its own fields | exactly what happened at one point, in full detail | population-wide questions without scanning everything; what other processes did | | Causal tree | one request's operations across processes | where a request's time went, which hop failed, what waited on what | fleet-wide rates, because typically only a fraction of requests are kept | The "cannot" column is the part candidates skip, and it is the part that matters. **Aggregation is destructive at emission, not at storage.** When a request is folded into a counter, its identity is discarded before anything is written down; no retention setting, no query and no later re-processing brings it back. Equally, a log record is written by one process and knows only what that process knew — two services logging the same user action produce two records with no inherent relationship. And a trace, because keeping every one is usually unaffordable, is a sample of the population rather than the population, so counting traces to get an error rate quietly counts only what survived the keep decision. ## Crossing a process boundary Metrics and logs are process-local artefacts: each process aggregates its own numbers and writes its own records. A trace is the only one of the three whose unit of data is defined by **the request** rather than by the process, which is why it is the shape that survives a fan-out across services. That property does not come free — it requires an identifier to be carried across every hop. Logs can be correlated after the fact only if someone deliberately puts a shared identifier into every record; that is a convention a team adopts, not something the log shape provides. ## The same minute, seen three ways A municipal parking-permit service renews permits online and holds itself to a 320 ms p99 budget. Between 09:12 and 09:31 one morning, 4,180 renewals go through and something is wrong. 1. **The aggregate** says the p99 rose to 604 ms and the failure ratio went from 0.2% to 3.1% at the north office. That is the whole population, it is cheap, and it is enough to know the budget is blown — and it will never say which renewal was slow. 2. **The records** for that window contain individual renewals: a plate, an office, a payment channel, a retry count, a 587 ms duration, an error string. Full detail for the events that were kept, and no view of the population without reading all of them. 3. **The tree** for one of those renewals shows 411 of its 587 ms sitting inside the card-charge hop, which itself waited on a lookup that returned late. That is the only shape that shows the waiting relationship between two processes. None of the three is a superset of the others. That is why all three exist. ## What this means when you instrument - **Anything you will want to slice by later must be a dimension chosen in advance.** An aggregate can only be broken down along dimensions that existed when it was written. - **Anything that is interesting per event belongs in a record**, not in a dimension — dimensions multiply series, fields on a record do not. - **Anything that must be followed across a boundary needs an identifier that travels with the request.** Without propagation there is no tree, only unrelated records in several places. - **A signal is only as good as the question it was shaped for.** Reaching for a shape that cannot answer your question is the most common reason a team owns three signals and still cannot explain an incident.

  • If you already write a structured record for every request, why emit metrics at all?
    Because the questions differ in cost and in completeness. Computing a rate by scanning records re-reads everything and gets slower as traffic grows, while an aggregate answers the same question in a fixed amount of work over any window. Aggregates also count the whole population, including requests whose per-event telemetry was thinned or dropped, which makes them the honest denominator.
  • Which of the three shapes stops working the moment a request crosses a process boundary, and why?
    Records do. Each process writes its own, and the shape carries no relationship to what another process wrote about the same request. Correlating them requires a shared identifier that someone chose to propagate and to include in every record. A trace has that identifier by construction — the propagated context is what makes the tree a tree rather than a pile of unrelated spans.
  • Could you drop metrics entirely if you traced every single request?
    Not practically. Keeping every trace is expensive in exactly the dimension that grows with traffic, and aggregate questions over spans cost far more to answer than reading a pre-aggregated series. You would also lose the property that makes aggregates trustworthy for alerting: they cover all traffic, whereas trace data covers only what the keep decision retained.

Metrics are the gauges on a depot wall, logs are each driver's diary entry, and a trace is one parcel's stamped itinerary from depot to doorstep.

saying these in an interview costs you the question

  • Says a metric can be drilled down to the individual slow request
  • Describes a trace as just logs that share a request identifier
  • Assumes every request's trace is stored by default
  • Thinks two services' logs correlate without a propagated identifier
  • Treats having all three signals switched on as observability