skip to content

Telemetry Signals

What the telemetry signals are and when to reach for each one. A classic screen for observability roles: candidates who can only recite 'three pillars' get probed on trade-offs and newer signals.

on this pageshow

questions

11

Which of metrics, logs and traces do you check first for a rising error rate, one customer's failed request, or one endpoint's doubled p99?

level: juniorimportance: must knowfreq 78%

answer

  1. Match the signal to the question's shape
  2. Population trend, one subject, or decomposition
  3. Aggregate to narrow, then one example
  4. Per-hop time attribution needs a trace

basics

~20 s

Metrics first for a rising error rate: they show whether it moved, when, and how far. Per-request records first for one named customer's request. A trace first for a single endpoint's latency, to see which downstream hop grew.

solid answer

~50 s

Pick the signal that matches the **shape** of the question. A moved error rate is a population-over-time question: start with aggregated counters, which tell you when it moved, how far, and which endpoint or status carries it. One named customer's failed request is a single-subject question: go to the per-request records and find it by its identifier, because an aggregate can only say roughly how many failed, never which. A p99 that doubled on one endpoint is a decomposition question: a trace of a slow request shows which downstream hop absorbed the extra time. The discipline is the second hop, not the first — aggregate to confirm and narrow, jump to one concrete request inside the bad window, then read that request's records for the reason. An investigation that never leaves the aggregate ends in a guess.

code

pseudocode · 9 lines
pseudocode
question: "error ratio moved at 09:14"
  -> aggregate: ratio by endpoint, status over 08:00..10:00
  -> narrow:    endpoint=submit_manifest, status=502
  -> example:   one request id in that window and slice
  -> follow:    that request's trace, then its records

question: "tour 4471's manifest would not submit"
  -> records:   filter by tour_id=4471, 09:00..09:30
  -> follow:    that request's trace if it crossed services

go deeper

for a junior

Be ready to name which signal you would open first for three everyday questions: a rate that moved, one customer's failed request, and one endpoint that got slower. Naming the signal and giving one reason for each is enough at this level.

for a middle

An interviewer expects the mechanics of the second hop: how you get from a bad window in an aggregate to one concrete request inside it, and what must already have been recorded for that jump to exist at all.

for a senior

Show the judgement of stopping and switching: when the aggregate has told you everything it can, when a trace is a stub rather than an answer, and how you avoid burning half an incident hunting inside the wrong signal.

for a principal

Own whether your estate can actually perform the hops you describe. If engineers routinely cannot get from a trend to one example request, that is a platform gap with a named emit-time fix, not an individual skill problem.

## Start from the question, not from the tool Every investigation begins with a sentence a human actually said: "the error rate is up", "this customer says their booking failed", "that endpoint got slower yesterday". The signal you reach for is decided by the *shape* of that sentence, not by which store you happen to be fluent in. Three shapes cover almost everything: - **A population-over-time question** — "did anything change, how much, and when?" Aggregated counters and timers answer this in one query, because that is exactly what they are: a summary over many events, cheap to keep at high resolution for a long time. - **A single-subject question** — "what happened to *this* request?" Only a per-request record can answer it. An aggregate can tell you roughly how many failed; it can never tell you *which* ones, because the identity of the individual event is precisely what aggregation discards. - **A decomposition question** — "the endpoint got slower; where did the extra time go?" A trace answers this because it records one request as a tree of timed hops across services, so the extra milliseconds have an address. ## The mapping, with its second hop | The question someone asked | Ask first | What it gives you | The second hop | |---|---|---|---| | "Error rate moved — is it real?" | Aggregated counters | When it moved, how far, which endpoint and status carry it | One example failing request inside the bad window | | "This customer's request failed" | Per-request records | The exact ordering, the exact message, the exact inputs | That request's trace, if the failure crossed a service boundary | | "p99 doubled on one endpoint" | A trace of a slow request | Which downstream hop absorbed the extra time | The records of that hop, for *why* it was slow | | "Is it happening to everyone?" | Aggregated counters | The blast radius and the affected slice | One example inside the affected slice | The first hop is rarely the interesting one. What separates a strong answer from a weak one is naming the **second hop** and what has to exist for it to be possible at all. ## A worked investigation An orchestra-touring logistics platform routes instrument freight, crew travel and venue slots across four services: `tour-routing`, `freight-booking`, `visa-docs` and `venue-slotting`. On one Tuesday morning three questions arrive within forty minutes. 1. The failure ratio on `freight-booking` moves from 0.4% to 3.1% at 09:14. That is a population question, and the aggregate has already answered the "is it real" part; the only remaining aggregate work is narrowing — which endpoint, which status, which region. 2. A tour manager writes in that the manifest for one specific tour would not submit at 09:22. That is a single-subject question. No aggregate will find it. You go to the records for that submission, by its identifier. 3. The p99 on `venue-slotting` rises from 240 ms to 1,880 ms while its error ratio stays flat. That is a decomposition question, and a trace of one slow request shows 1,610 ms of the increase sitting in a single call to `visa-docs`. Three questions, three first hops. Forcing all three into whichever signal the responder knows best is what turns a twenty-minute investigation into an afternoon. ## Why aggregate-first is usually right The default order is: 1. **Confirm and bound** with the aggregate: is it real, how big, since when, who is affected. 2. **Narrow** by whatever dimensions the aggregate carries, until you hold a small time window and a small slice. 3. **Pick one example** request inside that window and slice. 4. **Follow it** through its trace and its records until the reason is a fact rather than an inference. This order works because steps 1 and 2 are cheap and steps 3 and 4 are expensive. Reading individual records to answer a fleet-wide question is slow and, at any real volume, misleading: whatever you read first feels representative and usually is not. ## When the mapping inverts Three real cases turn the default order around, and a candidate who names one of them is showing operating experience rather than recall: - **The count you need was never made a metric.** If nobody instrumented "manifest rejected by the customs broker", then counting over records *is* your aggregate for this incident. The follow-up decision is whether that count deserves promoting, because you will ask for it again. - **The failure never got far enough to be timed.** A request rejected at the edge may leave a stub of a trace and one line in a record, so the record is both the first and the last hop. - **Coverage is uneven.** If one of the four services is not instrumented for tracing, per-hop attribution has to be reconstructed from timestamps in records — much weaker evidence, and it should be labelled as such when you report it. Knowing the mapping is the junior half of this topic. The senior half is knowing when your estate cannot perform the hop you just described, saying so out loud, and treating that as a gap to close rather than a reason to guess.

  • You have found the five-minute window where the error ratio moved. What has to already exist for you to get from there to one example request?
    Something that connects the summary to concrete records. At minimum, records retained at that granularity that you can filter by the same endpoint, status and window. Better is a pointer sampled onto the aggregate itself, so a point on the curve carries the identifier of one request that contributed to it. Without either, you are searching by timestamp and hoping the volume is low enough to make that tractable.
  • The error-rate aggregate never moved, but a customer insists their request failed. Where do you start, and why is that not a contradiction?
    Start from the per-request records, found by that request's identifier or by the customer identifier plus a tight window. One failure among tens of thousands of requests moves a fleet-wide ratio by less than its normal noise, so the aggregate is not lying — it is answering a different question. This is the standard case where the single-subject signal is the only one that can answer at all.
  • When is a query over records the right first hop even for a fleet-wide question?
    When the thing you need to count was never turned into a metric. If nobody instrumented a counter for 'manifest rejected by the customs broker', then counting over records is the only aggregate you have, and using it for this incident is correct. The real decision afterwards is whether that count is worth promoting to a metric, because a question asked once during an incident is usually asked again.

An aggregate is the building's fire alarm: it tells you something is burning somewhere. A trace tells you which floor, and the request's own records tell you what was on fire.

saying these in an interview costs you the question

  • Opens a record search for every question, including trends
  • Says traces are only for latency, never for errors
  • Assumes a pre-built view exists for any question asked
  • Cannot get from an aggregate window to one example request
  • Reads one record as evidence of a fleet-wide trend
  • Picks the signal by habit rather than by the question
open as a page

What do metrics, logs and traces each capture, and what does each one's data shape make it unable to answer?

level: juniorimportance: must knowfreq 84%

basics

~20 s

Metrics are numbers aggregated over an interval, so they show trends but never an individual request. Logs are discrete timestamped records of single events. Traces are causally linked trees showing one request's path across services.

open as a page

In a flame graph from a sampling profiler, what do a frame's width and the vertical stacking each mean?

level: middleimportance: must knowfreq 52%

basics

~20 s

Width is the share of collected stack samples containing that frame — its share of the sampled resource, not elapsed time and not a call count. Vertical stacking is caller-to-callee ancestry, merged across every sample.

open as a page

What does a sampling profiler record, and how does a profile differ from a trace or a metric?

level: juniorimportance: should knowfreq 45%

basics

~20 s

A sampling profiler interrupts the program at a fixed rate and records the call stack at that instant, merging identical stacks into counts. The result attributes a resource — CPU time or allocated bytes — to code paths rather than to requests.

open as a page

How does monitoring differ from observability, and what must a system already be emitting for the second to work?

level: middleimportance: should knowfreq 58%

basics

~20 s

Monitoring answers questions you predicted, from views and conditions built in advance for known failure modes. Observability is the property that lets you answer questions nobody anticipated, by slicing retained per-request detail along dimensions you did not choose beforehand.

open as a page

What is a wide structured event, and which questions can it answer that a pre-aggregated metric cannot?

level: middleimportance: should knowfreq 40%

basics

~20 s

A wide structured event is one record emitted per unit of work — a request, a job, a message — carrying every dimension you might later slice by. Pre-aggregation discards those rows, so questions nobody anticipated can no longer be asked.

open as a page

Why does a metrics bill scale with the number of distinct time series while logs and traces scale with request volume?

level: middleimportance: should knowfreq 56%

basics

~20 s

A metric emits one value per series per interval, so its cost tracks how many distinct series exist rather than how many requests arrive. Logs and traces are produced per request, so their cost tracks traffic directly.

open as a page

A team emits metrics, logs and traces at full fidelity. How would you set each signal's resolution rather than dropping one?

level: principalimportance: should knowfreq 46%

basics

~20 s

Full fidelity on all three costs more than it answers, but dropping a signal leaves a class of question unanswerable. Set each signal's resolution separately: dimension breadth, per-event detail, kept fraction, sized by what it is asked on a bad day.

open as a page

A per-request question answered out of aggregate metrics: how do you recognise the trap, and what should have been emitted?

level: seniorimportance: nice to knowfreq 26%

basics

~20 s

You are in the trap when your evidence is a shape rather than a record: inferring one request's fate from a bump in a curve. The fix is at emit time, with records that carry the identifiers you will later filter by.

open as a page

When is leaving a sampling profiler always on in production worth its cost, versus profiling on demand?

level: seniorimportance: nice to knowfreq 20%

basics

~20 s

Always-on profiling earns its cost when the incidents you must explain are already over — a restarted or rescheduled process cannot be profiled retroactively. On-demand capture suffices for reproducible workloads and costs nothing while idle.

open as a page

What does a metric aggregate lose that a per-request record keeps, and when is losing it the right trade?

level: seniorimportance: nice to knowfreq 33%

basics

~20 s

Aggregation discards identity, every attribute nobody made a dimension, and the ability to re-slice history: it keeps how many, never which ones. It stays the right trade when the question is a rate or trend answered at fixed cost.

open as a page