skip to content

Observability Foundations

The tool-agnostic ideas every monitoring stack is built on: telemetry signals, metric math, log pipelines, and tracing with sampling. Interviewers use these to check whether you understand telemetry itself or only one vendor's UI.

on this pageshow

explore

questions

65 · 4 sections

Which of metrics, logs and traces do you check first for a rising error rate, one customer's failed request, or one endpoint's doubled p99?

level: juniorimportance: must knowfreq 78%
basics
~20 s

Metrics first for a rising error rate: they show whether it moved, when, and how far. Per-request records first for one named customer's request. A trace first for a single endpoint's latency, to see which downstream hop grew.

open as a page

What do metrics, logs and traces each capture, and what does each one's data shape make it unable to answer?

level: juniorimportance: must knowfreq 84%
basics
~20 s

Metrics are numbers aggregated over an interval, so they show trends but never an individual request. Logs are discrete timestamped records of single events. Traces are causally linked trees showing one request's path across services.

open as a page

In a flame graph from a sampling profiler, what do a frame's width and the vertical stacking each mean?

level: middleimportance: must knowfreq 52%
basics
~20 s

Width is the share of collected stack samples containing that frame — its share of the sampled resource, not elapsed time and not a call count. Vertical stacking is caller-to-callee ancestry, merged across every sample.

open as a page

What does a sampling profiler record, and how does a profile differ from a trace or a metric?

level: juniorimportance: should knowfreq 45%
basics
~20 s

A sampling profiler interrupts the program at a fixed rate and records the call stack at that instant, merging identical stacks into counts. The result attributes a resource — CPU time or allocated bytes — to code paths rather than to requests.

open as a page

How does monitoring differ from observability, and what must a system already be emitting for the second to work?

level: middleimportance: should knowfreq 58%
basics
~20 s

Monitoring answers questions you predicted, from views and conditions built in advance for known failure modes. Observability is the property that lets you answer questions nobody anticipated, by slicing retained per-request detail along dimensions you did not choose beforehand.

open as a page

What identifies one time series in a dimensional metrics system, and how does adding a metric label change the total count?

level: juniorimportance: must knowfreq 72%
basics
~20 s

A time series is the metric name plus its complete label set; change any label value and it is a different series. Adding a label multiplies the series count by that label's distinct values rather than adding to it.

open as a page

What makes a metric a counter rather than a gauge, and why is a counter read as a rate over a window?

level: juniorimportance: must knowfreq 78%
basics
~20 s

A counter only ever increases and returns to zero when its publishing process restarts; a gauge is a level that moves either way. Because a counter's absolute value depends on process uptime, you read its rate over a window.

open as a page

Why can't you average the p99 latencies of ten replicas to get the fleet's p99?

level: middleimportance: must knowfreq 68%
basics
~20 s

A percentile is an order statistic over a set of observations, not a quantity that adds. Averaging per-replica p99s ignores how much traffic each served and describes no real request. Merge the underlying bucket counts first, then estimate the quantile once.

open as a page

What does a pull-based metrics collector get for free, and what does it demand of every target?

level: middleimportance: must knowfreq 74%
basics
~20 s

A pull-based collector already holds the list of what should exist, so every failed poll is itself a liveness signal, and one setting fixes the whole fleet's sampling resolution. In exchange each target must stay routable and answerable at any instant.

open as a page

Under the RED checklist, how do you instrument a request-serving service, and what counts as an error?

level: middleimportance: must knowfreq 64%
basics
~20 s

RED means emitting three things for a request-serving component: request rate, failed requests, and a duration distribution. Rate and errors come from one counter carrying an outcome dimension, and "error" must be defined in writing before the count means anything.

open as a page

What does each conventional log level from TRACE to FATAL mean, and what test separates WARN from ERROR?

level: juniorimportance: must knowfreq 66%
basics
~20 s

Log levels are a contract about who must act. ERROR means something failed and a human should look; WARN means something recovered but degraded; INFO records state changes; DEBUG and TRACE are developer detail, off by default in production.

open as a page

What does a log backend that full-text indexes every line at ingest buy you, and what does it cost?

level: juniorimportance: must knowfreq 62%
basics
~20 s

A full-text index built at ingest makes any word in any line searchable without declaring it first, and makes aggregations fast. You pay twice: CPU on the write path for every record, and index bytes stored alongside the logs.

open as a page

What does structured logging give a log consumer that a formatted message string cannot?

level: juniorimportance: must knowfreq 76%
basics
~20 s

Structured logging emits each line as named, typed fields instead of one formatted sentence. Ingest no longer has to guess where each value starts and ends, and queries can filter, compare and aggregate on a field rather than searching raw text.

open as a page

In a log aggregation pipeline, what does each stage — collect, parse, enrich, route, store — do, and which decisions can no later stage undo?

level: middleimportance: must knowfreq 70%
basics
~20 s

Collect reads bytes out of the producing process, parse turns text into named fields, enrich attaches context the emitter lacked, route chooses destinations, store indexes and serves. Collect losses are unrecoverable: an unread line exists in no other copy.

open as a page

What drives the cost of a centralized log platform, and which of those does shortening retention actually reduce?

level: middleimportance: must knowfreq 64%
basics
~20 s

Four things drive a log bill: bytes ingested, the parsing and index-building done on them at write time, stored bytes multiplied by replicas and retained days, and query load. Shortening retention shrinks only the third.

open as a page

What does a trace ID on every log line buy you, and what makes it go missing?

level: juniorimportance: must knowfreq 66%
basics
~20 s

A trace ID turns log search into an exact lookup: one filter returns every line every service wrote for one request. It goes missing when no span is active, when context is lost on an async hop, or when the logger never reads it.

open as a page

In distributed tracing, what identifiers does a span carry, and why is a span an interval rather than a point in time?

level: juniorimportance: must knowfreq 76%
basics
~20 s

A span carries a trace id shared by every span of one request, its own span id, and its parent's span id. It stores a start time and a duration because it measures an operation that lasts, not an instant.

open as a page

In head-based trace sampling, where is the keep-or-drop decision made, and why is it nearly free?

level: middleimportance: must knowfreq 62%
basics
~20 s

Head-based sampling decides at the start of a trace, before any span is recorded. The first service rolls the dice once and stamps the verdict onto the outgoing request, so a dropped trace costs nothing downstream.

open as a page

What can an automatic instrumentation agent see inside a running service, and what is it structurally blind to?

level: middleimportance: must knowfreq 62%
basics
~20 s

An automatic instrumentation agent hooks call sites it has a plugin for - web handlers, database drivers, message clients - so it produces spans at library boundaries only. Domain logic between those boundaries stays one unattributed gap in the trace.

open as a page

In distributed tracing, what must travel with an outbound request for the callee to continue the same trace, and what breaks when one hop drops it?

level: middleimportance: must knowfreq 70%
basics
~20 s

The request must carry the trace id, the id of the span making the call, and whether the trace is being recorded. A hop that drops them makes the callee start a brand-new trace, taking every service below it with it.

open as a page