Observability Foundations
The tool-agnostic ideas every monitoring stack is built on: telemetry signals, metric math, log pipelines, and tracing with sampling. Interviewers use these to check whether you understand telemetry itself or only one vendor's UI.
on this pageshowhide
explore
- Telemetry Signals11 questions
- Metrics, Logs, and Traces Trade-offs4 questions
- Events and Continuous Profiling4 questions
- Choosing the Right Signal3 questions
- Metrics Concepts19 questions
- Counter, Gauge, Histogram, Summary4 questions
- Labels and Cardinality4 questions
- Histogram Buckets and Percentile Pitfalls4 questions
- Push vs Pull Collection4 questions
- RED and USE Checklists3 questions
- Logging Concepts19 questions
- Structured Logging4 questions
- Log Levels in Production3 questions
- Aggregation Pipeline Stages4 questions
- Index-Based vs Label-Based Storage4 questions
- Retention and Cost Control4 questions
- Tracing and Sampling16 questions
- Spans, Traces, and Context Propagation4 questions
- Correlation IDs and Exemplars4 questions
- Head vs Tail Sampling4 questions
- Agent vs Library Instrumentation4 questions
- Backend Developerroleanchors this topic
- Data Engineerroleanchors this topic
- DevOps / SRE Engineerroleanchors this topic
- DevSecOps Engineerroleanchors this topic
- Full Stack Developerroleanchors this topic
- Java Backend Developerroleanchors this topic
- Java SDETroleanchors this topic
- Kotlin Backend Developerroleanchors this topic
- Kubernetesskillanchors this topic
- MLOps Engineerroleanchors this topic
- Network Engineerroleanchors this topic
- QA Engineerroleanchors this topic
- Software Architectroleanchors this topic
- Forward Deployed Engineerrole
questions
65 · 4 sectionsWhich of metrics, logs and traces do you check first for a rising error rate, one customer's failed request, or one endpoint's doubled p99?
basics
~20 sMetrics first for a rising error rate: they show whether it moved, when, and how far. Per-request records first for one named customer's request. A trace first for a single endpoint's latency, to see which downstream hop grew.
What do metrics, logs and traces each capture, and what does each one's data shape make it unable to answer?
basics
~20 sMetrics are numbers aggregated over an interval, so they show trends but never an individual request. Logs are discrete timestamped records of single events. Traces are causally linked trees showing one request's path across services.
In a flame graph from a sampling profiler, what do a frame's width and the vertical stacking each mean?
basics
~20 sWidth is the share of collected stack samples containing that frame — its share of the sampled resource, not elapsed time and not a call count. Vertical stacking is caller-to-callee ancestry, merged across every sample.
What does a sampling profiler record, and how does a profile differ from a trace or a metric?
basics
~20 sA sampling profiler interrupts the program at a fixed rate and records the call stack at that instant, merging identical stacks into counts. The result attributes a resource — CPU time or allocated bytes — to code paths rather than to requests.
How does monitoring differ from observability, and what must a system already be emitting for the second to work?
basics
~20 sMonitoring answers questions you predicted, from views and conditions built in advance for known failure modes. Observability is the property that lets you answer questions nobody anticipated, by slicing retained per-request detail along dimensions you did not choose beforehand.
What identifies one time series in a dimensional metrics system, and how does adding a metric label change the total count?
basics
~20 sA time series is the metric name plus its complete label set; change any label value and it is a different series. Adding a label multiplies the series count by that label's distinct values rather than adding to it.
What makes a metric a counter rather than a gauge, and why is a counter read as a rate over a window?
basics
~20 sA counter only ever increases and returns to zero when its publishing process restarts; a gauge is a level that moves either way. Because a counter's absolute value depends on process uptime, you read its rate over a window.
Why can't you average the p99 latencies of ten replicas to get the fleet's p99?
basics
~20 sA percentile is an order statistic over a set of observations, not a quantity that adds. Averaging per-replica p99s ignores how much traffic each served and describes no real request. Merge the underlying bucket counts first, then estimate the quantile once.
What does a pull-based metrics collector get for free, and what does it demand of every target?
basics
~20 sA pull-based collector already holds the list of what should exist, so every failed poll is itself a liveness signal, and one setting fixes the whole fleet's sampling resolution. In exchange each target must stay routable and answerable at any instant.
Under the RED checklist, how do you instrument a request-serving service, and what counts as an error?
basics
~20 sRED means emitting three things for a request-serving component: request rate, failed requests, and a duration distribution. Rate and errors come from one counter carrying an outcome dimension, and "error" must be defined in writing before the count means anything.
What does each conventional log level from TRACE to FATAL mean, and what test separates WARN from ERROR?
basics
~20 sLog levels are a contract about who must act. ERROR means something failed and a human should look; WARN means something recovered but degraded; INFO records state changes; DEBUG and TRACE are developer detail, off by default in production.
What does a log backend that full-text indexes every line at ingest buy you, and what does it cost?
basics
~20 sA full-text index built at ingest makes any word in any line searchable without declaring it first, and makes aggregations fast. You pay twice: CPU on the write path for every record, and index bytes stored alongside the logs.
What does structured logging give a log consumer that a formatted message string cannot?
basics
~20 sStructured logging emits each line as named, typed fields instead of one formatted sentence. Ingest no longer has to guess where each value starts and ends, and queries can filter, compare and aggregate on a field rather than searching raw text.
In a log aggregation pipeline, what does each stage — collect, parse, enrich, route, store — do, and which decisions can no later stage undo?
basics
~20 sCollect reads bytes out of the producing process, parse turns text into named fields, enrich attaches context the emitter lacked, route chooses destinations, store indexes and serves. Collect losses are unrecoverable: an unread line exists in no other copy.
What drives the cost of a centralized log platform, and which of those does shortening retention actually reduce?
basics
~20 sFour things drive a log bill: bytes ingested, the parsing and index-building done on them at write time, stored bytes multiplied by replicas and retained days, and query load. Shortening retention shrinks only the third.
What does a trace ID on every log line buy you, and what makes it go missing?
basics
~20 sA trace ID turns log search into an exact lookup: one filter returns every line every service wrote for one request. It goes missing when no span is active, when context is lost on an async hop, or when the logger never reads it.
In distributed tracing, what identifiers does a span carry, and why is a span an interval rather than a point in time?
basics
~20 sA span carries a trace id shared by every span of one request, its own span id, and its parent's span id. It stores a start time and a duration because it measures an operation that lasts, not an instant.
In head-based trace sampling, where is the keep-or-drop decision made, and why is it nearly free?
basics
~20 sHead-based sampling decides at the start of a trace, before any span is recorded. The first service rolls the dice once and stamps the verdict onto the outgoing request, so a dropped trace costs nothing downstream.
What can an automatic instrumentation agent see inside a running service, and what is it structurally blind to?
basics
~20 sAn automatic instrumentation agent hooks call sites it has a plugin for - web handlers, database drivers, message clients - so it produces spans at library boundaries only. Domain logic between those boundaries stays one unattributed gap in the trace.
In distributed tracing, what must travel with an outbound request for the callee to continue the same trace, and what breaks when one hop drops it?
basics
~20 sThe request must carry the trace id, the id of the span making the call, and whether the trace is being recorded. A hop that drops them makes the callee start a brand-new trace, taking every service below it with it.