When a single user-facing request fans out across five microservices, what mechanism lets a tool like Jaeger reconstruct the full call tree and show which of those five services caused the added latency, and what has to happen at each service boundary for that to work?
answer
- trace ID + span ID + parent span ID
- context propagation across headers
- W3C traceparent header
- span = one unit of work
- collector -> backend (Jaeger)
basics
~20 sEach call in the chain gets tagged with a shared trace ID, and each step gets its own ID linked back to whoever called it. Tools like Jaeger stitch these into one timeline showing where time actually went.
solid answer
~40 sDistributed tracing propagates a shared trace context - a trace ID for the whole request, plus a span ID per unit of work - across every service boundary, typically via HTTP headers (W3C traceparent) or messaging metadata. Each service, instrumented with an SDK like OpenTelemetry, creates a span recording its parent span ID, start/end time, and attributes (route, status, query), then exports it to a collector that forwards to a backend like Jaeger. Jaeger stitches spans sharing a trace ID into a waterfall tree, showing which hop dominated latency, whether calls ran sequentially or in parallel, and where errors originated. The critical requirement: every hop, including async ones like queues, must propagate the incoming context onward - one un-instrumented hop splits the trace into two disconnected pieces.
go deeper
Should understand that a trace ID follows a request across services so you can see the whole journey in one view, even without header-format detail.
Should describe the trace ID/span ID/parent span relationship, name OpenTelemetry and a backend like Jaeger, and know context must be propagated via headers.
Should explain what breaks when propagation is missing at one hop, discuss sampling strategies and their trade-offs, and read a waterfall view to diagnose latency.
Should discuss tail-based sampling architecture, collector scaling and cost at fleet scale, correlating traces with metrics/logs, and standardizing instrumentation across polyglot services via OpenTelemetry's vendor-neutral SDKs.
## The unit of work: the span **Distributed tracing** reconstructs the end-to-end path of a single logical request as it crosses multiple services by attaching a shared identifier to it at the very start and propagating that identifier through every subsequent call it triggers. The core unit is the **span**: a single timed unit of work, recording - what operation it represents, - when it started and ended, - and - critically - which other span, if any, called it (its **parent span ID**). Every span also carries a **trace ID**, which is the same for every span belonging to the same originating request, no matter how many services it passes through. ## How the context travels, and what the backend assembles When a request enters the system - say, at an API gateway - a new trace ID and a root span are created. As the gateway calls downstream service A, it attaches the trace ID and its own span ID (as the new "parent" reference) to the outgoing call, most commonly as an HTTP header following the **W3C Trace Context** standard (the `traceparent` header, plus an optional `tracestate` header for vendor-specific extensions). Service A, instrumented with a tracing SDK - **OpenTelemetry** is the current vendor-neutral standard - reads that incoming context, creates its own child span referencing it as parent, does its work, and if it in turn calls service B, propagates the same trace ID plus its own span ID onward. Each service, on completing its span, exports it asynchronously to a **collector**, which forwards the data to a storage/query backend - **Jaeger** is one of the most widely used open-source ones. Jaeger's job is purely to assemble every span sharing a trace ID into a tree, ordered by parent-child relationships and timestamps, and render it as a **waterfall**: a horizontal bar per span, positioned and sized by its start time and duration, nested under its parent, making it visually obvious which hop took the most time, whether calls happened sequentially or in parallel, and exactly where in the tree an error first appeared. ## Why logs and metrics can't answer this The reason this exists is that logs and metrics alone can't answer "why was this one request slow" in a system with many services. - **Metrics** are aggregated and lose per-request identity. - **Logs** from different services live in different streams with no inherent connection between them unless you manually correlate timestamps, which is unreliable under concurrent load. Tracing solves this specific problem: giving you the actual causal shape of one request's journey, with real timing, across process and network boundaries that would otherwise be opaque to each other. ## The hop that breaks the trace The critical mechanical requirement, and the thing that most commonly breaks in practice, is that every hop in the chain must both read the incoming trace context and propagate it to whatever it calls next - including calls that aren't a direct synchronous HTTP request, like publishing a message to a queue or kicking off a background job. If even one service in the middle forgets to forward the context (for example, using an HTTP client library that isn't instrumented, or manually constructing a request without copying the headers), the trace doesn't just have a "gap" - it splits into two entirely disconnected traces, each with spans but no relationship between them, and the missing service becomes an invisible black box in the visualization even if it's the actual source of the latency. ## Sampling: head against tail A key operational trade-off is **sampling**. Capturing every span of every request at high request volume is expensive - both in the network/CPU cost of exporting so much data and in storage/query cost at the backend - so most systems sample. | Strategy | When the keep decision happens | What that buys and costs | |---|---|---| | **Head-based sampling** | decides whether to record a trace at the very start (e.g., "record 1% of traces, chosen randomly") | cheap but risks missing the rare interesting traces purely by chance | | **Tail-based sampling** | defers the decision until the whole trace has been observed, buffering all spans until the end and then choosing to keep it based on outcome - always keeping error traces or slow traces regardless of the random sampling rate | captures the traces you actually care about but requires buffering full traces before deciding, adding memory and pipeline complexity | ## A worked example A worked example: a checkout request fans out from a gateway to an orders service, which calls inventory and pricing concurrently. In the Jaeger waterfall, the orders span might show a total duration of 220ms, with two child spans - inventory and pricing - each taking around 200ms but overlapping almost entirely, because they were issued in parallel rather than sequentially; without tracing, log timestamps alone would make it easy to misread this as 400ms of sequential work, when the actual critical path is the slower of the two parallel calls plus a small coordination overhead.
- What happens to a distributed trace if one service in the middle of the call chain doesn't propagate the incoming trace context to its own outbound calls?The trace breaks into two disconnected pieces - everything downstream of that gap starts a fresh trace ID or becomes orphaned spans, so the visualization can't show the full end-to-end waterfall. The missing service becomes an invisible black box even if it's the actual bottleneck.
- How does sampling affect what you can see in a tracing backend, and why is 100% sampling usually impractical at scale?Sampling decides which traces are actually recorded and exported; at high request volume, capturing every trace in full detail is expensive in storage and collector throughput. Systems use head-based sampling (decide at trace start, e.g., 1%) or tail-based sampling (decide after seeing the whole trace, keeping ones with errors or high latency) to keep the interesting traces without capturing everything.
- What's the difference between a span's tags/attributes and its logs/events in OpenTelemetry terms?Attributes are key-value metadata describing the span as a whole (e.g., http.method, http.status_code) set once and filterable across many spans. Span events are timestamped points within the span's duration (e.g., 'cache miss at 45ms') capturing something that happened during execution, closer to a log line scoped to that span.
- Why might a trace show two child spans running in parallel taking 200ms each, but the parent span duration is only 220ms instead of the naive 400ms expectation?Because the calls were issued concurrently rather than sequentially - the parent's duration reflects wall-clock time from when it issued the first call to when the last result returned, not the sum of children's durations. This is exactly the kind of thing a waterfall view makes visible that log timestamps alone would not.
Like a relay race where every runner writes down the shared race number plus their own leg's start and finish time on a shared clipboard - afterward you can reconstruct exactly which leg was slow, even though no single runner saw the whole race.
saying these in an interview costs you the question
- thinks logging timestamps alone gives you distributed tracing
- doesn't know trace context must be propagated across every hop
- confuses a span with a full trace
- assumes tracing works automatically with zero code/library changes
- no concept of sampling and its cost trade-off