skip to content

Walk through how a distributed tracing system such as Jaeger or Zipkin, instrumented via OpenTelemetry, tracks a single request as it flows through five microservices: what gets generated, what gets passed between services on the wire, and how does the backend reassemble it into one picture?

level: middleimportance: must knowfreq 80%

answer

  1. span = timed unit of work
  2. trace ID shared, span ID + parent ID per hop
  3. W3C traceparent header on the wire
  4. collector -> backend reassembles by trace ID
  5. Dapper -> Zipkin/Jaeger -> OpenTelemetry lineage

basics

~20 s

Each unit of work, like one service handling part of a request, creates a 'span' recording what it did and how long it took. All spans for one request share a trace ID, and each span points to its parent span, so a backend like Jaeger can stitch them into a single timeline showing the whole request's path and where time was spent.

solid answer

~50 s

OpenTelemetry instruments each service so that when it starts handling a unit of work, it opens a span: a record with a name, start/end time, tags, and a unique span ID. The first service in the chain generates a trace ID for the whole request. Every downstream call carries trace context, in practice the W3C traceparent header, encoding the trace ID, the calling span's ID (which becomes the new span's parent ID), and sampling flags. Each service extracts that header, creates its own child span under the received parent, does its work, and injects a new traceparent, with its own span as parent, into any further downstream calls it makes. Spans are exported asynchronously to a collector and stored keyed by trace ID; the backend, Jaeger or Zipkin, reassembles them into a tree or waterfall by following parent-child links, showing exactly which hop took how long and where errors occurred.

go deeper

for a junior

Should grasp that tracing shows the path and timing of a request across services, even without knowing the header format.

for a middle

Should be able to describe spans, trace IDs, and parent-child relationships, and know that context is passed via a header on each call.

for a senior

Should know the W3C Trace Context format at a conceptual level, be able to diagnose a broken trace from a missing propagation hop, and weigh instrumentation cost against value.

for a principal

Should be driving org-wide adoption of a single instrumentation standard, OpenTelemetry, and collector architecture so traces are comparable and joinable across every team's services.

## The unit of data Distributed tracing works by treating a single logical request as a tree of timed operations called **spans**, and by propagating enough context on every network call for each service to correctly place its own spans into that tree. The unit of data is the span: a record that has - a name (e.g. 'GET /cart' or 'db.query'), - a start timestamp, - a duration, - a status (ok or error), - a set of key-value attributes describing what happened, - and two identifiers, its own span ID and the trace ID of the request it belongs to. A **trace** is simply the full set of spans that share one trace ID, and the parent-child pointers between those spans, each span except the very first records a parent span ID, describe the causal structure: which operation triggered which other operation, and in what order. ## Hop by hop Concretely, the mechanism unfolds hop by hop. 1. The first service to touch the request, say an API gateway, has no incoming trace context, so its tracing library, OpenTelemetry's SDK in most modern stacks, mints a fresh trace ID and opens a **root span**. 2. If it needs to call an inventory service to serve the request, it injects the current trace context into the outgoing HTTP call, in practice as a W3C `traceparent` header of the shape `00-<trace-id>-<span-id>-<flags>`, where the span-id in that header is the gateway's current span, and the flags include a sampled bit. 3. The inventory service extracts that header on receipt, sees it belongs to an existing trace, and opens its own span as a child of the span ID it just read, with the same trace ID. 4. If the inventory service in turn calls a database or a pricing service, it repeats the same injection step with its own span now as the parent. 5. Each span, once finished, is exported, usually asynchronously and batched, over gRPC or HTTP, to a collector process, which forwards it to a storage backend, indexed by trace ID. A UI like the Jaeger or Zipkin web console then queries all spans for one trace ID and renders them as a waterfall or flame graph, ordered by start time and nested by parent-child relationship, which is what lets an engineer see at a glance that the pricing service's database call consumed 380ms out of a 420ms total request. ## Why grouping by log line is not enough The reason this exists, rather than just using logs and correlation IDs, is that logs alone give you grouping but not causal shape or precise timing. A trace answers exactly where the two seconds went, and in what order the five services called each other, in a way that scanning timestamped log lines across five services cannot reliably reconstruct, especially once retries, parallel fan-out calls, and asynchronous work are involved. Tracing also gives you span attributes as structured, queryable metadata (HTTP status code, DB statement, customer tier) that lets you slice traces, for example show me all traces where the payments span had status equal to error, far more cheaply than grepping unstructured logs. ## What it costs The trade-offs are real. - Every service needs instrumentation, either manual, wrapping code in spans, or automatic, agents that instrument common libraries like HTTP clients and JDBC drivers. - Every network hop, including message queues and especially any client the team doesn't fully control, must correctly propagate the trace-context header or the trace breaks into disconnected fragments. - Storing a full span tree for every single request at high request-per-second volumes is expensive in both storage and the write throughput the collector and backend must sustain, which is why almost no production system traces one hundred percent of traffic, this is the sampling problem, a distinct but closely related concern. - There is also a real engineering cost to running the collector pipeline itself and to keeping instrumentation libraries current as dependencies evolve. ## Failure modes Failure modes show up in a few recognizable shapes. - **A broken trace** happens when some hop, commonly a message queue consumer, a serverless function cold-starting, or a third-party API, fails to propagate the `traceparent` header, so the trace silently splits into two unconnected trace IDs and the UI shows an incomplete picture with a suspicious gap. - **A clock skew issue** happens when spans from different hosts have drifted system clocks, making the waterfall view show a child span apparently starting before its parent. - **Over-instrumentation**, spans for every trivial function call, produces noisy traces that are hard to read and expensive to store, while **under-instrumentation** leaves exactly the slow, opaque hop uninstrumented, which tends to be the one you needed visibility into during an incident. ## The lineage A concrete, well-known lineage: - Google's internal **Dapper** paper described exactly this span, trace-ID, parent-ID model at Google scale; - Twitter's **Zipkin** and Uber's **Jaeger** are open-source systems built on the same ideas; - and **OpenTelemetry** is the current CNCF-backed standard that unifies instrumentation APIs and the wire format, W3C Trace Context, so that a span emitted by a Java service and a span emitted by a Go service can be correctly stitched into one trace by any compliant backend.

  • What's the difference between a span and a trace?
    A span is a single timed unit of work with a name, duration, and attributes; a trace is the complete tree of all spans sharing one trace ID that together represent one end-to-end request. A trace typically contains many spans, one per meaningful operation across every service the request touched.
  • How does OpenTelemetry differ from Jaeger and Zipkin?
    OpenTelemetry is a vendor-neutral instrumentation standard, APIs, SDKs, and the wire protocol for propagating context and exporting spans; Jaeger and Zipkin are backends that store and visualize traces. Modern setups instrument with OpenTelemetry and export to whichever backend a team chooses, Jaeger, Zipkin, or a commercial APM, which is exactly the decoupling OpenTelemetry was built to provide.
  • What happens to a trace if one service in the chain doesn't propagate the traceparent header at all?
    The trace breaks at that hop: the unpropagated call starts a brand-new trace ID with no parent, so the backend shows two disconnected traces instead of one continuous one, and the causal link between them is lost unless some other correlating signal, like a shared correlation ID in logs, survives.

Think of a relay race where each runner (service) records their own start and finish time and which runner handed them the baton. Nobody needs to watch the whole race live; afterward, you can collect every runner's card, and because each card says which leg they ran and who handed them the baton, you can reconstruct the entire race's timeline and see exactly which leg was slow.

saying these in an interview costs you the question

  • Can't distinguish a span from a trace
  • Thinks the tracing backend infers ordering purely from timestamps rather than explicit parent-child span IDs
  • Unaware that trace context must be explicitly propagated on every network hop, including queues
  • Believes tracing replaces the need for logs entirely
  • Doesn't know any real system in this space (Jaeger, Zipkin, OpenTelemetry) even at a high level

context