skip to content

Spans, Traces, and Context Propagation

The anatomy of a trace: how spans nest into trees and how trace context crosses service boundaries in headers. The foundational tracing interview question — everything else in tracing builds on it.

on this pageshow

questions

4

In distributed tracing, what identifiers does a span carry, and why is a span an interval rather than a point in time?

level: juniorimportance: must knowfreq 76%

answer

  1. One value joins the whole request
  2. One field gives the trace its shape
  3. Three ids: shared, own, caller's
  4. Start plus duration, not one timestamp
  5. No parent reference means root

basics

~20 s

A span carries a trace id shared by every span of one request, its own span id, and its parent's span id. It stores a start time and a duration because it measures an operation that lasts, not an instant.

solid answer

~50 s

A span is the record of one unit of work — a request handler, an outbound call, a query. It carries three identifiers. The **trace id** is generated once at the entry point of a request and copied into every span produced anywhere in that request; it is the join key. The **span id** is unique to this span. The **parent span id** points at the span that caused the work, and it is the only thing that gives a trace its shape — the span with no parent is the root. Alongside those, a span carries a name, a start timestamp, a duration, key-value attributes, timestamped events and a status. It is an interval rather than an event because the measurement *is* the elapsed time: nesting, overlap and unexplained gaps are only drawable from a start plus a duration.

code

pseudocode · 10 lines
pseudocode
span {
  traceId      = "4c9a1f..."   // same on every span of this request
  spanId       = "b71e08..."   // unique to this span
  parentSpanId = "a20d5c..."   // absent on the root span
  name         = "GET /invoices/{id}"
  startTime    = 2026-09-06T09:14:02.118Z
  durationMs   = 43
  attributes   = { "http.status": 200, "tenant": "north-grid" }
  status       = OK
}

go deeper

for a junior

Be ready to name the three identifiers and say which one is shared across the request. Knowing that the root span is the one without a parent, and that a span records a duration rather than a moment, is enough at this level.

for a middle

Explain the mechanics: the trace id is copied unchanged at every hop, the parent link is stored on the child, and identifiers are random because they are minted without coordination. Be able to say what each field would cost you if it were absent.

for a senior

Show that you read waterfalls for what the interval model gives you — the parent's self-time, overlapping children that reveal accidental serialisation, and unexplained gaps before the first child. Mention monotonic timing for durations versus wall-clock start times.

for a principal

Own the conventions: what makes a good span name, which attributes are mandatory across the estate, and where the boundary sits between a value that belongs on a span and one that must never reach a metric label.

## What a span actually is A **span** is the record of one unit of work inside one process: an inbound request handler, an outbound HTTP or database call, a scheduled job, or any block of code someone judged worth timing. A **trace** is not a separate object that anybody writes; it is simply the set of every span that shares one identifier, drawn as a tree. Every capability a tracing system offers — the waterfall view, the service map, the latency breakdown — is derived from three identifier fields plus two clock readings carried on each individual span. ## The three identifiers | Field | Scope | Where the value comes from | What is lost without it | |---|---|---|---| | trace id | the whole request, in every process it touches | generated once, at the entry point, then copied unchanged into every span | nothing can be grouped: you own a pile of loose spans, not traces | | span id | this one span only | generated by the process that creates the span | children have nothing to point at, so the tree collapses into a list | | parent span id | the edge back to the span that caused this work | copied from the span currently executing, or from context received over the wire | the spans are in the right trace but have no shape or ordering | Three properties of those identifiers matter in an interview. 1. **The trace id is copied, never regenerated.** A service that mints its own trace id for incoming work has not joined the trace; it has started a second one. Copying is the entire mechanism by which spans emitted from 41 different processes end up on one screen. 2. **The identifiers are opaque random values, not counters.** Spans are created concurrently in many processes, in several languages, with no coordination and no central allocator, so uniqueness has to come from the width of the random value rather than from a sequence. 3. **The parent link lives on the child.** A parent span never learns what it caused; it ends and is exported without any list of children. That asymmetry is exactly what lets each service report on its own schedule and never talk to its peers about telemetry. The **root span** is simply the span that has no parent reference, and a healthy trace has precisely one. ## Why a span is an interval, not a point A log line answers *what happened at time T*. A span answers *what happened between T and T+d*, and that difference is the reason tracing exists as a separate signal. Storing a start timestamp plus a duration buys four things a point-in-time record cannot give: - **The duration is the measurement.** It is the number the record exists to carry; everything else on the span is context for it. - **Nesting becomes visible.** A child's bar is drawn inside its parent's bar. The parent's duration minus the time covered by its children is the work the parent did itself — plus any time it spent waiting in a queue that nobody instrumented. - **Concurrency becomes visible.** Two children whose bars overlap ran in parallel; two whose bars sit end to end ran in sequence. A fan-out that was meant to be concurrent but is accidentally serial is diagnosed by eye, with no arithmetic at all. - **Gaps become visible.** A 43 ms hole between a parent starting and its first child starting is real time spent somewhere the instrumentation does not reach — deserialisation, a lock, a pool checkout, a garbage-collection pause. A careful implementation measures the elapsed time from a **monotonic** clock, which only ever moves forward, so that a clock correction in the middle of the operation cannot produce a negative duration. The start timestamp itself has to be wall-clock time, because it is the only thing that positions this span against spans produced on other machines. ## What else the record carries - A **name**: a short, low-cardinality label for the operation. Embedding an identifier in the name — a customer number, a meter serial — gives every request a unique operation and destroys every aggregate view built over names. - **Attributes**: key-value pairs describing this particular execution, such as the route, the response code, or the tenant. High-cardinality values are far more affordable here than on a metric, because a span is one finite record rather than a permanent time series. - **Events**: timestamped points inside the interval, for the things that genuinely happen at an instant rather than over one. - **A status**: whether the operation succeeded or failed, which is what error views and error-biased retention read. ## What an interviewer is listening for The candidate should say that the trace id is the join key and is copied rather than reissued; that the child stores the parent link, so the tree is reconstructed from the leaves upward; that the root is defined by the *absence* of a parent; and that a span is an interval because duration, overlap and gaps are the whole product. A candidate who describes a span as "a log line with a trace id on it" has missed the point of the signal.

  • What identifies the root span of a trace, and what does it mean if a trace has two of them?
    The root is the span with no parent reference. Two roots in one trace means the tree was severed: usually a hop that failed to forward context, so a service started what it believed was a fresh trace. Renderers cope badly — they show disconnected fragments, or promote one root and hang the rest beside it.
  • Why are trace and span identifiers random values rather than sequential counters?
    They are minted concurrently by many processes, in different languages, with no coordination and no central allocator. A counter would need a shared allocator on the hot path of every request. Randomness of sufficient width gives practical uniqueness for free, and it also gives sampling schemes a value they can hash consistently on any host.
  • If duration is the measurement, why record a start timestamp at all?
    The start timestamp is what positions the span against other spans, including spans from other processes. It is what makes nesting, parallelism and gaps visible: the distance between a parent starting and its first child starting is time spent somewhere nobody instrumented, and you cannot see that from durations alone.

A trace is like a threaded conversation: the thread identifier groups the messages, each message has its own identifier, and the 'in reply to' field is what draws the tree.

saying these in an interview costs you the question

  • Says the span id is the same across the whole request
  • Thinks the parent stores a list of its children
  • Describes a span as a log line that happens at one instant
  • Believes the tracing backend assigns the identifiers
  • Assumes each service generates its own trace id
open as a page

In distributed tracing, what must travel with an outbound request for the callee to continue the same trace, and what breaks when one hop drops it?

level: middleimportance: must knowfreq 70%

basics

~20 s

The request must carry the trace id, the id of the span making the call, and whether the trace is being recorded. A hop that drops them makes the callee start a brand-new trace, taking every service below it with it.

open as a page

In distributed tracing, why is a trace assembled at read time, and what does a missing intermediate span do to the tree?

level: middleimportance: should knowfreq 48%

basics

~20 s

Each process exports its spans independently, so the backend stores them individually keyed by trace id and joins them into a tree only when someone queries. A missing intermediate span orphans everything beneath it and silently hides that hop's time.

open as a page

In a distributed trace, a child span starts before its parent and a queued job is nested under the wrong caller. What causes each?

level: seniorimportance: nice to knowfreq 26%

basics

~20 s

Clock skew explains the first: start times come from each host's own wall clock. The second is that a span's parent records whatever context was active at creation, which stops matching causality once work is queued or pooled.

open as a page