skip to content

In distributed tracing, why is a trace assembled at read time, and what does a missing intermediate span do to the tree?

level: middleimportance: should knowfreq 48%

answer

  1. Nobody knows when a request is finished
  2. Stored per span, joined on read
  3. Parent reference points at nothing
  4. Orphaned subtree, plausible-looking trace
  5. Missing hop's time lands on its caller

basics

~20 s

Each process exports its spans independently, so the backend stores them individually keyed by trace id and joins them into a tree only when someone queries. A missing intermediate span orphans everything beneath it and silently hides that hop's time.

solid answer

~50 s

No participant knows when a request is finished, so nothing can write a trace as a unit. Every process buffers its own finished spans and ships them in batches; they land out of order, from many machines, over a window of seconds. The backend indexes each span by trace id and, on a query, fetches them all and links each child to the span its parent reference names. If an intermediate span never arrives — an uninstrumented proxy, an export dropped under backpressure, a process killed before the span ended — its children point at an identifier the store has never seen. They become **orphans**: shown detached, re-parented to the root, or hidden entirely. The dangerous part is that the trace still looks plausible, and the missing hop's latency is silently attributed to its caller.

go deeper

for a junior

Recall that services export their own spans separately and that the tree you see is stitched together from them. Knowing that a trace can look incomplete simply because not everything has arrived yet is the useful takeaway.

for a middle

Explain the mechanics: spans indexed by trace id, children linked to parents at query time, orphans when a parent reference resolves to nothing. Be able to say why holding a trace until it is complete would need a stateful, consistently-sharded collection tier.

for a senior

Show you can diagnose from a partial trace: distinguish a hop that is uninstrumented from one whose export was dropped, and from a process that died with the span still open. Say what unattributed time in a parent actually means.

for a principal

Own the tradeoff between storage cost, retention and the completeness people expect from a trace view, and set the expectation across teams that a trace is a query result over independent records rather than a transaction.

## Spans are exported independently, by every process, at different times Consider a district-heating billing platform spread over 41 services. One customer request touches a dozen of them. Each process creates its own spans, holds finished ones in a small in-memory queue, and ships them in batches — often through a local collection agent and then a shared gateway tier — on its own schedule. Nothing coordinates those exports. Spans belonging to one request therefore arrive at the storage backend out of order, from many machines, over a window that can be several seconds wide, and occasionally minutes wide when a batch is retried after a gateway restart. The backend writes each span as it lands, indexed by its trace id. When somebody opens a trace, the store fetches every span carrying that trace id and links each one to the span named by its parent reference. **The tree is a view computed at query time, not a document that was written.** ## What a missing span does to the view Because the parent reference lives on the child, a child whose parent never arrived points at an identifier the backend has never seen: an **orphan**. Renderers differ in how they cope — some attach orphans to the root, some display disconnected fragments side by side, some hide them — but the information loss is the same in all of them. | What is missing | What the reader sees | What is actually lost | |---|---|---| | one intermediate span | its children hang detached, or are re-parented to the root | the hop's identity and duration; its time is silently attributed to the caller | | the root span | no top-level bar; the earliest orphan is often promoted to look like one | the end-to-end duration of the request | | a leaf span | nothing obviously wrong at all | a whole sub-operation, with no signal that it existed | | the trailing async work | the trace looks finished when it is not | anything that happened after the response was returned | The intermediate case is the dangerous one, because the trace still looks plausible. The gap between the caller's start and the surviving grandchildren reads as "the caller was slow", when in fact an uninstrumented proxy, a service whose export was dropped under backpressure, or a process that was killed while the span was still open sat in between. Spans that never end are never exported, so a crash removes exactly the span covering the crash. ## Why the assembly is deferred to read time 1. **Nothing signals that a trace is complete.** There is no end-of-trace event and no participant that knows the full membership. The root can finish and return a response while background work it triggered is still producing spans. 2. **Spans arrive late and out of order.** Anything written as a single unit would have to be found and rewritten on every late arrival, which is the most expensive access pattern a telemetry store has. 3. **Holding a trace together costs real money.** To buffer a whole trace you must route every span carrying the same trace id to the same process and hold it for a completeness window, which means a stateful, consistently-sharded collection tier. That machinery is exactly what a deferred retention decision has to pay for, and there is no reason to pay it merely to store. 4. **Almost no trace is ever read.** A fleet may record millions and a human opens a few hundred. Doing the join for the handful someone actually opens is cheap; doing it up front for all of them is waste. ## Practical consequences worth naming - A trace opened moments after the request may be genuinely incomplete and then fill in on refresh, so "spans are missing" often means "not yet". - A link followed straight from a log line into a tracing UI can briefly find nothing, because the span carrying the identifier had not been exported when the log line was written. - A very wide trace — say 2,318 spans from a fan-out over the whole estate — is expensive to assemble and many UIs truncate it, which looks identical to data loss. - Because assembly is a join on the trace id, a duplicated or reused trace id merges two unrelated requests into one nonsensical tree; identifiers must be wide and random for that reason. The senior version of this answer is that a trace is a *materialised query result over independent records*, not a transaction. Everything odd about traces in production — partial views, orphans, traces that grow after you open them — follows from that one fact.

  • A trace opened right after a request looks incomplete, then has more spans a minute later. Is that data loss?
    Usually not. Exporters batch, and a batch may be retried through several collection hops, so spans of one request land over a window. Because the tree is joined at query time, each read sees whatever has arrived. Genuine loss looks different: it persists, and it correlates with a queue-full or export-failure signal on a specific service.
  • What goes wrong if two unrelated requests end up with the same trace id?
    Assembly is a join on that key, so the backend merges them into one tree. You get a nonsensical trace with two roots, durations that overlap impossibly, and services that never spoke to each other appearing as siblings. It is why identifiers must be wide random values rather than anything derived from request content.

saying these in an interview costs you the question

  • Thinks each service sends its spans to the request's originator
  • Believes the backend waits for a trace to be complete
  • Says a missing span only removes one bar
  • Assumes spans arrive in causal order
  • Claims the root span writes the finished trace