skip to content

In distributed tracing, what must travel with an outbound request for the callee to continue the same trace, and what breaks when one hop drops it?

level: middleimportance: must knowfreq 70%

answer

  1. It rides in the request's metadata
  2. Which trace, which caller, whether recorded
  3. Send your own span id, not upstream's
  4. The callee becomes a new root
  5. The whole subtree moves, not one edge

basics

~20 s

The request must carry the trace id, the id of the span making the call, and whether the trace is being recorded. A hop that drops them makes the callee start a brand-new trace, taking every service below it with it.

solid answer

~50 s

Three things ride along with the request, in whatever metadata the transport offers — HTTP headers, RPC metadata, message headers: the **trace id**, the **span id of the span currently making the call** (which becomes the callee's parent), and the **recording decision**. Both ends must agree on the encoding; W3C Trace Context is today's standard and older estates still speak B3-style headers. The receiver creates its server span with the same trace id and that parent, then re-sends the context with its own span id substituted on every onward call. If a hop drops the context, the callee finds nothing and starts a *new* trace — and every service beneath it joins that new trace. You get two well-formed traces with no key joining them, while the caller's side shows the hop as an opaque box that took some time and did nothing.

code

pseudocode · 11 lines
pseudocode
// caller, immediately before sending
headers[TRACE_CONTEXT] = encode(currentTraceId,
                                currentSpanId,   // NOT the id received from upstream
                                isRecorded)

// callee, on receiving the request
ctx = decode(headers[TRACE_CONTEXT])
serverSpan = startSpan(name,
                       traceId      = ctx.traceId,
                       parentSpanId = ctx.spanId)
// no context received -> traceId is freshly generated, parentSpanId is empty

go deeper

for a junior

Recall that trace context rides in the request's own metadata, and that a service which receives none starts a fresh trace instead of joining one. Knowing that both sides must speak the same encoding is enough here.

for a middle

Explain the chain: read the context, create a server span under the same trace id with the received span id as parent, then re-send with your own span id substituted. Be precise that a service forwards its current span id, not the one it received.

for a senior

Demonstrate that you can reason about blast radius — a break re-roots an entire subtree, not one edge — and name the real culprits: header-stripping proxies, hand-written clients, broker hops, in-process async boundaries, half-finished format migrations.

for a principal

Own propagation as a fleet-wide invariant: a single agreed encoding, a migration plan for the seam, conformance tested at boundaries you do not own, and an explicit decision about which edges are allowed to be opaque.

## What has to cross the boundary For a callee to be part of the caller's trace rather than a trace of its own, three things must travel with the request itself: 1. **The trace id**, so both sides file their spans under the same key. 2. **The identifier of the span that is making the call**, which becomes the parent reference on the span the callee creates. Note the substitution that catches people out: a service sends *its own currently executing span id*, not the one it received from upstream. 3. **Whether this trace is being recorded**, so that a request either produces a complete trace or produces none, instead of a half-recorded one. They ride in whatever metadata channel the transport offers — headers on an HTTP request, request metadata on an RPC call, message headers or properties on a broker record. Both ends must agree on the encoding: the modern standard is W3C Trace Context, while older estates often speak the B3-style headers popularised by earlier tracing stacks, and a fleet part-way through a migration needs receivers that accept both. ## What the receiving side does The callee reads those fields, creates its server-side span with the **same** trace id and with its parent set to the received span id, and makes that span the current one for the duration of the request. From then on every span the callee creates inherits that trace id, and every outbound call it makes carries the context onward with its own span id substituted in. Propagation is therefore a chain: it works only if every link in it both reads and re-writes. ## What one dropped hop does This is the part candidates underestimate. A hop that fails to forward context does **not** cost you one edge in the tree. - The callee looks for context, finds none, and does the only sensible thing: it starts a **new trace** with a brand-new trace id and no parent. - Everything below it continues *that* trace, correctly and completely. A break at the third of eleven hops relocates eight hops' worth of spans. - You end up with two well-formed traces and no key that joins them. Searching the original trace id will never surface the second half; there is nothing on either side pointing at the other. - In the caller's trace, the outbound span still exists and still has its duration, so the hop looks like an opaque black box that took 380 ms and did nothing else. That is the misleading part: the view is not obviously broken, it is quietly shallow. - If the break is systematic — a shared gateway, a shared client library — a large share of the estate looks uninstrumented, and teams get asked to "add tracing" to services that already have it. A subtler variant drops only the recording decision while preserving the identifiers. The trace id survives, so the tree still assembles, but the downstream side decides independently whether to keep what it records, and traces come back with holes in them rather than split in two. ## Where it breaks in practice - **A proxy, gateway or WAF that forwards an allowlist of headers** and silently strips everything it does not recognise. - **A hand-written client** — very common when a team is moving off a hosted vendor mid-quarter and reimplements a call path in a hurry — that copies the headers somebody remembered. - **A queue or topic hop** where the producer never attaches context to the message, so the consumer has nothing to read even though both sides are instrumented. - **An in-process asynchronous boundary**, where the outbound call is made from a worker that no longer has the originating request's context attached to it. - **A partial format migration**, where half the fleet emits one encoding and the other half only understands the other. - **A retry or edge layer that rebuilds the request** from a stored copy of the body and a fixed header set. ## What you can conclude from the symptom Two disconnected traces whose root sits at a service in the middle of your topology is almost always a propagation break at the hop immediately above that service, not a missing instrumentation library in the service itself — the service is clearly instrumented, since it produced the root. The fix is at the boundary that dropped the context, and the test is whether the fields survive an end-to-end request through every proxy on the path, not whether each service works in isolation.

  • A team reports that half their services 'have no tracing'. Their traces are rooted at a service in the middle of the topology. What is your first hypothesis?
    Those services are instrumented — they produced the root you are looking at. The likely fault is the hop immediately above them dropping context: a proxy forwarding only an allowlist of headers, or a client that rebuilds the request. Test with one end-to-end call through every intermediary and check what actually arrives, rather than auditing each service in isolation.
  • How does this work across a message broker, where there is no request to attach anything to?
    The producer writes the context onto the message's own headers or properties, and the consumer reads it from there. The awkward part is semantics rather than transport: the consumer may run long after the producer finished, so a strict parent-child edge misrepresents it and a non-hierarchical reference between the two spans is usually the better model.
  • What happens when one half of a fleet speaks a different propagation format from the other?
    Requests crossing that seam behave exactly like a dropped hop — the receiver cannot read what arrived, so it starts a new trace. The usual remedy during a migration is to have receivers accept both encodings while senders emit both, then retire the old one only once nothing reads it.

saying these in an interview costs you the question

  • Says the caller forwards the span id it received upstream
  • Thinks a dropped hop loses only that one span
  • Believes the backend can rejoin the two halves later
  • Assumes context propagates because both services are instrumented
  • Confuses carrying context with sending spans to the next service