skip to content

Downstream services are producing spans, but instead of continuing the caller's trace they start new traces of their own. Walk through how you would diagnose a cross-service context-propagation failure.

level: seniorimportance: should knowfreq 54%

answer

  1. New trace id = propagation; holes in one trace = sampling
  2. Log header presence at ingress before reading any config
  3. W3C vs B3, and B3 single vs multi
  4. All-zero or duplicated traceparent is invalid
  5. Gateways/meshes rebuild requests from header allowlists

basics

~20 s

Prove first whether the context header arrives at the callee. If it does not, something between them stripped it or the caller never injected. If it does, the callee's propagator format or extraction is wrong. Distinguish new trace ids (propagation failure) from missing spans (sampling).

solid answer

~60 s

Split the question in two before touching config. **Does the header arrive?** Capture the inbound request at the callee — a debug log of `traceparent`/`b3` presence, or a packet/proxy trace. If absent: the caller may not be instrumented on that client path (a hand-built HTTP client, a raw socket, a background job), or an intermediary stripped it — gateways, service meshes, CDNs and any code that rebuilds requests from a header allowlist are the usual suspects. **Does the callee understand it?** A format mismatch is the classic case: the caller emits B3 (`b3` single header or the multi-header `X-B3-*` set) while the callee is configured for W3C only, or vice versa. Fix with a composite propagator on both sides while you standardise. Then check identity: an all-zero or malformed trace id is invalid and ignored, and duplicated `traceparent` headers invalidate extraction. Finally distinguish symptoms — *new trace ids* mean extraction failed; *same trace id with holes* means a sampler decision or an exporter drop, which is a different bug.

code

text · 10 lines
text
client -> gateway -> svcA -> svcB -> svcC

log at each ingress: request_id, present(traceparent), trace_id

  gateway  rid=R1  traceparent=yes  trace=4bf9...
  svcA     rid=R1  traceparent=yes  trace=4bf9...
  svcB     rid=R1  traceparent=NO   trace=9c02...   <-- break is on the A->B leg
  svcC     rid=R1  traceparent=yes  trace=9c02...

then ask on that leg only: did A inject? did anything between strip it?

go deeper

for a junior

Check whether the trace header arrives at the callee, and know that a missing or unrecognised header makes the service start its own trace.

for a middle

Add propagator format mismatch (W3C versus B3) and the composite-propagator fix, and separate propagation symptoms from sampling symptoms.

for a senior

Bisect the hop chain systematically, cover intermediaries, header validity rules, and in-callee span-start ordering, then propose regression tests.

for a principal

Set estate-wide propagation policy — one canonical format, a bounded dual-emit window, perimeter trust rules for inbound context, and a root-span-rate alarm so breaks are detected rather than reported.

## Establish the symptom precisely Two problems are routinely confused and they have disjoint causes. - **New trace id at the callee.** The callee did not extract a parent, so it started a root. This is propagation. - **Same trace id, missing spans in the middle.** Context propagated fine; something declined to record or export. This is sampling, exporter failure, or a dropped batch. Compare the trace ids on both sides of the suspect hop for the same logical request — correlate by request id, by timestamp, or by logging the trace id in application logs at each service. If the ids differ, continue with the propagation checklist below. If they match, stop and go look at samplers and exporters instead. ## Step 1: does the header reach the callee? Instrument the question rather than reasoning about config. Add a temporary debug log at the callee's ingress that records whether `traceparent`, `tracestate`, `b3` and `X-B3-TraceId` are present (log presence and the trace id, not full headers indiscriminately). Alternatively inspect at the proxy layer or capture traffic. **Header absent.** Candidates, roughly in order of frequency: - The caller's outbound path is not instrumented. Auto-instrumentation covers well-known clients; a hand-rolled client, a raw socket, a shell-out to a command-line HTTP tool, or a background job that constructs requests directly gets nothing. - The caller *is* instrumented but had no context to inject — an asynchronous hand-off dropped it inside the caller, so it injected an empty context. Check whether the caller's own span for that outbound call exists. - An intermediary removed it. API gateways that rebuild requests from an allowlist, service meshes with header policies, CDNs, WAFs, and load balancers configured to strip unknown headers all do this. Test by calling the callee directly, bypassing the intermediary, with a hand-crafted `traceparent`. **Header present.** Move to step 2. ## Step 2: format and configuration mismatch The most common cause of "header arrives, still a new trace" is that the two sides speak different propagation formats. Zipkin-lineage systems use B3, either as a single `b3: <traceid>-<spanid>-<sampled>` header or the multi-header `X-B3-TraceId` / `X-B3-SpanId` / `X-B3-Sampled` form; OpenTelemetry SDKs default to W3C `traceparent`/`tracestate`. A service configured only for W3C ignores a B3 header entirely, and the reverse is also true. Check the configured propagators on both sides (in OTel SDKs this is typically an environment setting listing the propagators to compose, or explicit code at SDK initialisation). During migration, configure a composite propagator that injects both formats and extracts whichever is present, then remove the legacy format once every service understands the new one. Note the cost: dual injection means more header bytes on every request, so treat it as a bounded migration state. Also check B3 sub-format mismatches — single-header versus multi-header — which fail the same way even when both sides say "we use B3". ## Step 3: validity and duplication Even a correctly named header can be rejected: - An all-zero trace id or span id is invalid by specification and must be ignored; some hand-written clients emit zeros as a placeholder. - Wrong lengths, uppercase hex, or a malformed version field make the header invalid. - A duplicated `traceparent` (two intermediaries each adding one) invalidates the context. - Case-sensitive header lookup in a custom getter can miss `Traceparent` from a non-conforming client. ## Step 4: extraction happened but was not used Occasionally the SDK extracts correctly and the application still roots a new trace. Causes: the server span is started *before* the extracted context is made current; framework instrumentation is registered after the handler that creates the span; or application code explicitly starts a span with a root parent, overriding what was extracted. Verify by logging the extracted trace id at ingress — if the log shows the caller's id but the exported span shows a different one, the problem is inside the callee's own span-start ordering, not on the wire. ## Step 5: think about trust, not just plumbing Sometimes the "failure" is deliberate. Perimeter services often strip inbound trace context so external callers cannot choose your trace ids or force the sampled bit. If traces reliably start at the edge gateway and nowhere earlier, confirm whether that is policy before you "fix" it. ## Prevention Make propagation a tested property rather than an observed one: a smoke test that issues a request with a known `traceparent` and asserts that every service in the path reports the same trace id; a dashboard of root-span rate per service (a service that should never be a root suddenly producing roots is the alarm); and a standard, centrally managed propagator configuration so no service quietly diverges.

  • Trace ids match across services but the middle service's spans are missing from the backend. Where do you look?
    Not at propagation — the context clearly arrived. Check the sampler first: if that service uses an independent ratio sampler instead of the parent-based default, it makes its own decision and drops traces the edge sampled. Then check the export path: exporter errors, a full batch queue dropping spans under load, or a collector rejecting the payload. SDK self-diagnostics and collector queue and refused-spans metrics distinguish those quickly.
  • Half your estate emits B3 and half emits W3C. What is the migration plan?
    Configure a composite propagator everywhere that *extracts* both formats first, so no context is lost regardless of which side a request comes from. Then enable dual injection so legacy consumers keep working, and migrate services to the target format. Once a dashboard confirms no service is still extracting from the legacy header, turn off legacy injection to reclaim the header bytes. Keeping dual emission permanently is a cost on every request in the estate.
  • Should a public-facing service accept `traceparent` from the internet?
    Usually not without thought. Accepting it lets any caller choose your trace ids — enabling collisions, deliberate pollution of existing traces, and forcing the sampled bit to make you record and store expensive traces on demand. The common posture is to strip inbound context at the perimeter and start a fresh trace, optionally recording the client-supplied id as an attribute, while accepting real context only from trusted partners over authenticated channels.

saying these in an interview costs you the question

  • Changing propagator configuration before establishing whether the header even arrives.
  • Confusing missing spans (sampling or export) with new trace ids (propagation).
  • Assuming both sides interoperate because both "use B3" — single-header and multi-header forms do not.
  • Overlooking gateways, meshes and CDNs that rebuild requests from a header allowlist.
  • Forgetting that an all-zero or duplicated `traceparent` is invalid and silently ignored.
  • Treating deliberate perimeter stripping as a bug without checking whether it is policy.

context