skip to content

Explain the X-Amzn-Trace-Id header that AWS X-Ray uses to correlate one request across services: which fields it carries, which components set it first, and what the Sampled flag values mean to a service that receives it.

level: middleimportance: should knowfreq 45%

answer

  1. three fields, semicolon separated
  2. the middle field of Root is a timestamp
  3. one field makes the whole trace agree
  4. question mark means nobody decided yet
  5. a dropped header means a second trace

basics

~20 s

X-Amzn-Trace-Id carries Root (the X-Ray trace id), Parent (the caller's segment or subsegment id) and Sampled (1, 0 or ?). AWS entry points such as an Application Load Balancer or a REST API Gateway stage add it; every hop must forward it or the trace breaks.

solid answer

~50 s

The header looks like `Root=1-67891233-abcdef012345678912345678;Parent=463ac35c9f6413ad;Sampled=1`. **Root** is the X-Ray trace id, whose middle field is the request start time in hex epoch seconds. **Parent** is the id of the caller's segment or subsegment, which the receiver copies into its own segment's `parent_id` so the tree can be rebuilt. **Sampled** is `1` (trace it), `0` (do not) or `?` (nobody has decided yet — the receiving service should apply its sampling rules and pass on the decision it made). An Application Load Balancer adds the header when a request arrives without one; API Gateway REST stages with tracing enabled do the same, and Lambda exposes the incoming value in the `_X_AMZN_TRACE_ID` environment variable. The failure mode is prosaic: any service that drops the header instead of forwarding it starts a brand-new trace, and your service map splits in two.

code

bash · 8 lines
bash
# What an ALB adds when a client sends no trace header
curl -s -D - -o /dev/null https://app.example.com/orders

# What an instrumented service forwards on the next hop
# X-Amzn-Trace-Id: Root=1-67891233-abcdef012345678912345678;Parent=463ac35c9f6413ad;Sampled=1
#   Root    -> the trace id, shared by every segment of this request
#   Parent  -> the caller's subsegment id, becomes parent_id downstream
#   Sampled -> 1 record, 0 do not, ? decide and pass on the result

go deeper

for a junior

Know the header's name and that it carries Root, Parent and Sampled, and that every service must pass it along for the trace to stay in one piece.

for a middle

Explain each field precisely — the timestamp inside Root, Parent becoming the receiver's parent_id, and the three Sampled values including what ? obliges the receiver to do.

for a senior

Diagnose a split trace: find the first hop where the header changes or disappears, recognise pass-through boundaries and unpatched clients, and be explicit about asynchronous hops where propagation must be hand-carried.

for a principal

Set the propagation contract for the estate — which edges mint trace ids, that sampling is decided once, and how async boundaries are handled — so that end-to-end tracing is a property teams can rely on rather than a per-service accident.

## The wire format X-Ray correlates work across processes with a single HTTP header, `X-Amzn-Trace-Id`, carrying semicolon-separated key/value pairs: ``` X-Amzn-Trace-Id: Root=1-67891233-abcdef012345678912345678;Parent=463ac35c9f6413ad;Sampled=1 ``` **Root** is the trace id and has a fixed structure: a version digit, the time the trace started as hex epoch seconds, and 96 bits of random hex — `1-67891233-abcdef012345678912345678`. Embedding the start time is not decoration; it lets X-Ray bucket and expire traces by time without opening them, and it is why you cannot invent an arbitrary trace id. **Parent** is the 64-bit hex id of the *caller's* segment or subsegment — usually the subsegment the caller opened for this outbound call. The receiver copies it into `parent_id` on the segment it creates, which is the entire mechanism by which a flat pile of independently-uploaded segments becomes a tree. **Sampled** is the sampling decision, propagated so that the whole trace agrees. Additional arbitrary key/value pairs are permitted in the header; unknown keys are forwarded rather than rejected. ## Who puts it there first Something at the edge has to mint the Root, because a browser or a third-party client will not send one: - An **Application Load Balancer** adds `X-Amzn-Trace-Id` with a `Root` to any request that arrives without one, and leaves an existing one alone. This happens regardless of whether you use X-Ray at all. - **API Gateway** REST API stages with active tracing enabled generate the trace id and pass it to the integration. - **Lambda** places the incoming header value in the `_X_AMZN_TRACE_ID` environment variable of the execution environment, which is where the SDK picks it up. - Your own **front-end service** mints one if it is the entry point and nothing upstream did. Downstream, the X-Ray SDKs inject the header automatically on outbound calls made through instrumented clients — an instrumented HTTP client, an AWS SDK client. It is the calls the SDK *cannot* see that break the chain. ## The three Sampled values This is the field candidates most often get wrong. - **`Sampled=1`** — the trace is being recorded. A receiver should record its segment too, *without* re-running its own sampling rules. Sampling has to be a whole-trace decision or you get holes in the middle. - **`Sampled=0`** — the trace is not being recorded. A receiver should not upload a segment for it, though it must still forward the header so that everyone downstream agrees. - **`Sampled=?`** — no decision has been made yet. The header exists (so the Root is fixed), but the component that created it did not decide. The receiving instrumented service applies its own sampling rules, records or does not record accordingly, and **passes on the concrete `1` or `0`** to everything it calls. A service running in *pass-through* mode — for example a Lambda function with tracing set to `PassThrough` rather than `Active` — does not record anything itself but still forwards the header untouched, so a downstream service that *is* tracing still joins the same trace. ## What breaks in practice Almost every "my trace is split into two traces" incident is one of a small set of causes: 1. **A hop that does not forward the header.** A hand-rolled HTTP client the SDK never patched; a queue or event hop where there is no HTTP header to carry it at all; a proxy configured to strip unknown headers. 2. **A pass-through boundary nobody knew about.** Work is handed to a component that neither records nor forwards, and the next service mints a fresh Root. 3. **A background thread or async continuation** that lost the in-process context, so the outbound call is made with no header to inject. 4. **Sampling disagreement**, where a middle service ignores an incoming `Sampled=1` and re-decides — producing a trace with a visible gap where that service should be. The diagnosis is nearly always the same: capture the actual header at each hop (log it at debug on entry and before each outbound call) and find the first hop where `Root` changes or the header vanishes. ## Non-HTTP hops There is no header on a queue message. If a request crosses SQS, SNS or an event bus, propagating trace context is something you do explicitly — carrying the value in a message attribute and re-establishing it on the consumer side — and many teams simply accept that the trace ends at the queue and a new one begins on the other side. Being honest about that boundary in an interview is better than claiming end-to-end tracing you do not have.

  • A service receives Sampled=1 but its own sampling rule says 5%. What should it do?
    Record the segment. The sampling decision belongs to the trace, not to each hop: once the entry point decided to sample, every downstream service records unconditionally, or you get a trace with a hole where the middle service should be. Local rules only apply when the decision is still open — `Sampled=?` or no header at all — and the result is then propagated as a concrete 1 or 0.
  • How does trace context cross an SQS queue, where there is no HTTP header?
    Not automatically. You have to carry the context yourself — typically as a message attribute the producer sets and the consumer reads to re-establish the trace — or accept that the trace ends at the queue and a fresh one starts on the consumer side. Interviewers like this one because the honest answer is that asynchronous hops are where most claimed end-to-end tracing quietly stops.
  • Why does the X-Ray trace id embed a timestamp rather than being fully random?
    So the service can locate and expire traces by time without reading their contents. The middle field of Root is the trace start time in hex epoch seconds, which lets X-Ray bucket traces into time windows for search and retention. It also means you cannot mint an arbitrary id: one whose timestamp is far from the present is rejected rather than stored.

saying these in an interview costs you the question

  • Thinking Sampled=? means sampling was disabled
  • Letting each service re-run its own sampling rules mid-trace
  • Assuming the SDK propagates context across a queue automatically
  • Believing the ALB only adds the header when X-Ray is enabled
  • Treating Parent as the trace id rather than the caller's segment id

context