skip to content

X-Ray Distributed Tracing

Following one request across Lambda, API Gateway, and downstream services. You instrument with the X-Ray SDK or ADOT, control what gets sampled, and read the service map to find the hop that is actually slow.

part ofAWSoverview, primer and where to startread it →
on this pageshow

questions

6

When instrumenting a service with the AWS X-Ray SDK, when should a key/value go on a segment as an annotation versus as metadata, and what does that choice change about finding the trace later?

level: middleimportance: must knowfreq 58%

answer

  1. one of the two is indexed
  2. filter expressions need a searchable key
  3. the index has a per-trace budget
  4. arbitrary JSON goes in the other one
  5. never the customer's email address

basics

~20 s

Annotations are indexed by X-Ray and searchable in filter expressions; metadata is stored with the segment but not indexed. Put the few fields you will search traces by in annotations, and everything else in metadata.

solid answer

~50 s

Both attach data to a segment or subsegment, but only **annotations are indexed**. An annotation is a key with a string, number or boolean value, and once set you can search for it in a filter expression — `annotation.order_id = "A-1234"` — which is how you go from a customer complaint to that customer's trace. **Metadata** takes arbitrary JSON of any size and shape, is stored with the segment and is visible when you open the trace, but is invisible to search. The design rule follows from the limits: annotations are indexed and therefore capped (X-Ray indexes up to 50 annotations per trace), so treat them as a deliberate, small set of search keys — tenant, order id, feature flag, deployment version — and use metadata as the roomy payload for request bodies, decision traces and diagnostic dumps. Never put secrets or personal data in either.

code

python · 18 lines
python
from aws_xray_sdk.core import xray_recorder, patch_all

patch_all()

@xray_recorder.capture("price_order")
def price_order(order):
    segment = xray_recorder.current_segment()
    # indexed: searchable as annotation.tenant_id / annotation.order_id
    segment.put_annotation("tenant_id", order["tenant"])
    segment.put_annotation("order_id", order["id"])
    segment.put_annotation("promo_applied", True)

    # not indexed: visible on the trace, invisible to search
    segment.put_metadata("pricing_inputs", {
        "line_items": order["items"],
        "rules_evaluated": ["vol_discount", "loyalty_tier"],
    })
    return compute_total(order)

go deeper

for a junior

Remember the one-line rule: annotations are indexed and searchable, metadata is stored and not. Know the two SDK calls that set them.

for a middle

Explain how an annotation reaches a filter expression as annotation.<key>, why values are limited to string, number or boolean, and why the indexed set is deliberately small.

for a senior

Show judgment about which keys earn an annotation — lookup keys versus grouping keys, the per-trace index budget across many services, and the rule that no personal data or secret goes into trace data at all.

for a principal

Own the annotation vocabulary as a cross-team contract: a short reviewed list where tenant_id means the same thing in every service, plus a redaction standard so the trace store never becomes an uncontrolled copy of customer data.

## Two attachment points, one difference that matters The X-Ray SDKs give you two ways to hang your own data on a segment or subsegment: ```python from aws_xray_sdk.core import xray_recorder segment = xray_recorder.current_segment() segment.put_annotation("tenant_id", "acme") # indexed, searchable segment.put_metadata("request_body", payload) # stored, not searchable ``` The two land in different places in the segment document — an `annotations` object and a `metadata` object — and X-Ray treats them completely differently on ingest. Annotations are extracted and **indexed**; metadata is stored as an opaque blob attached to the segment. ## What indexing buys you X-Ray's console and the `GetTraceSummaries` API take a **filter expression**. Built-in fields are always available: ``` service("checkout-api") { fault } http.status = 500 AND responsetime > 2 ``` Indexed annotations extend that vocabulary with your own domain: ``` annotation.tenant_id = "acme" AND responsetime > 2 annotation.order_id = "A-1234" ``` That is the whole point. Distributed tracing is only operationally useful if you can get from a *business* fact — this customer, this order, this experiment arm — to the trace. Without an annotation you are reduced to guessing at a time window and scanning. Annotation values are restricted to **string, number or boolean**, because they have to be indexable. Keys are restricted to alphanumerics and underscore. An object will not be accepted as an annotation value. ## What metadata is for Metadata takes arbitrary JSON — nested objects, arrays, long strings — organised into namespaces. You cannot search it, but it is fully visible when someone opens the trace, which makes it the right home for: - the request or response body (redacted), - the inputs to a pricing or routing decision, - feature-flag evaluation results, - anything you would otherwise have logged "just in case". The SDK's default namespace is `default`; the `aws` namespace is reserved for data the SDK itself records. ## Why you cannot annotate everything The obvious junior instinct is to annotate every field, since annotations are strictly more useful. Two things stop you. First, **the index is bounded**. X-Ray indexes up to 50 annotations per trace; beyond that the extra ones are still stored but you cannot rely on searching them. Fifty across an entire multi-service trace is a small budget once every service starts annotating freely. Second, **high-cardinality keys are a trap in the other direction**. An annotation whose value is unique per request — a full request id, a timestamp — is perfectly valid and genuinely useful for pinpoint lookup, but it is useless for grouping. Keys you want to *group* by (tenant, region, api version) should be low-cardinality; keys you want to *look up* by (order id) can be unique. Deciding which of the two a key is, before you add it, is the design conversation. A sensible convention is to reserve annotations for a short, reviewed list agreed across services — so that `annotation.tenant_id` means the same thing everywhere — and let metadata absorb everything else. ## Subsegments count too Both calls work on subsegments as well as segments, and annotations on a subsegment are indexed for the trace. Annotating the subsegment is often better hygiene: the fact that *this particular downstream call* was for tenant `acme` is more precise than a segment-level annotation when a request fans out. ## The rule that overrides both Neither annotations nor metadata are an appropriate place for **secrets or personal data**. Trace data is retained by X-Ray (traces are kept for 30 days) and is readable by anyone with X-Ray read access in the account, which is typically a much wider group than has access to production data stores. Annotating `customer_email` gives you a beautiful search experience and a compliance problem; annotate a pseudonymous customer id instead and join to the real identity in a system that has the right controls. If a field is sensitive and you truly need it in the trace, redact or hash it before it is attached. ## Reading it back When you find the trace, the console shows annotations in a dedicated section of the segment view and metadata beneath it; `BatchGetTraces` returns the whole segment document including both. So metadata costs you nothing at retrieval time — only at search time, where it does not exist.

  • You want to find every trace for one customer over the last hour. What has to have happened at instrumentation time?
    A stable customer identifier must have been attached as an *annotation*, not metadata, because only annotations are indexed and reachable from a filter expression such as `annotation.customer_id = "c-99"`. If it went into metadata, the data is in the traces but unsearchable, and your only recourse is scanning traces in the window by service and latency. Retrofitting is a deploy, so the annotation list is worth agreeing up front.
  • Is there a downside to annotating a value that is unique per request, such as a request id?
    Not for lookup — that is exactly what pinpoint search needs. The cost is that a unique value cannot be grouped or aggregated, so it tells you nothing on its own, and it consumes part of the per-trace annotation budget that X-Ray indexes. Keep such keys deliberate: one or two identifiers you will genuinely paste into a search box, and low-cardinality keys for anything you want to slice by.
  • Where should sensitive fields go if you need them for debugging?
    Neither place, unredacted. Trace data is retained for 30 days and readable by everyone with X-Ray read access, which is a wider audience than production data. Attach a pseudonymous identifier as the annotation and resolve it in a system with proper access control; if a payload must be attached at all, redact or hash the sensitive fields in the instrumentation before they reach the segment.

saying these in an interview costs you the question

  • Believing both annotations and metadata are searchable
  • Annotating every field because annotations are more useful
  • Putting a nested object in an annotation value
  • Storing customer email or tokens as annotations for easy search
  • Thinking metadata is lost rather than merely unindexed

context

open as a page

What changes when you switch an AWS Lambda function's tracing mode from PassThrough to Active, what segments does the resulting trace contain, and why might that function still appear as a trace of its own instead of joining its caller's?

level: seniorimportance: must knowfreq 62%

basics

~20 s

Active tracing makes the Lambda service record the invocation itself, producing a service segment plus a function segment with Initialization, Invocation and Overhead subsegments. The function joins its caller's trace only if the caller propagated the trace header and the execution role can write to X-Ray.

open as a page

In AWS X-Ray, what is the difference between a segment and a subsegment, and how does X-Ray turn the segments it receives into the service map shown in the console?

level: juniorimportance: should knowfreq 60%

basics

~20 s

A segment is what one service records for one request; subsegments break that segment into finer steps such as downstream calls. X-Ray links all segments sharing a trace id and aggregates them into the service map graph.

open as a page

Explain the X-Amzn-Trace-Id header that AWS X-Ray uses to correlate one request across services: which fields it carries, which components set it first, and what the Sampled flag values mean to a service that receives it.

level: middleimportance: should knowfreq 45%

basics

~20 s

X-Amzn-Trace-Id carries Root (the X-Ray trace id), Parent (the caller's segment or subsegment id) and Sampled (1, 0 or ?). AWS entry points such as an Application Load Balancer or a REST API Gateway stage add it; every hop must forward it or the trace breaks.

open as a page

An ECS service on Fargate is instrumented with the AWS X-Ray SDK, but no traces appear in the X-Ray console. Walk through how segment data actually reaches X-Ray from a container, and where you would look for the break.

level: seniorimportance: should knowfreq 48%

basics

~20 s

The SDK does not call X-Ray directly: it sends segments over UDP to a local collector — the X-Ray daemon or an ADOT collector sidecar on port 2000 — which batches them and calls PutTraceSegments. Check that the sidecar exists, that the app points at it, and that the task role grants X-Ray write.

open as a page

How does an AWS X-Ray sampling rule decide whether a given request is traced? Explain the reservoir and fixed-rate fields, how rule priority works, and how one reservoir is shared across many instances of the same service.

level: seniorimportance: nice to knowfreq 33%

basics

~20 s

An X-Ray sampling rule matches requests by service, host, method and URL path, then traces a per-second reservoir of them and a fixed percentage of the rest. Rules are evaluated by priority, lowest number first, and the reservoir is handed out to instances as quotas by the X-Ray service.

open as a page