skip to content

When instrumenting a service with the AWS X-Ray SDK, when should a key/value go on a segment as an annotation versus as metadata, and what does that choice change about finding the trace later?

level: middleimportance: must knowfreq 58%

answer

  1. one of the two is indexed
  2. filter expressions need a searchable key
  3. the index has a per-trace budget
  4. arbitrary JSON goes in the other one
  5. never the customer's email address

basics

~20 s

Annotations are indexed by X-Ray and searchable in filter expressions; metadata is stored with the segment but not indexed. Put the few fields you will search traces by in annotations, and everything else in metadata.

solid answer

~50 s

Both attach data to a segment or subsegment, but only **annotations are indexed**. An annotation is a key with a string, number or boolean value, and once set you can search for it in a filter expression — `annotation.order_id = "A-1234"` — which is how you go from a customer complaint to that customer's trace. **Metadata** takes arbitrary JSON of any size and shape, is stored with the segment and is visible when you open the trace, but is invisible to search. The design rule follows from the limits: annotations are indexed and therefore capped (X-Ray indexes up to 50 annotations per trace), so treat them as a deliberate, small set of search keys — tenant, order id, feature flag, deployment version — and use metadata as the roomy payload for request bodies, decision traces and diagnostic dumps. Never put secrets or personal data in either.

code

python · 18 lines
python
from aws_xray_sdk.core import xray_recorder, patch_all

patch_all()

@xray_recorder.capture("price_order")
def price_order(order):
    segment = xray_recorder.current_segment()
    # indexed: searchable as annotation.tenant_id / annotation.order_id
    segment.put_annotation("tenant_id", order["tenant"])
    segment.put_annotation("order_id", order["id"])
    segment.put_annotation("promo_applied", True)

    # not indexed: visible on the trace, invisible to search
    segment.put_metadata("pricing_inputs", {
        "line_items": order["items"],
        "rules_evaluated": ["vol_discount", "loyalty_tier"],
    })
    return compute_total(order)

go deeper

for a junior

Remember the one-line rule: annotations are indexed and searchable, metadata is stored and not. Know the two SDK calls that set them.

for a middle

Explain how an annotation reaches a filter expression as annotation.<key>, why values are limited to string, number or boolean, and why the indexed set is deliberately small.

for a senior

Show judgment about which keys earn an annotation — lookup keys versus grouping keys, the per-trace index budget across many services, and the rule that no personal data or secret goes into trace data at all.

for a principal

Own the annotation vocabulary as a cross-team contract: a short reviewed list where tenant_id means the same thing in every service, plus a redaction standard so the trace store never becomes an uncontrolled copy of customer data.

## Two attachment points, one difference that matters The X-Ray SDKs give you two ways to hang your own data on a segment or subsegment: ```python from aws_xray_sdk.core import xray_recorder segment = xray_recorder.current_segment() segment.put_annotation("tenant_id", "acme") # indexed, searchable segment.put_metadata("request_body", payload) # stored, not searchable ``` The two land in different places in the segment document — an `annotations` object and a `metadata` object — and X-Ray treats them completely differently on ingest. Annotations are extracted and **indexed**; metadata is stored as an opaque blob attached to the segment. ## What indexing buys you X-Ray's console and the `GetTraceSummaries` API take a **filter expression**. Built-in fields are always available: ``` service("checkout-api") { fault } http.status = 500 AND responsetime > 2 ``` Indexed annotations extend that vocabulary with your own domain: ``` annotation.tenant_id = "acme" AND responsetime > 2 annotation.order_id = "A-1234" ``` That is the whole point. Distributed tracing is only operationally useful if you can get from a *business* fact — this customer, this order, this experiment arm — to the trace. Without an annotation you are reduced to guessing at a time window and scanning. Annotation values are restricted to **string, number or boolean**, because they have to be indexable. Keys are restricted to alphanumerics and underscore. An object will not be accepted as an annotation value. ## What metadata is for Metadata takes arbitrary JSON — nested objects, arrays, long strings — organised into namespaces. You cannot search it, but it is fully visible when someone opens the trace, which makes it the right home for: - the request or response body (redacted), - the inputs to a pricing or routing decision, - feature-flag evaluation results, - anything you would otherwise have logged "just in case". The SDK's default namespace is `default`; the `aws` namespace is reserved for data the SDK itself records. ## Why you cannot annotate everything The obvious junior instinct is to annotate every field, since annotations are strictly more useful. Two things stop you. First, **the index is bounded**. X-Ray indexes up to 50 annotations per trace; beyond that the extra ones are still stored but you cannot rely on searching them. Fifty across an entire multi-service trace is a small budget once every service starts annotating freely. Second, **high-cardinality keys are a trap in the other direction**. An annotation whose value is unique per request — a full request id, a timestamp — is perfectly valid and genuinely useful for pinpoint lookup, but it is useless for grouping. Keys you want to *group* by (tenant, region, api version) should be low-cardinality; keys you want to *look up* by (order id) can be unique. Deciding which of the two a key is, before you add it, is the design conversation. A sensible convention is to reserve annotations for a short, reviewed list agreed across services — so that `annotation.tenant_id` means the same thing everywhere — and let metadata absorb everything else. ## Subsegments count too Both calls work on subsegments as well as segments, and annotations on a subsegment are indexed for the trace. Annotating the subsegment is often better hygiene: the fact that *this particular downstream call* was for tenant `acme` is more precise than a segment-level annotation when a request fans out. ## The rule that overrides both Neither annotations nor metadata are an appropriate place for **secrets or personal data**. Trace data is retained by X-Ray (traces are kept for 30 days) and is readable by anyone with X-Ray read access in the account, which is typically a much wider group than has access to production data stores. Annotating `customer_email` gives you a beautiful search experience and a compliance problem; annotate a pseudonymous customer id instead and join to the real identity in a system that has the right controls. If a field is sensitive and you truly need it in the trace, redact or hash it before it is attached. ## Reading it back When you find the trace, the console shows annotations in a dedicated section of the segment view and metadata beneath it; `BatchGetTraces` returns the whole segment document including both. So metadata costs you nothing at retrieval time — only at search time, where it does not exist.

  • You want to find every trace for one customer over the last hour. What has to have happened at instrumentation time?
    A stable customer identifier must have been attached as an *annotation*, not metadata, because only annotations are indexed and reachable from a filter expression such as `annotation.customer_id = "c-99"`. If it went into metadata, the data is in the traces but unsearchable, and your only recourse is scanning traces in the window by service and latency. Retrofitting is a deploy, so the annotation list is worth agreeing up front.
  • Is there a downside to annotating a value that is unique per request, such as a request id?
    Not for lookup — that is exactly what pinpoint search needs. The cost is that a unique value cannot be grouped or aggregated, so it tells you nothing on its own, and it consumes part of the per-trace annotation budget that X-Ray indexes. Keep such keys deliberate: one or two identifiers you will genuinely paste into a search box, and low-cardinality keys for anything you want to slice by.
  • Where should sensitive fields go if you need them for debugging?
    Neither place, unredacted. Trace data is retained for 30 days and readable by everyone with X-Ray read access, which is a wider audience than production data. Attach a pseudonymous identifier as the annotation and resolve it in a system with proper access control; if a payload must be attached at all, redact or hash the sensitive fields in the instrumentation before they reach the segment.

saying these in an interview costs you the question

  • Believing both annotations and metadata are searchable
  • Annotating every field because annotations are more useful
  • Putting a nested object in an annotation value
  • Storing customer email or tokens as annotations for easy search
  • Thinking metadata is lost rather than merely unindexed

context