skip to content

What is Baggage in OpenTelemetry, how does it differ from span attributes, and what risks come with adding fields to it?

level: middleimportance: must knowfreq 55%

answer

  1. Baggage = context that travels; attributes = one span, stays
  2. Not auto-copied onto spans — needs a processor
  3. 8192 bytes / 180 entries, dropped silently
  4. Inbound baggage is user input → strip or allowlist at edge
  5. Never for authorisation, never for PII

basics

~20 s

Baggage is a set of key/value pairs carried in the Context and propagated to every downstream service in the baggage header. Unlike span attributes, which stay on one local span, baggage travels everywhere — so it costs bytes on every hop, is untrusted when inbound, and is not automatically recorded on spans.

solid answer

~50 s

Baggage is application-defined key/value data that rides alongside the span context: it lives in the same Context object and is serialised into the W3C `baggage` header on every outbound call. Span attributes are the opposite — set locally on one span and exported straight to the backend, never propagated. Use baggage when a value known only at the entry point (tenant, deployment ring, request priority, experiment arm) must influence or annotate work several hops downstream. Three cautions. It is *not* automatically added to spans; something must copy it, usually a span processor, and that copy is where cardinality and PII problems appear — especially if it reaches metric labels. It is bytes on the wire for every request in the estate, and the specification caps the header at 8192 bytes and 180 entries. And inbound baggage from an untrusted client is attacker-controlled input that will be echoed into logs, spans and possibly to third parties, so strip or allowlist it at the perimeter.

code

text · 8 lines
text
edge service sets baggage {tenant=acme, ring=canary}
  -> outbound header  baggage: tenant=acme,ring=canary   (every hop, every call)

service C sets span attribute db.rows_returned=412
  -> exported with THAT span only. never on the wire again.

searchable as tenant=acme in the backend ONLY IF service C runs a processor:
  onStart(span): for k in ALLOWLIST: span.setAttribute(k, baggage[k])

go deeper

for a junior

Define baggage as key/values propagated with the request context, and contrast it with span attributes, which stay on one span.

for a middle

Add that it is not copied to spans automatically, name the size limits, and give one good use case and one bad one.

for a senior

Discuss perimeter stripping of untrusted inbound baggage, cardinality control when copying keys, and why behaviour must never depend on it.

for a principal

Treat the key set as a governed cross-service schema with owners, budgets and enforcement at the gateway, and argue where the implicit-API risk outweighs the observability benefit.

## What it is OpenTelemetry's Context is an immutable, propagation-agnostic container. Two things are commonly stored in it: the current span (from which the span context — trace id, span id, flags — is derived) and Baggage. Baggage is a simple ordered set of key/value string pairs, with optional per-entry metadata, defined by the W3C Baggage specification and serialised into a header of that name: ``` baggage: tenant=acme,ring=canary,priority=high ``` It is propagated by its own propagator, usually composed with the trace-context propagator, so a service configured for W3C emits both `traceparent` and `baggage`. The two are independent: baggage can exist without a sampled trace and vice versa. ## Baggage versus span attributes | | Baggage | Span attributes | |---|---|---| | Scope | the whole distributed request, every downstream hop | one span | | Transport | serialised into an HTTP/gRPC/broker header on each call | exported directly to the backend with the span | | Cost | bytes on every hop, plus header-size limits | payload size at export only | | Trust | inbound value may come from an untrusted caller | set by your own code | | Recorded automatically? | **No** | Yes, it is part of the span | That last row is the most common misconception. Putting `tenant=acme` in baggage does not make `tenant` searchable in your tracing backend. Something must read the baggage and set it as an attribute — typically a span processor that copies an allowlisted set of baggage keys onto every span it starts. Many SDKs and distributions ship such a processor, but it is opt-in and key-restricted for good reason. ## When baggage earns its place The legitimate use is *information known early that is needed late, by code that cannot look it up*. Examples: the tenant or customer tier, so a downstream service can label spans or make a shedding decision; a deployment ring or experiment arm, so telemetry can be split by cohort across services; a request priority for load shedding; a synthetic-traffic marker so load-test requests can be excluded from business metrics. The illegitimate uses are all variations of "use it as an RPC parameter". If a downstream service *behaves* differently based on baggage, you have created an implicit, untyped, unversioned API between services that no interface definition documents and no compiler checks. Worse, it is attacker-influenceable at any hop that does not sanitise. Anything load-bearing for correctness or authorisation belongs in the actual request contract; authorisation decisions must never read baggage. ## The risks in detail **Size and cost.** Baggage is added to every outbound request of every service in the path. A 200-byte baggage set on a fan-out of five services and ten calls each is real bandwidth and real header-parsing work, and it interacts badly with servers that cap total header size (a hop can start returning 431 or dropping requests when the accumulated set grows). The W3C specification limits the header to 8192 bytes and 180 entries, and implementations may drop entries when limits are hit — silently, from the perspective of the code that set them. **Cardinality.** If a processor copies baggage onto spans, high-cardinality values such as a user id are merely expensive in the trace store. If anything copies baggage into *metric* labels, high cardinality multiplies time series and can take down the metrics backend. Copy with an explicit allowlist of keys, never wholesale. **Confidentiality.** Baggage goes wherever the request goes, including to third-party APIs if your outbound HTTP client injects context indiscriminately, and into every log line that logs headers. Never put personal data, tokens or secrets in it. Consider restricting injection to internal destinations. **Trust.** Inbound baggage on a public endpoint is user input. Accepting it means an outsider can set `tenant=` or `priority=` values that your own processors will stamp onto spans and metrics — a spoofing and cardinality-bomb vector at once. The perimeter rule is: strip inbound baggage at the edge, or allowlist a small set of keys and validate their values, then set the trustworthy versions yourself from the authenticated principal. ## Governance Because baggage is global and cheap to add, it accretes. Treat the key set as a governed schema: a short documented list with an owner per key, a size budget, a perimeter allowlist enforced in the gateway, and an alert on total baggage header size at ingress. Removing a key later is hard precisely because you cannot see who depends on it — another argument for keeping behaviour out of it.

  • An engineer adds the end user's email to baggage so every service can log it. What is your objection?
    It is personal data broadcast to every hop, including any third-party endpoint the outbound client injects into, and it lands in every log and span that copies baggage — a privacy and retention problem that is very hard to unwind once distributed. It is also high cardinality if it ever reaches metric labels. If downstream services genuinely need user identity, pass an opaque user id in the documented request contract and resolve it locally, or attach it as a span attribute where it is actually needed.
  • Why does putting a value in baggage not make it searchable in your tracing backend?
    Baggage lives in the Context and is serialised to the `baggage` header; spans are exported with their own attributes only. Nothing copies one into the other by default. To make baggage searchable you enable a span processor that reads specific allowlisted baggage keys and sets them as attributes on spans as they start — keeping the allowlist explicit so high-cardinality keys do not bloat the trace store or leak into metrics.

Span attributes are notes written on one page of the file; baggage is a sticky note attached to the folder that every desk it visits can read and act on — useful, but everyone pays to carry it and anyone upstream can write on it.

saying these in an interview costs you the question

  • "Baggage automatically shows up as span attributes" — it does not; a processor must copy it.
  • Using baggage to carry authorisation or tenancy decisions that code trusts, turning attacker-controlled headers into policy input.
  • Putting user identifiers, emails or tokens in baggage and forgetting it is transmitted to every downstream, potentially external, service.
  • Copying all baggage keys into metric labels, creating a cardinality explosion.
  • Assuming baggage is unlimited — the header has size and entry caps and entries can be dropped silently.

context