As a platform owner, what governance and failure-mode concerns would you set for Baggage propagation across a large service fleet?
answer
- Baggage is transitive — every hop + every log line
- Allow-list + owners; config baseline is the control point
- No secrets/PII; strip at trust boundaries; don't trust inbound for authz
- Standardize W3C + field names or it silently breaks
- Budget header size on high-QPS / messaging fan-out
basics
~20 sRestrict baggage to a small, approved allow-list of small non-sensitive fields; forbid secrets/PII since baggage is forwarded everywhere and logged. Standardize field names and propagation format fleet-wide, strip baggage at public trust boundaries, and watch header-size/overhead on high-volume paths.
solid answer
~50 sBaggage is powerful but dangerous at scale because it's forwarded to *every* downstream service and typically copied into logs. My governance: (1) An **approved allow-list** of baggage fields (tenantId, requestId, cohort) with owners; nothing propagates unless declared, so I control it via a shared config baseline (`remote-fields`/`correlation.fields`). (2) **No secrets or PII** — baggage leaks widely and lands in MDC/logs and third-party spans; treat it as public. (3) **Strip baggage at trust boundaries** — inbound from the internet and outbound to third parties must not carry internal baggage. (4) **Standardize the propagation format** (W3C) and **field names/casing** across the fleet, or propagation silently breaks. (5) **Budget the overhead** — every field is bytes on every hop and every message header; cap count and size, especially on high-QPS and messaging paths. (6) Monitor for context-loss (async, un-instrumented clients) and cardinality of correlated fields in logs.
go deeper
Know baggage should stay small and shouldn't hold secrets.
Explain that baggage propagates everywhere and is logged, so limit fields and avoid PII.
Handle spoofing, strip at boundaries, and manage async context loss and header overhead.
Set fleet-wide policy: approved allow-list with owners, standardized format/names, edge stripping, cost budgets, and monitoring for silent propagation breakage.
**Why baggage needs governance.** Unlike a header you set on one call, baggage is **transitively propagated**: once in scope it rides on *every* downstream request and message for the whole trace, and (if correlated) is written into the **MDC/logs** of every service. That reach makes it excellent for correlation and hazardous for security, cost, and reliability. A platform owner treats baggage as a shared, blast-radius-wide resource. **1) Allow-list & ownership.** Because Spring only propagates fields declared in `management.tracing.baggage.remote-fields` (and correlates those in `correlation.fields`), configuration *is* the control point. Ship a **shared config baseline** (a starter/BOM or config server defaults) defining the approved set — e.g. `tenantId`, `requestId`, `experimentCohort` — each with a documented owner and purpose. Teams can't silently add fields that then spread fleet-wide. Review additions like you'd review a new public API. **2) Security — treat baggage as public.** - **No secrets** (tokens, keys) and **no PII** you wouldn't broadcast: baggage flows to every internal service, is copied into logs (a wide exfiltration surface), and may reach third-party tracing backends. - **Strip at trust boundaries.** Inbound baggage from untrusted callers can be **spoofed** (a client could set `tenantId` to another tenant) — never trust inbound baggage for authorization; derive authz from authenticated identity, not baggage. Strip or overwrite internal baggage at the edge, and strip it again on egress to external partners so internal context doesn't leak. - **Injection/log-forging.** Values reach the MDC; sanitize to avoid log-injection (newlines/control chars) and don't let user-controlled baggage pollute log parsing. **3) Reliability — silent failure modes.** - **Format & naming drift.** Mismatched propagation format (W3C vs B3) or inconsistent field **names/casing** across services causes *silent* context/baggage loss — no error, just broken correlation. Standardize both fleet-wide (prefer **W3C**). - **Context loss in async code.** Raw threads, un-instrumented executors, reactive boundaries, and manual message handling drop baggage/trace context. Mandate context-propagating executors / Reactor context propagation / `ContextSnapshot`, and test it. - **Scope leaks.** `BaggageInScope` must be closed (try-with-resources) or values bleed across requests on pooled threads. **4) Cost & performance.** Each remote field is extra bytes on **every** hop's headers and **every** message header. On high-QPS HTTP and high-volume topics this adds real bandwidth and can hit **header-size limits** (proxies/brokers reject oversized headers). Keep the field count small and values short; be especially conservative on messaging fan-out. Correlated fields also raise log volume and can add **high-cardinality** dimensions if you index them. **5) Observability of the mechanism itself.** Monitor for: traces that split (propagation broken), consumers starting as roots (missing extraction), and unexpected baggage fields appearing (drift). Consider synthetic checks that assert a known baggage field survives end-to-end. **When baggage is the right tool.** A *few* stable, non-sensitive correlation keys that many services need to log/branch on. It is **not** a substitute for passing request data explicitly, a config channel, or an authz mechanism. **Summary posture.** Small approved allow-list; public, non-sensitive values only; standardized format + names; stripped at edges; overhead-budgeted; and continuously monitored for silent breakage.
- A team wants to use an inbound `tenantId` baggage field to decide which tenant's data to return. What's your concern?Baggage is client-settable and spoofable, so using it for authorization lets a caller impersonate another tenant. Authorization must come from authenticated identity (the validated principal/JWT claims), not from baggage. Baggage is fine for correlation/logging, never for access-control decisions, and inbound baggage should be validated or stripped at the edge.
- How would you prevent internal baggage from leaking to a third-party API you call?Strip or filter baggage on egress at the trust boundary — e.g. a dedicated outbound client that doesn't propagate internal remote-fields, or a gateway/mesh policy that removes baggage headers on external routes. The same applies inbound from the public internet: don't accept or trust externally supplied baggage.
saying these in an interview costs you the question
- Using baggage for authorization decisions (it's spoofable)
- Putting PII/secrets in baggage because 'it's just internal'
- Assuming every team can add baggage fields freely without fleet-wide impact
- Ignoring header-size/bandwidth cost on high-volume and messaging paths
- Not stripping baggage at internet/third-party boundaries