Describe the inject/extract carrier model OpenTelemetry uses to move context across a process boundary, and how you would apply it to a transport that is not HTTP, such as a message broker's record headers.
answer
- Propagator = format, setter/getter = carrier
- fields() → clear stale traceparent on reused objects
- Headers not payload body
- Batch consume → links, not one forced parent
- PRODUCER / CONSUMER span kinds
basics
~20 sA text-map propagator exposes inject(context, carrier, setter) and extract(context, carrier, getter). The setter and getter abstract the carrier, so the same propagator writes into HTTP headers, gRPC metadata or broker record headers. For messaging you inject when producing and extract when consuming.
solid answer
~50 sOpenTelemetry separates *what* is serialised (the propagator: W3C trace context, B3, Baggage) from *where* it is written (the carrier), joined by a tiny setter/getter interface. `inject` asks the propagator to serialise the current context and calls `setter.set(carrier, key, value)`; `extract` calls `getter.get(carrier, key)` and returns a new context containing the remote span context and baggage. Because the carrier is opaque, the same propagator serves any key/value-capable transport — HTTP headers, gRPC metadata, Kafka record headers, an AMQP property map, or a field you serialise yourself. A composite propagator lets you emit and accept several formats during a migration. Two messaging-specific points: inject at produce time so the context is captured with the message rather than at flush, and on the consumer side a batch has many producers, so you typically start one consumer span and attach each message's extracted context as a *link* rather than forcing a single parent.
code
text · 10 linesHTTP: inject(ctx, request, (req,k,v) -> req.setHeader(k,v))
Broker: inject(ctx, record, (rec,k,v) -> rec.headers.add(k, utf8(v)))
consume(batch):
span = start("process", parent = local_ctx, kind = CONSUMER)
for msg in batch:
remote = extract(root, msg.headers, headerGetter)
span.addLink(remote.spanContext) # many producers, one batch
... process ...
span.end()go deeper
Name the two operations and say the setter/getter make the propagator work over any key/value transport, with headers as the place context goes.
Add composite propagators, produce-time injection, and the distinction between parent and link for messaging.
Cover getter pitfalls (case, duplicates, binary values), stale-field clearing, and the operational consequences of traces that span queues and retries.
Set the estate-wide policy: one canonical format with a bounded dual-emit migration window, conventions for asynchronous boundaries, and where to break a trace rather than let it grow unbounded.
## The abstraction Cross-process propagation needs two orthogonal decisions: the wire *format* of the context, and the *place* it is written. OpenTelemetry keeps them separate. A `TextMapPropagator` owns the format. It has three operations: ``` inject(context, carrier, setter) # serialise current context into the carrier extract(context, carrier, getter) -> context # parse a carrier into a new context fields() -> [names it writes] # so callers can clear stale values ``` The carrier is whatever object holds the key/value pairs — a header map, a metadata object, a list of record headers. The setter and getter are small adapters supplied by the instrumentation: ``` setter.set(carrier, key, value) getter.get(carrier, key) -> value or null getter.keys(carrier) -> all keys ``` That is the whole contract. It is why the same W3C propagator works unchanged over HTTP, gRPC, and a broker: only the four-line adapter differs. `fields()` is easy to overlook and matters in practice: when you reuse a message or request object, you should clear the propagator's fields before injecting, otherwise a stale `traceparent` from a previous use can survive and parent the new operation to the wrong span. ## Composite propagators A composite (or multi) propagator wraps several propagators. On inject it writes all of their formats; on extract it tries each in order and takes the first valid result. Two standard uses: combining trace context with baggage (almost always what you want), and interoperating during a migration — accept both W3C and B3 while some services still emit only the old format, and emit both until every consumer understands the new one. The cost is header bytes on every request, so it is a migration state, not a destination. ## Getter and setter details that bite - **Case-insensitivity.** HTTP header names are case-insensitive; a getter over a plain case-sensitive map will miss `TraceParent` sent by a non-conforming client. Getters over HTTP carriers must fold case. - **Multiple values.** Headers can repeat. Trace context requires a single `traceparent`; the spec says a duplicated header makes the context invalid, so a getter that silently returns the first value can produce a context the sender never intended. - **Binary carriers.** Broker headers are often byte arrays rather than strings; the adapter must decode with a fixed charset (UTF-8) on both sides. - **Immutable carriers.** Some client libraries expose request objects that are already frozen at the point instrumentation runs, which is why injection has to happen at the right lifecycle stage rather than wherever it is convenient. ## Applying it to a message broker Producing: 1. Start a producer span while the current context is the one that caused the send. 2. Inject the context into the message's *headers*, using a setter that appends to the record's header list. Headers, not the payload body: injecting into the body changes the message schema, forces every consumer to understand tracing, and defeats brokers or bridges that only forward headers. 3. Send. Do this at produce time, not at flush or batch-completion time — by then the ambient context is a different request's. Consuming: 1. For each message, extract a context from its headers with the same propagator. 2. If the consumer handles one message per unit of work, use the extracted context as the parent: the trace continues naturally across the queue. 3. If the consumer processes a *batch*, the messages usually come from many different traces. Forcing one parent would be a lie. The convention is to start one processing span whose parent is the local consumer context, and attach each message's extracted span context as a **link** — a typed reference that says "related to, but not a child of". Backends render links as navigable cross-trace edges. Span kinds should be `PRODUCER` and `CONSUMER` rather than client/server, which is what lets a backend show the asynchronous edge correctly. ## Consequences of propagating through a queue A trace that crosses a queue can span minutes or hours: the producer span ends immediately, the consumer span begins when the message is picked up. Some backends assemble such traces poorly, or have already flushed and indexed the earlier part. Deep retry chains and dead-letter requeues can also extend one trace id indefinitely, or fan a single trace into thousands of spans. Two mitigations: prefer links over parenting when the asynchronous gap is large or the fan-in is wide, and consider starting a fresh trace with a link back to the origin when a message is replayed from a dead-letter queue much later. ## Interview summary Say that the propagator owns the format and the setter/getter pair owns the carrier, which is why the model is transport-agnostic; then walk the produce/consume path, call out headers-not-body, and explain links for batch consumption.
- Why attach links instead of parenting when consuming a batch of messages?A parent asserts that one operation caused this one, and a span can have only one parent. A batch typically contains messages from many unrelated traces, so any single parent choice would be wrong for the rest and would splice unrelated traces together. Links express "related to" without implying causation or merging traces, so each origin trace stays intact and the backend can still offer navigation from the batch span to every contributing message.
- A team injects the trace context into the message payload instead of the headers. What problems follow?The message schema now carries transport concerns, so every producer and consumer must be updated together and non-instrumented consumers may fail to deserialise. Brokers, bridges and routing rules that inspect only headers cannot see or forward the context, and schema-registry validation may reject the extra field. Headers are the designed extension point precisely because they are metadata that intermediaries can carry without understanding the body.
saying these in an interview costs you the question
- Thinking a propagator is HTTP-specific rather than carrier-agnostic through setter/getter.
- Injecting context at flush or batch-send time, capturing whatever request happens to be current then.
- Forcing a single parent onto a batch of messages from many different traces.
- Putting the context in the message body instead of headers.
- Forgetting case-insensitive header lookup, so context from some clients is silently missed.
- Reusing a request or record object without clearing the propagator's fields, inheriting a stale traceparent.