skip to content

Why can OpenLLMetry's flattened gen_ai.prompt attributes overwhelm a tracing backend?

level: seniorimportance: should knowfreq 40%

answer

  1. One message becomes several indexed attributes
  2. Two limits, count and size
  3. Failures are silent, not errors
  4. Default cap of 128 span attributes
  5. Payload bytes dominate the tracing bill

basics

~10 s

Each chat message becomes its own indexed attributes — gen_ai.prompt.0.role, gen_ai.prompt.0.content and so on. Long conversations and retrieved context produce many large string attributes, which hit span attribute limits and dominate ingestion cost.

solid answer

~50 s

OpenLLMetry does not store a chat array as one structured field; it flattens it into indexed attributes — `gen_ai.prompt.0.role`, `gen_ai.prompt.0.content`, `gen_ai.prompt.1.role`, and the mirror `gen_ai.completion.0.*` keys. Two consequences follow. First, **count**: a long conversation or a tool-calling loop yields two or more attributes per message, and OpenTelemetry SDKs cap span attributes (128 by default via `OTEL_SPAN_ATTRIBUTE_COUNT_LIMIT`), so attributes past the cap are dropped — usually the tail of the conversation, silently. Second, **size**: a RAG prompt carrying retrieved chunks can be tens of kilobytes in a single string attribute, so payloads dominate span bytes, drive ingestion cost, and get truncated by whatever value-length limit the SDK or backend applies. Mitigations are to disable or redact content, cap what you attach, put payloads in a store built for blobs and reference them from the span, and sample which requests carry full content.

code

json · 9 lines
json
{
  "gen_ai.request.model": "gpt-4o",
  "gen_ai.prompt.0.role": "system",
  "gen_ai.prompt.0.content": "You are a support agent...",
  "gen_ai.prompt.1.role": "user",
  "gen_ai.prompt.1.content": "<18 KB of retrieved context + question>",
  "gen_ai.completion.0.role": "assistant",
  "gen_ai.completion.0.content": "Based on the policy..."
}

go deeper

for a junior

Know that message text is stored as numbered attributes like gen_ai.prompt.0.content rather than as one list, and that big prompts therefore make big spans.

for a middle

Explain both limits — attribute count and attribute value length — and that exceeding them drops or truncates data without raising an error, so a conversation can appear shorter than it was.

for a senior

Show the production handling: chart dropped attributes, set SDK and backend limits deliberately, sample which requests carry full payloads, and reference payloads stored elsewhere instead of embedding them on every span in an agent loop.

for a principal

Own the economics and the data-classification consequences — payload bytes set the tracing bill and the sensitivity tier of the whole trace store, so decide per service whether traces carry evidence or merely point at it.

## The shape on the wire A chat request is a list of message objects. Span attributes, by contrast, are a flat map of string keys to primitive values. OpenLLMetry bridges that gap by flattening: message `i` of the prompt becomes `gen_ai.prompt.{i}.role` and `gen_ai.prompt.{i}.content`, and the response messages become `gen_ai.completion.{i}.role` and `gen_ai.completion.{i}.content`. It is a reasonable encoding — any backend can display it without understanding chat — but it turns one logical payload into many attributes, and that is where the operational problems come from. ## Problem one: attribute count OpenTelemetry SDKs enforce a limit on how many attributes a span may carry; the specified default is 128, configurable through `OTEL_SPAN_ATTRIBUTE_COUNT_LIMIT`. At two attributes per message you exhaust that in the sixties of messages — sooner once tool calls, function arguments and your own business attributes share the same span. What makes this nasty is the failure mode: attributes beyond the limit are **dropped**, and the span still exports successfully. The trace looks fine. A count of dropped attributes may be recorded, but nobody reads it. You discover the problem when a debugging session shows a conversation that ends mid-way and you assume the application truncated it. Long agent loops are the usual trigger, because each iteration appends to the message list and the last call in the loop carries the whole history. ## Problem two: attribute size The other axis is bytes per attribute. A retrieval-augmented prompt inlines the retrieved chunks into a message, so a single `gen_ai.prompt.0.content` value of 20-80 KB is unremarkable. Multiply by request volume and content is typically the majority of your LLM tracing bill, since these products price on ingested data. Along the way you meet truncation: OpenTelemetry has a configurable value-length limit (`OTEL_ATTRIBUTE_VALUE_LENGTH_LIMIT`), and backends impose their own caps on attribute size and total span size, above which they truncate or reject. A rejected span is worse than a truncated one — you lose the metadata too. There is also a transport dimension. Big attributes make export batches large, which raises memory pressure in the exporter queue and increases the chance of dropped batches under load, exactly when tracing matters most. ## Problem three: portability The flattened indexed keys are OpenLLMetry's encoding. A backend that has never seen them will render `gen_ai.prompt.3.content` as an ordinary string attribute rather than as a conversation turn, and any query you write against the numbered keys is coupled to that shape. Treat conversation rendering as a feature of the backend you chose, not as something guaranteed by the standard attribute vocabulary — and be aware that the upstream conventions have continued to evolve in how message content is carried, so encodings can change under you across instrumentation versions. ## Problem four: what is in those strings Content attributes carry raw user text, which means the data-classification of your whole trace store is set by the most sensitive prompt you record. Attribute-level redaction is harder than it sounds because the payload is free text, not fields — you cannot drop a column, you have to scan and rewrite. This is why the blunt on/off content toggle exists and why many teams use it. ## What to do about it 1. **Decide whether you need full content at all.** For many services, the model, token counts, latency and an application-side request id are enough, with payloads living in a system that already has retention and access controls. The span links to the evidence rather than being the evidence. 2. **Cap what you attach.** If you are constructing spans yourself, record the last N messages plus lengths, or a summary plus a reference, instead of the entire history on every call in a loop. 3. **Sample content, not just traces.** Full payloads on a small, deliberately chosen percentage of requests plus metadata on all of them usually keeps the debugging value and removes most of the bytes. 4. **Know your limits before you hit them.** Check the SDK's attribute count and value-length settings, and the backend's per-span and per-attribute caps, and set them consciously rather than discovering them through mysteriously short conversations. 5. **Watch for silent drops.** If your tracing backend surfaces dropped-attribute counts, chart them; a rising line is the early warning for exactly this failure. 6. **Redact in a span processor** when policy allows content in principle but not verbatim — hash or mask before export. ## The interview answer in one breath Flattening turns one conversation into 2N indexed attributes, so you meet a count limit and a size limit at the same time; both fail quietly by dropping or truncating rather than erroring; payload bytes dominate cost; and the fixes are to reference rather than embed, cap and sample what you attach, and set your limits deliberately.

  • An agent loop's final span shows only the first part of the conversation. What is your first hypothesis?
    That the span hit the SDK's attribute count limit — 128 by default — and the later indexed prompt attributes were dropped rather than exported. Each message contributes at least a role and a content attribute, so a long loop exhausts the budget fast. Check the dropped-attribute count if the backend exposes it, raise OTEL_SPAN_ATTRIBUTE_COUNT_LIMIT deliberately, or stop attaching the full history on every iteration.
  • How would you keep the debugging value of prompts without paying for full payloads on every request?
    Sample the content dimension separately from the trace dimension: keep metadata on 100% of calls and full payloads on a small, deliberate slice — plus an override that captures everything for a specific user or session while you reproduce a bug. Alternatively store payloads once in a blob or application store and put only its id on the span, so the trace points at the evidence.
  • Why is a rejected span worse than a truncated one?
    Truncation costs you the tail of a payload but you keep the span, so cost, latency and error dashboards are unaffected. A rejection at ingestion removes the whole span, taking model, token counts and status with it — the request disappears from your metrics entirely, which both hides the incident and quietly biases every aggregate toward small requests.
  • Are the indexed gen_ai.prompt.N keys something every tracing backend understands?
    No. Flattening into numbered keys is the instrumentation's encoding, and a backend without specific support renders them as ordinary string attributes rather than a conversation view. Queries written against the numbered keys are coupled to that encoding, which has continued to evolve upstream — so treat conversation rendering as a property of the backend you chose, not a guarantee of the convention.

saying these in an interview costs you the question

  • Assumes attributes over the limit cause a visible error
  • Thinks the chat array is stored as one structured field
  • Ignores that payload size drives most of the tracing bill
  • Believes every backend renders the indexed keys as a conversation
  • Adds the entire message history to every span in an agent loop

context