Which OpenTelemetry GenAI span attributes carry the model name and token counts?
answer
- Standard names, so any backend can chart it
- Two model attributes, not one
- Tokens split by direction
- input_tokens and output_tokens, not prompt/completion
- gen_ai.request.model vs gen_ai.response.model
basics
~10 sgen_ai.request.model records the model you asked for, gen_ai.response.model the one that answered, and gen_ai.usage.input_tokens plus gen_ai.usage.output_tokens the token counts. Standard names let any backend chart cost without reading your code.
solid answer
~40 sThe GenAI semantic conventions fix a small vocabulary that instrumentation writes onto every model-call span. Identity comes from `gen_ai.provider.name` (which provider API was called) and `gen_ai.operation.name` (chat, embeddings, and so on). The model appears twice on purpose: `gen_ai.request.model` is what your code asked for, `gen_ai.response.model` is what the provider says actually served the call — often a dated snapshot id. Sampling parameters land as `gen_ai.request.temperature`, `gen_ai.request.max_tokens`, `gen_ai.request.top_p`. Usage is two integers, `gen_ai.usage.input_tokens` and `gen_ai.usage.output_tokens` (earlier drafts of the convention spelled these prompt_tokens and completion_tokens). Nothing here is prompt text and nothing here is money: message content is a separate opt-in concern, and cost is computed by the backend from tokens times its own price table. Because the names are standard, a dashboard you did not write can group spend by model.
code
json · 13 lines{
"name": "chat gpt-4o",
"attributes": {
"gen_ai.provider.name": "openai",
"gen_ai.operation.name": "chat",
"gen_ai.request.model": "gpt-4o",
"gen_ai.response.model": "gpt-4o-2024-08-06",
"gen_ai.request.temperature": 0.2,
"gen_ai.request.max_tokens": 512,
"gen_ai.usage.input_tokens": 431,
"gen_ai.usage.output_tokens": 88
}
}go deeper
Be able to name the four everyday attributes without hesitation: gen_ai.request.model, gen_ai.response.model, gen_ai.usage.input_tokens, gen_ai.usage.output_tokens, and say that standard names are what make a shared dashboard possible.
Explain why the model appears twice, why cost is absent, and that the token attributes were renamed from prompt/completion to input/output — then show how a cost view is assembled from tokens grouped by model.
Show the operational instincts: missing usage attributes on streamed calls, a stale price table behind a wrong cost chart, and grouping quality metrics by response model so a silent alias rollover shows up as a new series.
Own the argument for standardising on the upstream vocabulary across every service and language rather than per-team custom keys, and for keeping money out of the span so pricing stays a query-time concern you can restate retroactively.
## Why a fixed vocabulary exists Every provider SDK returns its own shape: one calls the field `usage.prompt_tokens`, another `usage.input_tokens`, a third nests it under a different key entirely. If each tracing integration invented its own span attribute names, every dashboard, alert and cost report would have to be rewritten per provider and per library. The OpenTelemetry GenAI semantic conventions solve exactly that: they define a small set of attribute keys that any instrumentation — OpenLLMetry included — writes onto the span that wraps a model call. The payoff is that a backend which has never seen your source code can still answer "how many output tokens did each model burn yesterday". ## The identity attributes - `gen_ai.provider.name` — which provider API was invoked, with values such as `openai`, `anthropic`, `aws.bedrock`. This attribute replaced the older `gen_ai.system`, which is now deprecated. - `gen_ai.operation.name` — the kind of call: chat completion, text completion, embeddings. This is what lets you exclude embedding spans from a chat-latency chart. ## The two model attributes `gen_ai.request.model` is the identifier your code passed in — for example `gpt-4o`. `gen_ai.response.model` is what the provider reports having used, which is frequently a pinned, dated snapshot behind an alias. Recording both is deliberate: an alias silently rolling to a new snapshot is one of the most common explanations for "the prompt started behaving differently and nothing changed on our side". Group your quality metrics by the response model, not the request model, if you want that shift to be visible. ## Request parameters `gen_ai.request.temperature`, `gen_ai.request.max_tokens` and `gen_ai.request.top_p` record the sampling knobs. They are cheap, low-cardinality numbers, and they let you correlate an odd or slow response with the settings that produced it — useful when a config change ships through a feature flag rather than through code. ## Usage: the two attributes everything financial is built on `gen_ai.usage.input_tokens` and `gen_ai.usage.output_tokens` are plain integers. Almost every cost view in every LLM observability product is these two numbers multiplied by a per-model price and grouped by model. Note the naming history: earlier drafts of the convention used `gen_ai.usage.prompt_tokens` and `gen_ai.usage.completion_tokens`, mirroring the OpenAI response shape. The current names are input/output, which generalise across providers that do not use prompt/completion language. Dashboards built against the old spelling silently return nothing rather than erroring, so a chart that suddenly reads zero after a dependency bump is worth checking against the current names. ## What is deliberately absent Three things are not in this core set: 1. **Cost.** No attribute carries dollars. Prices change and vary by contract, so the convention records the physical quantity (tokens) and leaves money to the backend's price table. That also means a stale price table, not a bad span, is a common cause of a wrong cost chart. 2. **Prompt and completion text.** Message content is treated as a separate, sensitive, opt-in concern rather than part of the always-on attribute set — it is the payload, not the metadata. 3. **Provider-specific extras.** Anything one vendor exposes and others do not tends to live outside the standard names, so treat such keys as non-portable. ## Where these come from in practice With OpenLLMetry you rarely write these attributes yourself. Its auto-instrumentation wraps the provider client and sets them on the span it creates around the call, reading the values out of the request arguments and the response object. You read them in queries: sum output tokens by `gen_ai.request.model`, chart p95 duration filtered by `gen_ai.provider.name`, alert when a model's error rate spikes. If you hand-roll a span around a call the instrumentation does not cover, setting the same keys yourself is what makes it show up in the same charts. ## Common trip-ups Confusing request and response model when grouping metrics; assuming token attributes are always present (a streamed response may not carry usage unless the client asked for it, in which case the attribute is simply missing rather than zero); and expecting `gen_ai.usage.input_tokens` to break out cached versus fresh input tokens, which it does not — it is one number, and billing may treat parts of it differently.
- Why does the convention record tokens rather than cost directly on the span?Because price is not a property of the call. Rates change, differ per contract and per region, and are renegotiated after the trace was written. Tokens are the physical, immutable quantity; the backend multiplies them by its own price table at query time. That also means a wrong cost chart is usually a stale price table rather than bad instrumentation.
- A cost dashboard built a year ago suddenly shows zero tokens after a dependency upgrade. What is the first thing you check?Whether the query still uses the old attribute spelling. The convention moved from gen_ai.usage.prompt_tokens and completion_tokens to gen_ai.usage.input_tokens and output_tokens. Query languages return an empty series for an attribute that does not exist rather than raising, so a rename looks exactly like "no traffic". Compare a raw span from before and after the upgrade.
- Which model attribute should quality metrics be grouped by, and why?gen_ai.response.model. The request model is often an alias that the provider resolves to a dated snapshot; when that alias rolls forward, output quality can shift while your code and your request attribute stay identical. Grouping by the response model makes the rollover visible as a new series on the chart instead of an unexplained drift in the old one.
saying these in an interview costs you the question
- Thinks the span carries a cost or dollar attribute
- Says request.model and response.model always hold the same value
- Uses gen_ai.usage.prompt_tokens as the current name
- Assumes prompt text is part of the standard attribute set
- Believes token attributes are always present on every span