skip to content

Why follow the OpenTelemetry GenAI semantic conventions when tracing LLM calls?

level: seniorimportance: should knowfreq 37%

answer

  1. names as an interoperability contract
  2. one query across providers and languages
  3. pre-stable, so names still move
  4. dual emission during a rename
  5. own namespace for your own attributes

basics

~20 s

They fix the attribute names on a model-call span — provider, requested model, token usage, finish reason — so any backend can render and query your LLM spans without per-vendor mapping. As of mid-2026 they are still pre-stable, so pin a version.

solid answer

~50 s

Standard names are what make LLM traces portable. The GenAI semantic conventions define a common vocabulary for a model-call span: `gen_ai.operation.name`, the provider in `gen_ai.system`, `gen_ai.request.model` versus `gen_ai.response.model`, request settings such as `gen_ai.request.temperature`, `gen_ai.response.finish_reasons`, and token counts under `gen_ai.usage.*`. Adopt them and one query works across services written in different languages and calling different providers, off-the-shelf dashboards light up, and you can change backend without re-instrumenting. The honest caveat as of mid-2026 is that `gen_ai.*` is still **pre-stable**: the attributes moved into a dedicated semantic-conventions repository at v1.42.0 in June 2026, with no 1.0 and no announced stabilisation date, and names have already churned. So pin a convention version explicitly, use the dual-emission opt-in (`OTEL_SEMCONV_STABILITY_OPT_IN`) while a rename rolls out, normalise old names in the collector rather than editing every service, and keep your own attributes in your own namespace so a future rename cannot collide with them.

code

yaml · 15 lines
yaml
span:
  name: chat some-chat-model
  kind: CLIENT
  attributes:
    gen_ai.operation.name: chat
    gen_ai.system: some_provider
    gen_ai.request.model: some-chat-model
    gen_ai.response.model: some-chat-model-2026-05-01
    gen_ai.request.temperature: 0.2
    gen_ai.request.max_tokens: 1024
    gen_ai.response.finish_reasons: [stop]
    gen_ai.usage.input_tokens: 3120
    gen_ai.usage.output_tokens: 214
    myapp.feature: trip_planner
    myapp.prompt.version: "2026.06.1"

go deeper

for a junior

Know that standard attribute names exist for LLM spans and that using them is preferred over inventing your own, so dashboards and tools understand your traces.

for a middle

Be able to name the core model-call attributes — operation, provider, requested versus responding model, finish reason, token usage — and explain what portability they buy.

for a senior

Demonstrate migration discipline: pin the convention version, dual-emit through a rename, normalise old names in the collector, and keep custom attributes namespaced.

for a principal

Own the decision to depend on a pre-stable standard: what you standardise across teams, what you allow locally, and how you absorb renames without a fleet-wide rewrite.

## The problem the conventions solve Every team that instruments LLM calls invents attribute names on day one: `model`, `model_name`, `llm.model`, `openai.model`. Individually all fine. Collectively they mean that a dashboard built for one service is useless for the next, a query for 'all calls to the expensive model across the estate' has to enumerate spellings, and switching observability backends means touching every service. Semantic conventions are the fix: a published vocabulary, so that instrumentation written by different people — and by library authors who have never seen your system — produces spans that mean the same thing. ## What the GenAI conventions cover The part you use daily is the model-call span. Conventionally it is named after the operation and the model, and carries: - `gen_ai.operation.name` — what kind of call this is (for example chat or embeddings). - `gen_ai.system` — which provider or system served it. - `gen_ai.request.model` — the model you asked for, and `gen_ai.response.model` — the model that actually answered. These differ more often than people expect, when a provider serves an alias or a pinned snapshot. - Request settings such as `gen_ai.request.temperature` and `gen_ai.request.max_tokens`. - `gen_ai.response.finish_reasons` — why generation stopped, which is the single most useful attribute for diagnosing truncated output. - `gen_ai.usage.*` — input and output token counters. There are also conventions for tool and agent operations, and a separate, more contested area: where the prompt and completion *content* goes. Content has moved between span attributes, span events and log records across revisions, which is precisely the kind of churn you have to plan for. ## The stability picture as of mid-2026 This matters more than the attribute list. The `gen_ai.*` conventions are **not stable**. In June 2026 they were split out of the main semantic-conventions repository into a dedicated one at v1.42.0, and there is still no 1.0 and no committed stabilisation date. Concretely: an attribute you emit today may be renamed in a release six months out, and instrumentation libraries will adopt the rename at different times, so a single trace can contain both spellings. Mature teams handle this with four habits: 1. **Pin and record the convention version.** Treat it as a dependency, not a background fact, and put the version on the span or resource so you can tell later which spelling to expect. 2. **Use dual emission during a rename.** The ecosystem's mechanism is the `OTEL_SEMCONV_STABILITY_OPT_IN` environment variable, which lets instrumentation emit old names, new names, or both for a migration window. Both means dashboards keep working while you cut over. 3. **Normalise in the pipeline, not in the services.** A collector-side transform that maps yesterday's names onto today's is one change instead of forty, and it can be removed once every service has moved. 4. **Namespace your own attributes.** Anything you invent — feature name, tenant, prompt version, verdict — goes under a prefix you own. Never squat inside `gen_ai.*`, because a future convention release may define your name with different semantics. ## What you still have to decide yourself The conventions are a naming contract, not an instrumentation design. They do not tell you how to nest agent, tool and retrieval spans, what your correlation ids should be, or how much payload to capture. They also cannot make a badly shaped trace useful: perfectly named attributes on a single span per turn are still a single span per turn. Treat conventions as the interoperability layer beneath your own span taxonomy. ## When to deviate Deviating is sometimes right, and should be a decision rather than an accident. If a convention has no attribute for something you genuinely need, add it in your own namespace and move on. If your provider exposes a concept the conventions have not modelled, do the same. What you should not do is rename convention-covered concepts to something you prefer, because you are then paying the cost of a standard — learning it — while getting none of the benefit. ## The interview point The strong answer has two halves. First, why standardising names pays: portability across tools, languages and providers, and dashboards that outlive any one vendor choice. Second, an honest account of the state of these particular conventions in 2026 — pre-stable, actively moving, with a documented dual-emission escape hatch — and the migration discipline that follows from it. Candidates who only recite attribute names miss the half that actually bites in production.

  • A trace contains both the old and the new spelling of a GenAI attribute. How did that happen and what do you do?
    Different instrumentation in the same process adopted the convention rename at different versions, or a dual-emission window is deliberately open. Short term, make dashboards tolerant by querying both spellings, or normalise to one name in the collector so downstream sees a single vocabulary. Then align library versions and close the window. The mistake is editing dashboards service by service while the rename is still in flight.
  • Why record both the requested model and the model that responded?
    Because they differ. Providers serve aliases and pinned snapshots, so an alias you requested can resolve to a version that changed under you — one of the most common causes of 'it got worse overnight'. Recording only the request hides the switch entirely; recording both lets you group behaviour by what actually served the call and spot the day the resolution changed.
  • If the conventions do not define an attribute you need, what do you do?
    Add it under a prefix you own, document it, and keep it out of the gen_ai.* namespace. Squatting inside a pre-stable namespace risks a future release defining the same name with different semantics, which silently corrupts historical queries. Your own namespace also makes it obvious later which attributes are yours to migrate if the conventions eventually cover the concept.

saying these in an interview costs you the question

  • Assuming the GenAI conventions are stable and frozen
  • Inventing your own attribute names for concepts the conventions already cover
  • Putting custom attributes inside the gen_ai.* namespace
  • Renaming attributes service by service instead of normalising in the pipeline
  • Recording only the requested model and never the model that answered

context