skip to content

OpenLLMetry

OpenTelemetry semantic conventions applied to LLM and agent calls, so model spans land in the same backend as the rest of your telemetry. It is the answer when an interviewer asks how LLM observability fits an existing tracing stack instead of replacing it.

on this pageshow

questions

12

What does Traceloop.init() do so existing OpenAI calls emit spans in OpenLLMetry?

level: juniorimportance: must knowfreq 74%

answer

  1. one call at startup, no call-site edits
  2. the SDK patches libraries you already import
  3. provider SDKs, frameworks, vector stores
  4. patching applies to library methods, not your instances
  5. calls before init are never traced

basics

~20 s

Traceloop.init() starts OpenLLMetry's tracing pipeline and monkey-patches the LLM client libraries already installed in the process — OpenAI, Anthropic, LangChain, LlamaIndex, vector-store clients — so their calls emit spans. Your existing call sites stay unchanged.

solid answer

~40 s

`Traceloop.init(app_name="my-service")` is the single setup call. It builds the tracing pipeline the SDK exports through, then runs the OpenTelemetry instrumentation packages that ship with `traceloop-sdk`. Those instrumentors patch methods on the client libraries they find installed — the OpenAI and Anthropic SDKs, LangChain and LlamaIndex, and vector-store clients — so a normal `client.chat.completions.create(...)` produces a span with model, parameters and token usage without you touching the call site. The patching is applied to the library's own methods, so it does not matter whether you constructed your client object before or after `init`; what matters is that calls happen *after* `init` has run. Practically that means calling it once at application startup, before serving traffic. `app_name` becomes the service identity the spans are grouped under in your backend.

code

python · 11 lines
python
from openai import OpenAI
from traceloop.sdk import Traceloop

Traceloop.init(app_name="recipe-bot")

client = OpenAI()
response = client.chat.completions.create(
    model="gpt-4o-mini",
    messages=[{"role": "user", "content": "Suggest a pasta dish."}],
)
print(response.choices[0].message.content)

go deeper

for a junior

Be able to say that one startup call, Traceloop.init(app_name=...), is the whole integration and that your existing model calls do not change. Know that the SDK patches the client libraries for you.

for a middle

Explain the patching mechanic: instrumentors wrap methods on the provider and framework libraries, so instances you already hold are covered, but only calls made after init are traced. Name the three families covered — provider SDKs, frameworks, vector stores.

for a senior

Show where init belongs in a real startup path and what you check first when spans are missing: init ordering, whether the instrumentation package is installed, and whether the library was blocked. Be ready to describe verifying with a console exporter.

for a principal

Own the argument for OpenLLMetry over a proprietary LLM-only tool: it emits OpenTelemetry spans into the tracing stack you already operate, so LLM latency and cost sit in the same trace as the request that caused them, with one exporter policy rather than two.

## The one-line integration OpenLLMetry's whole pitch is that you do not rewrite your application to get LLM traces. You install `traceloop-sdk`, add one call at startup, and the model calls you already make start producing spans: `Traceloop.init(app_name="recipe-bot")` Everything else in this leaf — decorators, exporter choice, batching, selective instrumentation — is a refinement of that one call. ## What init sets up Two things happen. First, the SDK establishes the tracing pipeline that finished spans travel through: a tracer, a span processor, and an exporter with a destination. By default that destination is Traceloop's own endpoint (configured with an API key), but you can hand it any OpenTelemetry span exporter or point it at your own OTLP endpoint instead. Second, it activates the *instrumentation* packages bundled with the SDK — separate `opentelemetry-instrumentation-*` distributions, one per library — that know how to trace a specific client. ## Monkey-patching, not a wrapper client This is the mechanic interviewers are checking you understand. OpenLLMetry does **not** give you a special client class to use instead of `openai.OpenAI`. Each instrumentor wraps functions on the target library at import time: it locates the method that performs the request and replaces it with a wrapper that opens a span, calls the original, records the response, and ends the span. Because the patch lands on the library's *class methods*, every instance you hold — including one constructed before `init` ran — goes through the wrapper from that moment on. What is never captured is a call that already completed before `init` executed. The consequence for code layout: call `init` once, as early as you reasonably can in process startup (module import of your app entrypoint, a FastAPI lifespan/startup hook, the top of a script), not lazily inside a request handler. ## What is covered out of the box Coverage spans three families rather than one: - **Provider SDKs** — the OpenAI and Anthropic Python clients and other model providers. - **Frameworks** — LangChain and LlamaIndex, so chains, retrievers and agent steps appear as spans in their own right rather than as a flat list of model calls. - **Vector stores** — clients such as Pinecone, Chroma, Qdrant and Weaviate, so the retrieval half of a RAG request is visible next to the generation half. Each of those is a separate instrumentation package. If the package for a library is not installed, that library simply is not traced — there is no error, just missing spans. Which instrumentors run can also be narrowed explicitly through `init`'s `instruments` and `block_instruments` arguments. ## What the spans contain The emitted spans follow the OpenTelemetry GenAI semantic conventions: which provider and model was called, request parameters such as temperature, and token usage on the response. Prompt and completion text is recorded by default, which is exactly what makes the traces useful and also exactly what a regulated team may need to switch off — the SDK exposes a toggle for content capture. ## What init deliberately does not do `init` gives you *call-level* spans. It has no way to know that five model calls and two retrievals were one logical "answer the support ticket" operation, because that grouping exists only in your code. That is why the SDK also ships `@workflow`, `@task`, `@agent` and `@tool` decorators: you put them on your own functions to create named parent spans that the auto-generated LLM spans nest inside. It also does not replace your existing tracing. OpenLLMetry is OpenTelemetry underneath, so its spans join the same traces as your HTTP and database spans if the rest of your stack is OTel-instrumented — the LLM call becomes another child of the request span rather than a parallel, disconnected timeline. ## The common failure When someone reports "I called `init` and see nothing", the usual causes are: calls made before `init` ran, the instrumentation package for that library not installed, the library blocked in `init`, or a short-lived process exiting before buffered spans were exported. Verifying with a console exporter locally settles which of those it is in about a minute.

  • If a colleague constructs the OpenAI client at module import and calls Traceloop.init() later in main(), do they lose tracing?
    No, as long as no model calls happened in between. The instrumentors patch methods on the client library's classes, not on your object, so an instance created earlier routes through the wrapper once init has applied it. Only calls that actually executed before init are missing. It is still better practice to init first, so nothing depends on that ordering subtlety.
  • How does OpenLLMetry relate to an OpenTelemetry stack you already run?
    It is built on OpenTelemetry rather than beside it — the LLM spans are ordinary OTel spans following the GenAI semantic conventions. If your web framework and database clients are already instrumented, the model call becomes a child span in the same trace, and you can export everything to the same collector or backend instead of maintaining a separate LLM-only tool.
  • You call init and your OpenAI calls are traced, but your HTTP calls to a self-hosted model server are not. Why?
    Auto-instrumentation covers known client libraries. A raw `requests` or `httpx` call to your own inference endpoint is not a recognized LLM client, so no GenAI span is produced — at best you get a generic HTTP span if you have HTTP instrumentation installed. Wrap that call in your own `@task`-decorated function, or add manual instrumentation, if you want it represented as a model step.

saying these in an interview costs you the question

  • Thinking you must replace the OpenAI client with a Traceloop client
  • Believing every call site needs a decorator before anything is traced
  • Assuming init retroactively traces calls made earlier in the process
  • Calling init inside each request handler instead of once at startup
  • Assuming a library is traced even when its instrumentation package is absent

context

open as a page

Which OpenTelemetry GenAI span attributes carry the model name and token counts?

level: juniorimportance: must knowfreq 68%

basics

~10 s

gen_ai.request.model records the model you asked for, gen_ai.response.model the one that answered, and gen_ai.usage.input_tokens plus gen_ai.usage.output_tokens the token counts. Standard names let any backend chart cost without reading your code.

open as a page

Why add OpenLLMetry's @workflow and @task decorators when init already traces model calls?

level: middleimportance: must knowfreq 66%

basics

~20 s

Auto-instrumentation only sees individual library calls. The decorators create named parent spans for your own multi-step logic, so a retrieval plus three model calls appears as one named operation you can time, compare and filter on rather than four unrelated spans.

open as a page

What does setting TRACELOOP_TRACE_CONTENT=false change in OpenLLMetry spans?

level: middleimportance: must knowfreq 58%

basics

~20 s

It stops OpenLLMetry recording prompt and completion message text on spans. Metadata still flows: provider, model, token counts, parameters, latency and errors. Content capture is on by default, so this is the switch a regulated team flips.

open as a page

When should you pass disable_batch=True to OpenLLMetry's Traceloop.init(), and what does it cost?

level: middleimportance: should knowfreq 48%

basics

~20 s

Pass disable_batch=True in short-lived processes — serverless handlers, CLI scripts, notebooks, tests — where the process can end before buffered spans are sent. It exports each span as it finishes, which is safe but adds export work on the calling path, so leave it off in long-running services.

open as a page

How do you make OpenLLMetry send traces to your own backend instead of Traceloop's cloud?

level: middleimportance: should knowfreq 44%

basics

~20 s

Pass an OpenTelemetry span exporter to Traceloop.init(exporter=...), or point the SDK at your own OTLP endpoint with the api_endpoint and headers arguments or the TRACELOOP_BASE_URL and TRACELOOP_HEADERS environment variables. LLM spans then land in whatever backend already receives your traces.

open as a page

Why does gen_ai.provider.name replace gen_ai.system on GenAI spans?

level: middleimportance: should knowfreq 42%

basics

~20 s

gen_ai.system was ambiguous — readers could not tell whether it meant the API being called or the model's vendor. The convention renamed it to gen_ai.provider.name, which names the provider API. gen_ai.system still ships but is deprecated.

open as a page

An OpenLLMetry service shows workflow spans but no LLM spans. How do you diagnose it?

level: seniorimportance: should knowfreq 45%

basics

~20 s

Workflow spans arriving proves the pipeline and exporter work, so the fault is auto-instrumentation, not export. Check that the provider's instrumentation package is installed, that it is not blocked in Traceloop.init(), that the client library version is still one the instrumentor patches, and that the call really goes through that library.

open as a page

How do the instruments and block_instruments arguments of Traceloop.init() control OpenLLMetry?

level: seniorimportance: should knowfreq 36%

basics

~20 s

They narrow auto-instrumentation. Passing instruments={...} turns the default all-on behaviour into an allowlist of only those libraries; block_instruments={...} keeps the default and subtracts specific ones. Both take members of the SDK's Instruments enum and are used to cut span noise, avoid duplicate patching, and quarantine a misbehaving instrumentor.

open as a page

Why can OpenLLMetry's flattened gen_ai.prompt attributes overwhelm a tracing backend?

level: seniorimportance: should knowfreq 40%

basics

~10 s

Each chat message becomes its own indexed attributes — gen_ai.prompt.0.role, gen_ai.prompt.0.content and so on. Long conversations and retrieved context produce many large string attributes, which hit span attribute limits and dominate ingestion cost.

open as a page

Your gen_ai.usage token attributes don't match the provider bill — what do you check?

level: seniorimportance: should knowfreq 36%

basics

~20 s

Check for missing usage on streamed calls, spans lost to sampling, the same call counted twice by overlapping instrumentation, retries and fallbacks recorded as extra spans, and cached or reasoning tokens the two attributes do not break out.

open as a page

How do you set a content-capture policy for gen_ai.prompt attributes across environments?

level: principalimportance: should knowfreq 28%

basics

~20 s

Classify each service's data, then decide per deployment whether spans may hold message text at all. Because the capture switch is process-wide and defaults to on, enforce it in deployment configuration, and give services that cannot capture a reference-based alternative.

open as a page