skip to content

An OpenLLMetry service shows workflow spans but no LLM spans. How do you diagnose it?

level: seniorimportance: should knowfreq 45%

answer

  1. decorator spans arriving exonerates the exporter
  2. the fault is patching, not export
  3. package installed? blocked? version drifted?
  4. is the call even going through that client?
  5. reproduce with a console exporter and one call

basics

~20 s

Workflow spans arriving proves the pipeline and exporter work, so the fault is auto-instrumentation, not export. Check that the provider's instrumentation package is installed, that it is not blocked in Traceloop.init(), that the client library version is still one the instrumentor patches, and that the call really goes through that library.

solid answer

~50 s

Split the problem the moment you hear it: decorator spans reach the backend, so init ran, the exporter is configured and the network path works. Everything downstream of span creation is exonerated, which leaves patching. Work through four causes in order. First, the instrumentation package for that provider may not be installed in the image — auto-instrumentation fails silently when a package is absent. Second, the library may be excluded by `instruments` or `block_instruments` in `Traceloop.init()`, possibly by an allowlist that predates the dependency. Third, a client-library upgrade may have moved the methods the instrumentor patches, so the wrapper never attaches. Fourth, the call may bypass the instrumented library entirely — raw `httpx` to an inference endpoint, a gateway, or a framework path that is not covered. Confirm locally with a console exporter and a minimal script before touching production.

code

python · 14 lines
python
from openai import OpenAI
from opentelemetry.sdk.trace.export import ConsoleSpanExporter
from traceloop.sdk import Traceloop

Traceloop.init(
    app_name="probe",
    exporter=ConsoleSpanExporter(),
    disable_batch=True,
)

OpenAI().chat.completions.create(
    model="gpt-4o-mini",
    messages=[{"role": "user", "content": "ping"}],
)

go deeper

for a junior

Know that decorator spans and model spans come from different mechanisms: decorators are your code, model spans come from auto-instrumentation, so one can work while the other does not.

for a middle

Be able to list the concrete causes — instrumentation package missing, library blocked at init, client version drift, call bypassing the instrumented library — and check the installed packages in the running image rather than the dependency file.

for a senior

Lead with the split: workflow spans arriving proves export, so investigate patching. Reproduce with a console exporter and a one-call script, bisect on the client version, and know that instrumentation fails silently by design.

for a principal

Treat silent instrumentation regressions as a systemic risk. Require a deploy-time smoke check asserting a model span is emitted, pin instrumentation and client versions together, and make coverage gaps visible rather than discovered when a cost dashboard goes flat.

## Read the evidence before you start The detail that makes this tractable is *which* spans are missing. If nothing at all arrived, the suspects are init ordering, exporter destination and process lifetime. Here decorator spans do arrive, which means `Traceloop.init()` executed, a tracer exists, the exporter has a reachable destination, credentials are accepted and spans are being flushed. A candidate who starts by checking the endpoint has not read the evidence. So the failure is in *auto-instrumentation*: the provider call is not producing a span in the first place. ## Cause 1: the instrumentation package is not installed Each library is traced by its own `opentelemetry-instrumentation-*` distribution. If that distribution is missing from the runtime image, that library is simply untraced — no exception, no warning in the application. This happens more than it should: a slimmed production image, a dependency group excluded from the install, a lockfile that resolved differently on the build machine. Check what is actually installed in the running container, not what the dependency file says. ## Cause 2: the library is excluded at init `Traceloop.init()` accepts `instruments` (an exclusive allowlist) and `block_instruments` (a subtraction). An allowlist written a year ago will silently exclude a provider added last month, and a temporary block from an incident can outlive the incident. Read the actual `init` call and any environment-specific config that builds it — this is a one-line cause that can burn an afternoon because nobody thinks to look at the config that was "already working". ## Cause 3: client version drift Instrumentors wrap specific methods on specific classes of the client library. A major upgrade can rename, move or restructure those code paths, and the patch then finds nothing to wrap. The tell is timing: tracing worked, then a dependency bump landed, then LLM spans stopped — with no application error, because instrumentation is designed to fail quietly rather than break the app. Check the client version against what the instrumentation release supports, and pin while you wait for a compatible release. The same shape of failure appears when the client is imported and used in a way that bypasses the patched entry point. ## Cause 4: the call does not go through the instrumented library Coverage is per known client, not per network call. If your service reaches the model through raw `httpx`, a bespoke internal wrapper, an unsupported SDK, or a self-hosted inference endpoint, there is no GenAI span to create — you may see a generic HTTP span if HTTP instrumentation is present, which some people misread as "tracing is on, LLM spans are broken". Similarly, a framework may call the provider along a path the framework instrumentation does not cover. The fix here is not debugging: it is wrapping that call in your own `@task`-decorated function or adding manual instrumentation. ## The verification loop Do not iterate in production. Reproduce with the smallest possible script: `Traceloop.init(app_name="probe", exporter=ConsoleSpanExporter(), disable_batch=True)`, one model call, run it. Two outcomes: - **A model span prints.** Instrumentation works in that environment, so the difference is environmental — the deployed image, its installed packages, its `init` arguments, or the code path production actually takes. - **No model span prints.** You have reproduced the fault in the smallest scope. Now bisect: pin the client library to its previous version, check the installed instrumentation package, and remove any instruments arguments. A console exporter with immediate export is the right tool because it removes batching, network and backend from the experiment entirely. ## Ordering, and why it is usually a red herring here A popular guess is "init ran too late". It is worth ruling out — calls made before `init` are never traced — but it rarely explains a *steady-state* production symptom, because a service that keeps serving requests will have initialized long before the traffic you are looking at. Ordering explains missing spans at startup, in tests, and in scripts. If spans are missing for every request all day, look at patching instead. ## What good looks like in the answer Name the split first (export proven, patching suspect), then the four causes, then the console-exporter reproduction. Add one preventive measure: a smoke check on deploy that asserts a model span was produced, so a silent instrumentation regression from a dependency bump is caught by CI rather than discovered weeks later when someone asks why the cost dashboard flatlined.

  • Why does missing instrumentation never raise an error in the application?
    Telemetry is designed to fail open. An instrumentor that cannot find the method it patches, or whose package is absent, does nothing rather than crash your service — an observability library taking down production would be a far worse failure. The cost is that regressions are silent, which is exactly why a deploy-time smoke check asserting a model span was produced is worth having.
  • How would you catch this class of regression automatically?
    Add a test or post-deploy smoke check that runs one model call with an in-memory or console exporter and asserts an LLM span with the expected attributes was emitted. It runs in seconds and turns a silent breakage from a dependency bump into a failed build. Without it, the usual detection path is someone noticing a cost or latency dashboard has quietly gone flat.
  • Traces show a generic HTTP span to the provider's host but no model span. What does that tell you?
    That the request is leaving through HTTP instrumentation rather than the provider client's instrumented path — typically a raw httpx or requests call, an internal wrapper, or an unsupported SDK. There is nothing to fix in OpenLLMetry's configuration; the call is not one of the libraries it patches. Wrap it in your own decorated function or add manual instrumentation if you need model-level detail.
  • Would checking the exporter endpoint be a reasonable first step here?
    No, and saying so is part of the answer. Decorator spans reached the backend through the same exporter, so destination, credentials and network are already proven. Spending time there ignores the evidence in the symptom. Endpoint checks belong to the different symptom where no spans arrive at all.

saying these in an interview costs you the question

  • Checking the exporter endpoint when decorator spans already arrive
  • Assuming init ordering explains a steady-state production symptom
  • Expecting a failed patch to raise an error in the application
  • Debugging directly in production instead of reproducing with one script
  • Treating a generic HTTP span to the provider as proof the LLM path is instrumented

context