skip to content

What does Langfuse's drop-in OpenAI wrapper capture automatically?

level: juniorimportance: should knowfreq 56%

answer

  1. change the import, nothing else
  2. it produces a model-call observation
  3. usage read from the response
  4. cost follows from model plus tokens
  5. streaming usage must be requested

basics

~20 s

Importing the OpenAI client from langfuse.openai instead of the openai package makes every call emit a GENERATION observation on its own — model, parameters, messages, completion, token usage, latency and inferred cost — with no other code change.

solid answer

~50 s

You replace `import openai` with `from langfuse.openai import openai` (or import `OpenAI` / `AsyncOpenAI` from there) and leave every call site untouched. The wrapper delegates to the real client and, around each call, records a GENERATION observation carrying the model name, the request parameters, the messages as input, the completion as output, token usage read from the provider's response, latency, and cost derived from the model and that usage. It nests like any other observation, so calling it inside an `@observe`-decorated function puts the generation under that function's span rather than in a trace of its own. Its value over a hand-rolled span is that GENERATION-specific fields — model, tokens, cost — are populated correctly, which is exactly what people get wrong by hand. With streaming you only get exact usage if you ask the provider to return it; otherwise usage may be estimated or absent.

code

python · 16 lines
python
from langfuse import observe, get_client
from langfuse.openai import openai


@observe(name="answer-question")
def ask(question: str) -> str:
    completion = openai.chat.completions.create(
        model="gpt-4o-mini",
        messages=[{"role": "user", "content": question}],
        temperature=0.2,
    )
    return completion.choices[0].message.content


print(ask("what is a trace?"))
get_client().flush()

go deeper

for a junior

Know that you swap the import to langfuse.openai and every OpenAI call is then traced automatically, with model, messages, completion and token usage recorded for you.

for a middle

Explain why it beats a hand-written span: it creates a GENERATION with provider-reported usage, so cost and token dashboards are correct rather than estimated.

for a senior

Raise the operational caveats — streamed calls need usage explicitly requested, message content is captured by default, and the wrapper is in-process rather than a proxy, so it does not add a failure mode to model calls.

for a principal

Decide the standard: wrapper-based capture for supported SDKs, a documented manual pattern for the rest, and a content-capture policy applied before teams start swapping imports across services.

## The integration in one line Langfuse ships a wrapper around the OpenAI Python client. You change the import and nothing else: ``` from langfuse.openai import openai ``` Every `openai.chat.completions.create(...)` call then goes through the real SDK as before, and Langfuse records a GENERATION observation around it. The wrapper also exposes the class-based clients, so `from langfuse.openai import OpenAI, AsyncOpenAI` works for code that instantiates a client rather than using the module-level one. ## What lands on the observation - **model** — the model string you requested. - **model parameters** — temperature, max tokens, top_p and friends, as sent. - **input** — the messages array. - **output** — the completion, including tool calls when the model returns them. - **usage** — input and output token counts, read from the provider's response rather than guessed. - **latency** — measured around the call, including time to first token for streams. - **cost** — derived from the model and the usage numbers against Langfuse's model pricing definitions. - **errors** — a failed call is still recorded, with the error surfaced on the observation, which is what makes rate-limit and timeout patterns visible. ## Why not just wrap it in a span yourself You can, and for a provider Langfuse does not wrap you will. But the hand-rolled version is where the two classic mistakes happen: people record a model call as a generic SPAN (so the tokens and money never reach the analytics), and people fill in usage from their own estimate rather than the provider's response (so the numbers drift from the bill). The wrapper does both correctly, for free, at the cost of one import. It also keeps working when call sites change. A hand-rolled span records the parameters someone remembered to copy across; the wrapper records what was actually sent. ## Nesting The wrapper creates its observation under whatever is currently active. In practice that means: - Called inside an `@observe`-decorated function, the generation becomes that function's child. This is what you want — the span gives business context ("answer-question"), the generation gives the model detail. - Called with nothing active, the generation becomes a trace of its own with a single node. Fine for a script; in an application it is a sign that you have model-call visibility without any application structure around it. ## Streaming Streaming complicates usage. OpenAI's streaming responses do not include token counts unless you request them — `stream_options={"include_usage": True}` on the call. Without that, the wrapper has no authoritative usage to record and you may see estimated or missing token counts on streamed generations, and therefore unreliable cost. If cost accuracy matters and you stream, ask for usage explicitly. ## Limits worth stating - It covers the OpenAI client surface. Other providers reached through their own SDKs need their own integration or manual instrumentation. - It captures message content by default. That is the point of it, and also the thing a regulated team has to think about before switching the import — the content of every prompt and completion goes to the trace store unless redaction is configured at the deployment level. - It is not a proxy. Requests still go straight from your process to the provider; Langfuse sees a copy of the metadata asynchronously, so the wrapper does not sit in the failure path of your model call. ## Interview framing Lead with the one-line swap and the fact that it produces a GENERATION rather than a generic span. Then show you know why that matters — model, provider-reported usage and cost, which is what cost dashboards are built on. Finish with the two caveats an experienced user hits: streaming usage needs to be requested, and message content is captured by default.

  • Why does a streamed call sometimes show no token usage in Langfuse?
    Because OpenAI does not include usage in a streamed response unless you ask for it with stream_options={"include_usage": True}. Without that final usage chunk there is no authoritative count to record, so token numbers may be estimated or absent and the derived cost becomes unreliable. If you stream and care about spend accuracy, request usage explicitly on every streamed call.
  • Does the wrapper put Langfuse in the path of your model request?
    No. It is an SDK-level wrapper, not a proxy: your process still calls the provider directly, and Langfuse receives a copy of the metadata asynchronously in the background. So Langfuse being slow or unreachable degrades your observability, not your model calls — which is a different tradeoff from gateway-style integrations that sit inline.
  • You call a provider Langfuse does not wrap. What do you do instead?
    Instrument it by hand as a GENERATION rather than a plain span, and fill in the fields the wrapper would have: model name, the request parameters, messages as input, completion as output, and usage taken from the provider's own response rather than estimated. Get the type and the usage source right and the call behaves like a wrapped one in every dashboard.

saying these in an interview costs you the question

  • Thinks the wrapper proxies traffic through Langfuse
  • Expects it to trace non-OpenAI SDKs too
  • Assumes streamed calls always report exact token counts
  • Believes it needs a decorator to record anything
  • Says it only logs prompts, not usage or cost

context