skip to content

What does LangSmith's wrap_openai add that a plain @traceable does not?

level: middleimportance: must knowfreq 66%

answer

  1. generic decorator versus provider-aware wrapper
  2. who knows the token counts?
  3. llm run type, messages, usage
  4. patches the client object once
  5. streamed chunks folded into one output

basics

~20 s

wrap_openai patches an OpenAI client so every model call becomes an llm run carrying the exact messages, model parameters and token usage the API returned — which is what LangSmith needs to show token counts and cost. A plain @traceable records only your function's inputs and outputs.

solid answer

~40 s

`from langsmith.wrappers import wrap_openai` returns a wrapped OpenAI client whose methods are patched: each `chat.completions.create` (or `responses.create`) call emits a run of type `llm` with the message list, the request parameters such as `model` and `temperature`, and the token usage the provider returned. Because the run is typed `llm` and carries usage, LangSmith renders it as a chat conversation and can attribute token counts and cost to it. `@traceable` around your own function records whatever you passed in and returned as opaque JSON — accurate, but the platform has no idea it was a model call, so there is no cost accounting and no message rendering. The two compose: wrap the client once, decorate your orchestration functions, and the model calls nest inside them. A sibling `wrap_anthropic` does the same for the Anthropic client.

code

python · 17 lines
python
from langsmith import traceable
from langsmith.wrappers import wrap_openai
from openai import OpenAI

client = wrap_openai(OpenAI())


@traceable(name="answer_question")
def answer_question(question: str) -> str:
    resp = client.chat.completions.create(
        model="gpt-4o-mini",
        messages=[{"role": "user", "content": question}],
    )
    return resp.choices[0].message.content


print(answer_question("Give me one fact about tides."))

go deeper

for a junior

Know that you wrap the OpenAI client once with wrap_openai and keep using it normally, and that this is what makes model calls show up as chat conversations in LangSmith.

for a middle

Explain why token usage and cost need the provider response envelope, and how a wrapped client's llm run nests under your own @traceable run to form one trace.

for a senior

Discuss what the wrapper puts in the trace store by default — full prompts and completions — and the redaction decision that follows, plus how streaming and async clients behave.

for a principal

Own the tradeoff between an in-process wrapper and a proxy-based approach: request-path risk, provider coverage, and who is accountable for cost attribution across teams and models.

## The problem the wrapper solves `@traceable` is generic. It knows a function was called, with what arguments, and what it returned. It has no idea whether that function talked to a language model, and it certainly does not know how many tokens the provider billed you for. Cost and token analytics in LangSmith are computed from the run's usage numbers plus the model name, and those only exist if something captured the provider's response envelope. That is the job of a provider wrapper. ## What wrap_openai does ``` from langsmith.wrappers import wrap_openai from openai import OpenAI client = wrap_openai(OpenAI()) ``` The call returns the same client object with its completion methods patched. From then on, every request through that client produces an `llm`-typed run containing: - the full **message list** sent, in role/content form, which the UI renders as a conversation rather than a JSON blob; - the **invocation parameters** — `model`, `temperature`, `max_tokens`, tool/function definitions, and so on; - the **response**, including tool calls and finish reason; - the **token usage** from the provider's response (prompt/input and completion/output counts), which is what drives the cost column and the token charts. It also handles **streaming**. A streamed response arrives as many chunks; the wrapper aggregates them into one logical output so the run shows a single completed message instead of hundreds of fragments, and still ends the run with correct timing. It applies to both sync and async clients, and it accepts a `tracing_extra` argument if you want fixed metadata, tags or a project override on every model call made through that client. ## How it composes with @traceable The two are complementary, not alternatives. In a typical application you wrap the provider client once at startup and decorate your own orchestration code: - `@traceable` on `answer_question(...)` gives you the business-level run: the user's question in, the final answer out. - The wrapped client gives you a child `llm` run per model call underneath it, with the prompt actually sent after templating and the tokens actually spent. Because nesting is implicit through the active-run context, you get one tree: business call at the root, retrieval and tool runs beside it, model calls as leaves. That tree is what makes debugging tractable — you can see the exact prompt that produced a bad answer, not the arguments to the function that built it. ## What you lose without it If you call the raw OpenAI client from inside a `@traceable` function, you still get a trace, but: - the run is typed `chain`, so no cost and no token attribution; - the prompt appears only if you happened to return or log it, and the actual rendered prompt is often assembled deep inside your function where the decorator cannot see it; - streaming responses show up as whatever your function chose to return; - you cannot open the call in the playground to iterate on the prompt, because the platform does not have a structured model call to reconstruct. The fallback if you cannot patch the client — say a provider with no wrapper — is to put `@traceable(run_type="llm")` on a thin function that performs the call and to shape its inputs and outputs to look like a model call, including a usage field. That is more work and easier to get subtly wrong, which is precisely why the wrapper exists. ## Practical notes Wrap once, at client construction, and inject the wrapped client everywhere; wrapping repeatedly in a hot path is wasteful and confusing. If tracing is disabled by environment, the wrapped client behaves exactly like the unwrapped one, so it is safe to leave in place across environments. Note also that the wrapper puts prompt and completion **content** into the trace store by default — that is the point of it, and it is also the thing a regulated team has to think about before switching it on, since redaction then has to be configured explicitly rather than inherited from your own logging discipline. Finally, do not confuse the wrapper with a proxy. It is an in-process patch on the SDK object; requests still go directly from your process to the provider, and LangSmith receives an asynchronous copy of the metadata. Nothing sits in the request path, so a LangSmith outage cannot fail your model call.

  • Does wrap_openai put anything in the request path between your service and OpenAI?
    No. It patches the SDK client object in your process, so the HTTP request still goes straight to the provider. The run is shipped to LangSmith asynchronously on a background thread. That means a LangSmith outage degrades your observability, not your product — unlike a proxy-based tool, which sits between you and the provider.
  • How would you trace a provider that has no LangSmith wrapper?
    Put `@traceable(run_type="llm")` on a thin function that makes the call, and shape its inputs and outputs like a model call — a messages list in, the completion plus a usage object out — so the platform can render it and attribute tokens. It is more manual and easier to get wrong, which is the argument for using a supplied wrapper where one exists.
  • What happens to a streamed completion in the trace?
    The wrapper aggregates the streamed chunks into one logical output so the run shows a single assembled message rather than hundreds of fragments, and the run ends when the stream completes, so latency reflects the full stream. Without it, you would see whatever fragments your own code chose to return.

saying these in an interview costs you the question

  • Thinks @traceable alone gives token counts and cost
  • Believes wrap_openai proxies traffic through LangSmith
  • Wraps the client again on every request instead of once
  • Claims cost is estimated from the prompt string length
  • Assumes streaming responses cannot be traced

context