How do you enable LangSmith tracing for a LangGraph app, and what does a trace show?
answer
- configuration, not instrumentation
- two environment variables and you are live
- a run per node execution, in order
- prompts and tool arguments are captured
- uploads are async — flush before exit
basics
~20 sSet LANGSMITH_TRACING=true plus LANGSMITH_API_KEY, and optionally LANGSMITH_PROJECT; no code change is needed. Each graph run becomes a nested trace: a root run, a child run per node execution, and model and tool runs with prompts, latency and token counts.
solid answer
~50 sTracing is configuration, not instrumentation. Setting `LANGSMITH_TRACING=true` and `LANGSMITH_API_KEY` in the environment attaches a global tracer, and every compiled graph invocation in that process is recorded; `LANGSMITH_PROJECT` groups the runs, and the older `LANGCHAIN_TRACING_V2`/`LANGCHAIN_API_KEY`/`LANGCHAIN_PROJECT` names are still honoured in the 1.x line. A trace is the same run tree the fine-grained event stream exposes live, but persisted: root run for the invocation, a child run per node per superstep, and beneath them the model calls with their rendered prompts, outputs, latency and token usage, plus tool calls with arguments and results. Because the sequence of node runs is visible, a trace is how you answer "why did it route there" after the fact. Add `run_name`, `tags` and `metadata` to the run config to make traces findable, and pass a stable `thread_id` in `configurable` so a conversation's runs group together.
code
python · 15 linesimport os
os.environ["LANGSMITH_TRACING"] = "true"
os.environ["LANGSMITH_API_KEY"] = "ls-..."
os.environ["LANGSMITH_PROJECT"] = "checkout-agent"
result = graph.invoke(
{"topic": "cats"},
config={
"run_name": "checkout-turn",
"tags": ["prod", "tier-2"],
"metadata": {"request_id": "req-8813"},
"configurable": {"thread_id": "user-42"},
},
)go deeper
Know that tracing is switched on with environment variables rather than code, and that a trace shows the nodes that ran plus the prompts and responses of each model call.
Explain the run tree — root run, a run per node execution including repeats in a loop, nested model and tool runs with latency and token counts — and how tags and metadata make runs findable.
Show operational judgment: async uploads and flushing in serverless, hiding inputs and outputs for regulated data, sampling for volume, and using a persisted trace to explain a routing decision after the connection is long closed.
Own the observability policy — what leaves the process, retention and residency, how tracing coexists with your own metrics and alerting, and what the team is expected to attach to every run so incidents are answerable.
## Turning it on There is no code to write. The relevant environment variables: - `LANGSMITH_TRACING=true` — the master switch. - `LANGSMITH_API_KEY` — credentials. - `LANGSMITH_PROJECT` — the project runs land in; defaults to a default project if unset. - `LANGSMITH_ENDPOINT` — point at a self-hosted or regional instance. The older `LANGCHAIN_TRACING_V2`, `LANGCHAIN_API_KEY` and `LANGCHAIN_PROJECT` names remain supported, which is why you still see them in older material; prefer the `LANGSMITH_*` spelling in new work. Because the switch is environment-level, the usual mistakes are environmental too: tracing enabled in a CI job that then sends test prompts to the production project, or a container that inherits the key but not the project name. ## What a trace contains One graph invocation produces one root run. Underneath it: - **A run per node execution.** Not per node *definition* — a node visited three times in a loop appears three times, in order, which is what makes cyclic agent behaviour legible. - **Model runs** nested inside the node that called them, showing the fully rendered prompt (after every template substitution), the raw response, latency, and token counts for input and output. - **Tool runs** with the arguments the model produced and the value returned or the exception raised. - **Errors** attached to whichever run failed, with the traceback, so a failure is localised to a node rather than to "the graph". That structure is why tracing answers the routing question. When a conditional edge sent a run somewhere surprising, the trace shows the state the routing decision saw and the model output that produced it — which is almost never what the developer assumed. ## Making traces findable A production project accumulates thousands of runs a day; an untagged trace is unfindable. Attach metadata at invocation time through the run config: - `run_name` — a human label for the root run. - `tags` — coarse buckets: environment, tenant tier, experiment arm. - `metadata` — structured facts you will filter on: user id, request id, model version, feature flag state. Carry your own request id in metadata so a support ticket in your logs maps to a trace in one search. Pass a stable `thread_id` inside `configurable` and a conversation's runs group together, which is what turns "the bot got confused" into a readable multi-turn story. ## Operational realities **Uploads are asynchronous and best-effort.** Tracing is deliberately non-blocking, which means a process that exits immediately after a run can drop the tail of its trace data. In serverless handlers and short-lived CLI jobs, flush the LangSmith client before returning. **Traces contain prompts and outputs.** That is the point, and it is also a data-governance problem: user text, retrieved documents and tool arguments all leave your process. The `LANGSMITH_HIDE_INPUTS` and `LANGSMITH_HIDE_OUTPUTS` environment variables suppress payloads while keeping structure and timing, which is often the right setting for a regulated workload. Otherwise, redact before it reaches state or the prompt. **Volume and cost.** Every node execution is a run. A busy agent service generates a lot of them; sample deliberately rather than discovering the ingest volume from an invoice. **It is not a replacement for your own telemetry.** Application logs, metrics and error tracking still own alerting and SLOs. Tracing owns explaining one specific run. ## Tracing versus streaming Streaming and the live event stream are for the connection that is open right now; tracing is the durable record of a run nobody was watching. Production incidents are almost always the second case — the user is gone, the socket is closed, and all you have is what was persisted. That is the argument for enabling tracing in production even though it costs money, and it is the answer interviewers are listening for when they ask how you would debug a run that took an unexpected path.
- A node runs three times in a loop. How does that appear in the trace?As three separate child runs under the root, in execution order, each with its own model and tool runs underneath. Traces record executions rather than definitions, which is exactly what makes a cyclic agent readable: you can see the loop's state on each pass and identify the iteration where the routing decision went wrong.
- Why can a short-lived serverless handler lose part of a trace?Because trace uploads are asynchronous and non-blocking by design, so the process can return before the background sender has shipped the last runs. Flush the LangSmith client before the handler returns, and treat any tracing gap in serverless as this problem until proven otherwise rather than assuming the runs were never recorded.
- Your workload cannot send user text to a third party. Can you still trace?Yes, in reduced form. The LANGSMITH_HIDE_INPUTS and LANGSMITH_HIDE_OUTPUTS environment variables suppress the payloads while keeping the run tree, timings, token counts and errors, so you retain structural debugging without exporting content. A self-hosted or regional endpoint via LANGSMITH_ENDPOINT is the other lever, alongside redacting before text ever reaches state.
- How do you connect a customer complaint in your logs to the right trace?Put your own correlation id into the run config metadata at invocation time, and a stable thread_id in configurable so the conversation's runs group. Then the support ticket's request id is a single filter in the tracing UI. Relying on timestamps alone fails as soon as the service handles more than a few runs a second.
saying these in an interview costs you the question
- Believes tracing requires wrapping every node by hand
- Assumes traces upload synchronously and never drop
- Ignores that prompts and tool arguments leave the process
- Uses tracing as the alerting and metrics system
- Never sets project, tags or a correlation id in metadata