In LangSmith, what does the run_type on @traceable actually change?
answer
- a classification, not decoration
- chain is the default
- which class earns a cost number?
- documents render only if typed right
- also a filter dimension in list_runs
basics
~20 srun_type classifies a run — chain (the default), llm, tool, retriever, prompt, parser or embedding. It changes how LangSmith renders the run, whether tokens and cost are attributed to it, and it becomes a filter dimension in the project view and in list_runs.
solid answer
~40 s`run_type` is a label you set on `@traceable(run_type="...")` (or that a wrapper sets for you). The platform treats the classes differently rather than just colouring them. An `llm` run is rendered as a chat conversation and is the only class from which token counts and cost are computed, so a model call typed `chain` silently disappears from your cost charts. A `retriever` run renders its output as a list of documents when the output is document-shaped, which is what makes RAG traces readable. `tool` runs render as tool invocations in an agent trace. `chain` is the default and means "a step that orchestrates other steps". The type is also a first-class filter: you can narrow a busy project to just retriever runs, or pass `run_type=` to `Client.list_runs` when pulling runs programmatically.
code
python · 17 linesfrom langsmith import traceable
@traceable(run_type="retriever")
def search_docs(query: str) -> list[dict]:
return [
{"page_content": "Tides follow the moon.", "metadata": {"id": "d1"}},
{"page_content": "Spring tides are largest.", "metadata": {"id": "d2"}},
]
@traceable(run_type="tool")
def word_count(text: str) -> int:
return len(text.split())
print(search_docs("tides"), word_count("one two three"))go deeper
Know the common values — chain, llm, tool, retriever — and that you set them with @traceable(run_type="..."), with chain as the default.
Explain the three things the type actually controls: how the run is rendered, whether tokens and cost are attributed to it, and that it is a filter dimension in the UI and in list_runs.
Argue for typing discipline across a codebase, since the aggregate questions you want to ask later — retriever latency, tool-call counts, per-model spend — only work when every team types its steps consistently.
Own the convention: which types the organisation uses, how it is enforced in review, and the fact that mistyped runs are not retroactively fixable, so the cost of drift compounds.
## The vocabulary Every run carries a `run_type` string. The recognised values are `chain`, `llm`, `tool`, `retriever`, `prompt`, `parser` and `embedding`. When you decorate a function with `@traceable` and say nothing, the run type is `chain`. Provider wrappers set `llm` for you; LangChain and LangGraph instrumentation assigns types automatically to their own components. It is tempting to treat this as cosmetic metadata. It is not — three concrete behaviours hang off it. ## 1. Rendering The UI switches presentation on the type: - **llm** — the inputs are read as a message list and displayed as a conversation with roles, tool calls and finish reason, rather than a raw JSON payload. This is also the type that can be opened in the prompt playground, because the platform can reconstruct a structured model request from it. - **retriever** — the output is read as a list of documents, each with page content and metadata, and rendered as a scannable list. If your retrieval step returns document-shaped output but is typed `chain`, you get an unreadable wall of JSON instead. - **tool** — displayed as a tool invocation with its arguments, which is what makes an agent trace legible: you can see the sequence of tool calls without expanding every node. - **chain** — the generic container rendering, inputs and outputs as JSON. ## 2. Token and cost attribution Cost is computed for `llm` runs from the token usage recorded on the run plus the model name. Only that class participates. The practical failure is a hand-rolled model call wrapped in `@traceable` with no `run_type`: the trace looks fine, latency is right, and the project's cost dashboard shows nothing, because as far as the platform is concerned no model was called. If you must trace a model call by hand, type it `llm` and shape inputs and outputs like a model request and response, including usage. ## 3. Filtering and querying Run type is one of the standard filter dimensions. In the project view you can restrict to a type; programmatically, `Client.list_runs(project_name=..., run_type="retriever")` pulls just those runs. This matters when you want to answer questions about one layer of the stack — "what is p95 latency of my retriever step", "how many tool calls does a typical agent turn make" — without wading through everything else. Combined with `is_root=True` you can separate whole-request runs from internal steps. ## Choosing a type A simple rule works for most code: - The function calls a language model directly → `llm`. - The function fetches candidate context from a store or a search API and returns documents → `retriever`. - The function is something an agent may choose to call — an API lookup, a calculator, a database query → `tool`. - The function turns a raw model output into a structured value → `parser`. - The function renders a template into messages → `prompt`. - The function computes vectors → `embedding`. - Anything that orchestrates the above → `chain`. Do not over-think the rarely used types. The distinctions that pay off in practice are `llm` (cost and playground), `retriever` (readable RAG traces) and `tool` (readable agent traces); the rest are mostly organisational. ## Consistency is the real value A project where every retrieval step is typed `retriever` supports questions the tool can answer for you — filter to retriever runs, sort by latency, see which corpus is slow. A project where half of them are typed `chain` because someone forgot supports nothing; you are back to reading traces one at a time. Because the type is set at the decorator, it is cheap to standardise in code review and worth doing early, since changing it later does not retype the runs already stored. ## What it is not The run type is not a permission boundary, not a sampling dimension, and not a substitute for tags and metadata. It classifies the *kind* of work; tags and metadata describe *this particular* execution — which user, which prompt version, which environment. You want both, and they answer different questions.
- You typed a hand-written model call as chain by mistake. What breaks?The trace still records inputs, outputs, latency and errors, so debugging mostly works — but the run contributes nothing to token or cost aggregation, is rendered as raw JSON rather than a conversation, cannot be opened in the playground, and is invisible when you filter the project to llm runs. Retyping the decorator fixes future runs; the stored ones keep the old type.
- What shape does a retriever run's output need for the document rendering to work?A list of document-shaped objects, each with page content and a metadata mapping. Returning a list of bare strings or a single concatenated context blob gives you a JSON dump instead of a document list, which is why RAG traces from hand-rolled retrieval often read so poorly compared with framework-instrumented ones.
- Is run type a substitute for tags and metadata?No — they answer different questions. Run type says what kind of work this is, and is the same for every execution of that step. Tags and metadata describe this particular execution: environment, prompt version, user, session. You filter on both together, for example all retriever runs tagged for the canary deployment.
saying these in an interview costs you the question
- Treats run_type as a purely cosmetic label
- Expects cost numbers from a model call typed chain
- Thinks run type is assigned automatically for hand-written code
- Confuses run type with tags or metadata
- Believes changing the decorator retypes runs already stored