skip to content

How do tags and metadata make LangSmith runs findable, and where do you set them?

level: middleimportance: should knowfreq 58%

answer

  1. labels for slicing versus facts for identifying
  2. low cardinality versus high cardinality
  3. decorator, per call, or request scope
  4. tracing_context wraps a whole block
  5. has(tags, ...) in the filter box

basics

~20 s

Tags are short labels for coarse slicing (environment, variant); metadata is arbitrary key/value data for identifiers (user id, prompt version, session). Set either on @traceable, per call via langsmith_extra, or for a whole request with tracing_context. Both are searchable in the project view and via Client.list_runs filters.

solid answer

~50 s

Without them a busy project is a firehose of runs you can only scroll. **Tags** are a list of short strings — `"prod"`, `"canary"`, `"variant-b"` — meant for low-cardinality slicing. **Metadata** is a dict of key/value pairs meant for the high-cardinality facts about one execution: user id, request id, prompt version, model revision, tenant. You attach them three ways: statically on the decorator, `@traceable(tags=[...], metadata={...})`; per call, by passing `langsmith_extra={"tags": [...], "metadata": {...}}` to the decorated function; or for a whole request scope with `with tracing_context(tags=[...], metadata={...}):`, which applies to every run created inside the block, so a request-scoped user id lands on the whole trace rather than one node. Both are filterable in the project view, and the same filter string can be passed to `Client.list_runs(filter='has(tags, "canary")')`. Metadata keys `session_id`, `thread_id` or `conversation_id` additionally group traces into a thread view.

code

python · 14 lines
python
from langsmith import traceable
from langsmith.run_helpers import tracing_context


@traceable(tags=["retrieval"], metadata={"index": "docs-v3"})
def search_docs(query: str) -> list[str]:
    return ["doc-1", "doc-2"]


with tracing_context(
    tags=["prod", "canary"],
    metadata={"user_id": "u_123", "session_id": "conv_9", "prompt_version": "v7"},
):
    search_docs("tides")

go deeper

for a junior

Know that tags are a list of short labels and metadata is a key/value dict, and that you can pass both to @traceable so runs can be filtered later.

for a middle

Explain the cardinality rule and the three attachment points — decorator, per-call langsmith_extra, and a tracing_context block covering a whole request.

for a senior

Design the labelling scheme before the incident: request-scope middleware, release sha and prompt version on every run, thread ids for multi-turn, and a filter query you can hand to on-call.

for a principal

Own it as a schema across teams — an agreed vocabulary of tags and metadata keys, plus a rule against smuggling personal data into metadata, since inconsistent labelling silently caps what the platform can ever answer.

## Why this is the difference between a trace store and a haystack A production LangSmith project accumulates thousands of runs an hour. The debugging questions you actually get are not "show me the last run" — they are "show me what happened for the customer who complained at 14:20", "compare the canary prompt against the current one", "find every run of the summariser that errored since the deploy". None of those are answerable from run names and timestamps. They are answerable from the labels you put on runs at the time they were created, which means the decision has to be made before the incident, not during it. ## Tags Tags are a flat list of strings on a run. Treat them as **low-cardinality dimensions** — values from a small closed set that you will want to slice by repeatedly: - environment: `"prod"`, `"staging"` - rollout: `"canary"`, `"baseline"` - experiment arm: `"variant-a"`, `"variant-b"` - surface: `"mobile"`, `"web"` A tag is a yes/no membership test. In the filter language that is `has(tags, "canary")`, composable with `and(...)` and `or(...)`. Do not put a user id or a request id in tags: you end up with an unbounded tag vocabulary that is useless as a facet. ## Metadata Metadata is a JSON-serialisable dict on the run, and this is where identifiers and versions belong: - `user_id`, `tenant_id`, `request_id` - `prompt_version`, `model_revision`, `git_sha` - feature-flag state, retrieval index name You filter on key/value pairs, so `prompt_version: "v7"` gives you every run of a specific prompt release, and `git_sha` lets you attribute a regression to a deployment. Anything a support ticket might quote at you belongs in metadata — that is the field you will search when someone says "my request failed". ## The three places to set them **On the decorator** — for facts that are true of every execution of that function: ``` @traceable(tags=["retrieval"], metadata={"index": "docs-v3"}) ``` **Per call** — pass `langsmith_extra` as a keyword argument to the decorated function; the decorator consumes it and your function never sees it. Useful when the caller knows something the callee does not. **Per request scope** — `with tracing_context(tags=[...], metadata={...}):` wraps a block, and everything traced inside it carries the values. This is the right hook for a web framework: in middleware, open a `tracing_context` with the authenticated user, the request id and the release sha, then call the handler. Every run in the resulting trace is labelled without any of your business functions knowing that LangSmith exists. `tracing_context` also takes `project_name` and `enabled`, so the same hook can route or suppress traffic. ## Searching afterwards The project view exposes a filter box using the same query language you pass to the SDK: `has(tags, "canary")`, equality on run fields, `and(...)`/`or(...)` composition, and time bounds. Programmatically: ``` Client().list_runs(project_name="checkout-assistant", filter='has(tags, "canary")', is_root=True) ``` `is_root=True` restricts to whole-request runs rather than every nested step, which is usually what you want when counting or exporting. `run_type=` narrows to one class of step. The results are what you triage, and — separately — what a team later promotes into datasets, which is the practical reason people care about labelling quality: unlabelled traffic cannot be selected on any axis worth selecting on. ## Threads Multi-turn applications need traces grouped by conversation, not just by request. LangSmith reads the metadata keys `session_id`, `thread_id` or `conversation_id` and groups traces carrying the same value into a thread view, so you can read a conversation end to end. It costs one metadata entry set in your request-scope context, and retrofitting it after the fact is impossible for traffic already recorded. ## Anti-patterns Encoding facts into the run **name** (`answer_question_user_1234`) — names are for identifying the step, and a per-user name makes aggregation impossible. Dumping the entire request object into metadata, which bloats every run and can smuggle personal data into the trace store. Using tags for identifiers, producing thousands of one-off tags. And labelling only the root run when the interesting failure is three levels down — a request-scope `tracing_context` avoids exactly that.

  • Why is a user id a bad tag but a good metadata key?
    Tags are membership labels meant to be a small closed set you can facet by; a user id makes the tag vocabulary unbounded, so the facet lists thousands of one-use values and helps nobody. As a metadata key/value pair it stays a searchable field, you can filter on the specific value when a ticket names it, and other users' runs stay findable by tag.
  • Where in a web service would you set request-scoped tracing labels?
    In middleware, before the handler runs: open a `tracing_context` carrying the authenticated user id, the request id, the release sha and the environment tag, and call the handler inside it. Every run created downstream inherits the labels, so the business code never mentions LangSmith and no nested run ends up unlabelled.
  • How do you get multi-turn conversations to read as one thread rather than separate traces?
    Put a stable conversation identifier in run metadata under `session_id`, `thread_id` or `conversation_id`; LangSmith groups traces sharing that value into a thread view you can read end to end. Set it in the request-scope context so every turn carries it — it cannot be added retroactively to traffic already recorded.

saying these in an interview costs you the question

  • Puts user or request ids into tags
  • Encodes variable facts into the run name instead
  • Labels only the root run and not nested steps
  • Dumps the whole request payload into metadata
  • Assumes labels can be added to runs after the fact

context