skip to content

LangSmith

LangChain's hosted platform for tracing runs, promoting real traffic into datasets, evaluating against them, and monitoring in production. It is the reference example of the trace-to-dataset-to-eval loop interviewers expect you to be able to describe.

on this pageshow

explore

questions

24

How do you create a LangSmith dataset and add examples with the Python SDK?

level: juniorimportance: must knowfreq 62%

answer

  1. a named collection of examples
  2. two dicts per row
  3. inputs plus optional reference outputs
  4. create_dataset then create_examples
  5. dataset.id is what you pass on

basics

~10 s

Instantiate langsmith.Client, call client.create_dataset(dataset_name=...) to get a Dataset, then client.create_examples(dataset_id=dataset.id, examples=[...]) where each example is a dict with an inputs dict and, usually, a reference outputs dict.

solid answer

~50 s

A LangSmith dataset is a named collection of **examples**, and an example is essentially a pair of dicts: `inputs` (what your application will be called with) and optionally `outputs` (the reference answer you want to compare against). With the Python SDK you do `client = Client()`, then `dataset = client.create_dataset(dataset_name="support-qa", description="...")`, then `client.create_examples(dataset_id=dataset.id, examples=[{"inputs": {...}, "outputs": {...}}, ...])`. `create_example` adds one at a time. The keys inside `inputs` are yours to choose, but they must match what the target function you later evaluate expects to read, and the keys inside `outputs` must match what your evaluators expect as the reference. The most valuable examples usually don't come from your imagination — you promote them from real traces, reading a run with `client.read_run(run_id)` and writing its `inputs`/`outputs` into the dataset (the UI exposes the same thing as an "add to dataset" action on a trace).

code

python · 20 lines
python
from langsmith import Client

client = Client()
dataset = client.create_dataset(
    dataset_name="support-qa",
    description="Curated support questions with reference answers",
)
client.create_examples(
    dataset_id=dataset.id,
    examples=[
        {
            "inputs": {"question": "How do I reset my password?"},
            "outputs": {"answer": "Use the Forgot password link on the sign-in page."},
        },
        {
            "inputs": {"question": "Where do I find invoices?"},
            "outputs": {"answer": "Under Billing in account settings."},
        },
    ],
)

go deeper

for a junior

Be able to name the two pieces of an example — an inputs dict and an optional reference outputs dict — and write the create_dataset then create_examples sequence from memory without looking it up.

for a middle

Explain the key-naming contract: dataset input keys, what the target function reads, and what evaluators read as reference and prediction all have to line up, and mismatches surface as errored rows rather than low scores.

for a senior

Show how you keep the dataset alive from real traffic — promoting failing traces with hand-written references, tagging examples with metadata that lets you slice later, and checking what customer text you are allowed to retain in a hosted dataset.

for a principal

Own the question of what the dataset represents for the org: who is allowed to add rows, whether references are human-reviewed, and how you avoid a set that only tests what the current prompt already does well.

## What a dataset is in LangSmith A dataset is a named, server-side collection of **examples**. An example is a small record with an `inputs` dictionary, an optional `outputs` dictionary (the reference, sometimes called the expected answer or ground truth), and optional `metadata`. Nothing about the dataset is tied to a particular model, prompt, or framework — it is just the input/reference pairs. That separation is the whole point: the dataset is the fixed thing, and every prompt or model you try is run *against* it so the results are comparable. ## Creating the dataset ``` from langsmith import Client client = Client() dataset = client.create_dataset(dataset_name="support-qa", description="...") ``` `Client()` picks up credentials from the environment (`LANGSMITH_API_KEY`), so you rarely pass them in code. `create_dataset` returns a `Dataset` object; the field you actually need afterwards is `dataset.id`. Dataset names are unique within a workspace, so re-running the snippet against an existing name raises rather than silently reusing — scripts that must be idempotent typically try `client.read_dataset(dataset_name=...)` first and create only on failure. ## Adding examples ``` client.create_examples( dataset_id=dataset.id, examples=[{"inputs": {"question": "..."}, "outputs": {"answer": "..."}}], ) ``` `create_examples` takes a list of dicts and writes them in one bulk request, which matters when you are seeding hundreds of rows — a loop over `create_example` costs one HTTP round trip each. `create_example` is the single-row version and is what you reach for when promoting one interesting trace. ## The key-naming contract The SDK does not validate your key names, which is where beginners lose an afternoon. Three places have to agree: 1. **`inputs` keys** must match what the target function reads. If your example is `{"question": ...}` and your target does `inputs["query"]`, every row raises `KeyError` and the whole experiment shows up as errors rather than low scores. 2. **`outputs` keys** must match what your evaluators read as the reference. 3. **Your target's return keys** must match what evaluators read as the prediction. Pick the names once, write them in the dataset description, and keep them stable. Renaming an input key later means rewriting every example, because a dataset is not schema-migrated for you. ## Promoting real traces Synthetic examples written by the same person who wrote the prompt tend to test what the prompt already does well. The examples worth having are the ones your application actually got wrong. Given a run id from a trace, you can read it and turn it into an example: ``` run = client.read_run(run_id) client.create_example(inputs=run.inputs, outputs=run.outputs, dataset_id=dataset.id) ``` Note what this gives you: the run's *actual* output becomes the reference. That is only correct when the output was good. For a failure case you keep the inputs and hand-write the `outputs` to be what the system *should* have said — otherwise you have pinned the bug in place as ground truth. ## Examples without references `outputs` is optional. A dataset of inputs alone is perfectly usable when your evaluators are reference-free — a judge that scores tone, a checker that asserts the response parses as valid JSON, a rule that asserts no phone number appears. This matters because production traffic never comes with ground truth attached; a reference-requiring metric simply cannot be run over it. Decide which kind of dataset you are building before you seed a thousand rows. ## Metadata is worth filling in Each example can carry `metadata`. Putting the source (`"from_production_incident_4412"`), a customer segment, or a difficulty label there costs nothing at write time and later lets you filter and slice, and lets you find the row again when a result looks wrong. ## Common mistakes - Storing raw prompt strings as the input instead of the variables. If the input is the fully-rendered prompt, you can never evaluate a *new* prompt against the dataset — the prompt is baked into the data. Store the variables. - Putting personally identifiable customer text into a hosted dataset without checking whether you are allowed to retain it. A dataset is durable storage, unlike a sampled trace. - Treating the dataset as write-once. It should grow every time production surprises you.

  • When would you deliberately create examples with no outputs at all?
    When your evaluators are reference-free — a judge scoring tone or helpfulness, a structural check that the response is valid JSON, a rule that no email address leaks. Production traffic arrives without ground truth, so a dataset promoted straight from live traces often has inputs only. Reference-requiring checks like exact match simply cannot run on those rows, so decide which style of dataset you are building before seeding it.
  • You promote a traced run straight into a dataset with its own output as the reference. What is the risk?
    You pin the current behaviour as the definition of correct. That is fine for a run you verified was good, but for a failure case it is exactly backwards — the regression becomes the target. For failures, keep the run's inputs and hand-write the reference outputs to what the system should have produced, ideally with a human reviewing it.
  • Why bulk-write with create_examples instead of looping over create_example?
    create_examples sends the whole list in one request, while a loop pays an HTTP round trip per row. Seeding a few hundred examples one at a time is noticeably slow and much more likely to fail partway, leaving a half-populated dataset. Use the single-row call for the one-off case of promoting an individual trace.

saying these in an interview costs you the question

  • Thinking a dataset stores prompts or model settings
  • Storing the rendered prompt instead of the input variables
  • Assuming outputs is mandatory on every example
  • Expecting the SDK to validate input key names
  • Believing a dataset can only be seeded by hand

context

open as a page

What signature and return value does a custom LangSmith evaluator function need?

level: juniorimportance: must knowfreq 78%

basics

~20 s

A LangSmith evaluator is an ordinary Python function that declares any subset of the arguments inputs, outputs and reference_outputs, and returns a dict such as {"key": "exact_match", "score": 1}. You pass it in the evaluators list of evaluate().

open as a page

How do you attach an end-user thumbs-down to a LangSmith run from your app?

level: juniorimportance: must knowfreq 52%

basics

~20 s

Capture the traced call's run id, then call Client.create_feedback with that run id, a feedback key such as "user_score", and a score. The key names the metric, the score is the number, and comment or correction can carry the user's words or the fixed answer.

open as a page

How do you enable LangSmith tracing for a plain Python function with @traceable?

level: juniorimportance: must knowfreq 78%

basics

~10 s

Set LANGSMITH_TRACING=true and LANGSMITH_API_KEY in the environment, then decorate the function with @traceable from the langsmith package. Each call is logged as a run in the LANGSMITH_PROJECT project, nested under any traced caller.

open as a page

In LangSmith's evaluate(), what must the target function accept and return?

level: middleimportance: must knowfreq 74%

basics

~20 s

The target takes one argument — a single example's inputs dict — and returns a dict of its outputs. LangSmith calls it once per example in the dataset and records each call as a run inside one named experiment.

open as a page

In LangSmith, what does an automation rule do to the runs in a tracing project?

level: middleimportance: must knowfreq 62%

basics

~20 s

A LangSmith automation rule watches a tracing project, keeps the runs matching its filter, samples a percentage of those, and sends each sampled run to an action: an online evaluator, an annotation queue, or a dataset.

open as a page

What does LangSmith's wrap_openai add that a plain @traceable does not?

level: middleimportance: must knowfreq 66%

basics

~20 s

wrap_openai patches an OpenAI client so every model call becomes an llm run carrying the exact messages, model parameters and token usage the API returned — which is what LangSmith needs to show token counts and cost. A plain @traceable records only your function's inputs and outputs.

open as a page

Why can't a LangSmith online evaluator score live traffic against a reference answer?

level: seniorimportance: must knowfreq 48%

basics

~20 s

Production runs carry an input and an output but no ground truth, so a metric that needs an expected answer has nothing to compare against. Online rules need reference-free signals; the reference only exists on dataset examples, which is where reference-based metrics belong.

open as a page

In LangSmith, how do you pick an online evaluator's sampling rate at high volume?

level: seniorimportance: must knowfreq 55%

basics

~20 s

Work backwards from the judge bill: sampled runs equal eligible runs times the rate, and each sampled run costs at least one judge model call. Narrow the rule's filter first, then set a rate that yields a few hundred to a few thousand scored runs a day.

open as a page

How do you run a LangSmith experiment over only one split of a dataset?

level: middleimportance: should knowfreq 40%

basics

~10 s

Splits are labels attached to examples inside a single dataset. Fetch just the labelled ones with client.list_examples(dataset_name=..., splits=["test"]) and pass that as the data argument of evaluate(), since data accepts any iterable of examples.

open as a page

In LangSmith's evaluate_comparative, what does the evaluator receive and return?

level: middleimportance: should knowfreq 40%

basics

~20 s

evaluate_comparative takes two or more already-finished experiments over the same dataset and, for each example, hands the evaluator that example's inputs plus a list of outputs — one per experiment, in order. The evaluator returns a ranking, a score per experiment position.

open as a page

How do you score a finished LangSmith experiment with a new evaluator?

level: middleimportance: should knowfreq 42%

basics

~20 s

Call evaluate_existing with the finished experiment's name or ID and the new evaluators, or aevaluate_existing for async ones. It fetches the recorded runs, applies the evaluators, and writes the new feedback onto the same experiment without executing your application again.

open as a page

What does a LangSmith summary evaluator score, and what is it passed?

level: middleimportance: should knowfreq 55%

basics

~20 s

A summary evaluator scores an experiment as a whole rather than one row. Passed through the summary_evaluators argument, it runs once after all rows finish and receives the full lists of inputs, outputs and reference_outputs, returning a single key/score for the run.

open as a page

What is a LangSmith annotation queue, and how do runs end up in one?

level: middleimportance: should knowfreq 45%

basics

~20 s

An annotation queue is a review worklist of production runs. A human opens it, sees one run's input and output at a time, applies the queue's rubric, and their verdict is written back as human-sourced feedback on that run.

open as a page

In LangSmith, what does the run_type on @traceable actually change?

level: middleimportance: should knowfreq 52%

basics

~20 s

run_type classifies a run — chain (the default), llm, tool, retriever, prompt, parser or embedding. It changes how LangSmith renders the run, whether tokens and cost are attributed to it, and it becomes a filter dimension in the project view and in list_runs.

open as a page

How do tags and metadata make LangSmith runs findable, and where do you set them?

level: middleimportance: should knowfreq 58%

basics

~20 s

Tags are short labels for coarse slicing (environment, variant); metadata is arbitrary key/value data for identifiers (user id, prompt version, session). Set either on @traceable, per call via langsmith_extra, or for a whole request with tracing_context. Both are searchable in the project view and via Client.list_runs filters.

open as a page

Your LangSmith experiment over 2,000 examples keeps hitting provider 429s — how do you tune the run?

level: seniorimportance: should knowfreq 46%

basics

~20 s

Lower max_concurrency on evaluate(), which caps how many examples run in parallel. Count calls per example — one target call plus one per LLM evaluator — since that multiplier, not the example count, sets your request rate. Use aevaluate for async targets and shrink the run to a split while iterating.

open as a page

What does num_repetitions do in LangSmith's evaluate(), and what does it cost?

level: seniorimportance: should knowfreq 34%

basics

~20 s

num_repetitions runs every example N times inside one experiment instead of once, so you can see the spread of a non-deterministic system rather than a single sample. It multiplies both target calls and evaluator calls by N, so cost and wall clock scale linearly with it.

open as a page

How do you keep judge cost sane when adding LLM evaluators to a LangSmith experiment?

level: seniorimportance: should knowfreq 52%

basics

~20 s

Count first: judge calls equal rows times repetitions times judged evaluators. Then cut the multipliers — put free deterministic checks first, keep one judged metric rather than four, emit several sub-scores from a single judge call, run the expensive judge on a sample or on a nightly schedule.

open as a page

Prompts traced to LangSmith contain PII — how do you keep content out of runs?

level: seniorimportance: should knowfreq 45%

basics

~20 s

Strip payloads before they leave the process: construct langsmith.Client(hide_inputs=True, hide_outputs=True) — or set LANGSMITH_HIDE_INPUTS/LANGSMITH_HIDE_OUTPUTS — to drop them wholesale, or pass process_inputs/process_outputs to @traceable to redact selected fields. Structure, latency, errors and token counts still ship.

open as a page

Teammates added 40 examples to a LangSmith dataset between two experiments — are the runs comparable?

level: principalimportance: should knowfreq 30%

basics

~20 s

Not directly — the two experiments ran over different populations, so part of the score difference is the new examples, not the change under test. LangSmith versions datasets on every write, so pin gated runs to a tagged version and re-baseline the old configuration whenever the set grows.

open as a page

In LangSmith, how do you split coverage between an online judge and human review?

level: principalimportance: should knowfreq 36%

basics

~20 s

Two different budgets: dollars for the judge, reviewer-hours for the queue. Send the judge wide and cheap across all traffic for a trend line, and send humans narrow and deep into the slices where being wrong is expensive and where their labels calibrate the judge.

open as a page

At high request volume, how do you decide which LangSmith runs to trace at all?

level: principalimportance: should knowfreq 34%

basics

~20 s

Sample at the root: LANGSMITH_TRACING_SAMPLING_RATE keeps a fraction of traces whole, and tracing_context(enabled=...) lets code decide per request from cheap signals. Trace internal, canary and high-value traffic fully, sample the bulk, and route classes of traffic to separate projects.

open as a page

In the langsmith SDK, when do you need RunEvaluator instead of a plain function?

level: middleimportance: nice to knowfreq 30%

basics

~20 s

Rarely. The SDK wraps any plain function into a RunEvaluator for you. Subclass RunEvaluator, or use the @run_evaluator decorator, when the evaluator needs constructor-held state such as a configured judge client, or must implement evaluate_run against the Run and Example objects directly.

open as a page