skip to content

Datasets & Experiments

Promoting interesting traces into a dataset, then running a target function over it to produce a comparable experiment. This is the loop teams use to prove a prompt or model change was an improvement, not a vibe.

on this pageshow

questions

6

How do you create a LangSmith dataset and add examples with the Python SDK?

level: juniorimportance: must knowfreq 62%

answer

  1. a named collection of examples
  2. two dicts per row
  3. inputs plus optional reference outputs
  4. create_dataset then create_examples
  5. dataset.id is what you pass on

basics

~10 s

Instantiate langsmith.Client, call client.create_dataset(dataset_name=...) to get a Dataset, then client.create_examples(dataset_id=dataset.id, examples=[...]) where each example is a dict with an inputs dict and, usually, a reference outputs dict.

solid answer

~50 s

A LangSmith dataset is a named collection of **examples**, and an example is essentially a pair of dicts: `inputs` (what your application will be called with) and optionally `outputs` (the reference answer you want to compare against). With the Python SDK you do `client = Client()`, then `dataset = client.create_dataset(dataset_name="support-qa", description="...")`, then `client.create_examples(dataset_id=dataset.id, examples=[{"inputs": {...}, "outputs": {...}}, ...])`. `create_example` adds one at a time. The keys inside `inputs` are yours to choose, but they must match what the target function you later evaluate expects to read, and the keys inside `outputs` must match what your evaluators expect as the reference. The most valuable examples usually don't come from your imagination — you promote them from real traces, reading a run with `client.read_run(run_id)` and writing its `inputs`/`outputs` into the dataset (the UI exposes the same thing as an "add to dataset" action on a trace).

code

python · 20 lines
python
from langsmith import Client

client = Client()
dataset = client.create_dataset(
    dataset_name="support-qa",
    description="Curated support questions with reference answers",
)
client.create_examples(
    dataset_id=dataset.id,
    examples=[
        {
            "inputs": {"question": "How do I reset my password?"},
            "outputs": {"answer": "Use the Forgot password link on the sign-in page."},
        },
        {
            "inputs": {"question": "Where do I find invoices?"},
            "outputs": {"answer": "Under Billing in account settings."},
        },
    ],
)

go deeper

for a junior

Be able to name the two pieces of an example — an inputs dict and an optional reference outputs dict — and write the create_dataset then create_examples sequence from memory without looking it up.

for a middle

Explain the key-naming contract: dataset input keys, what the target function reads, and what evaluators read as reference and prediction all have to line up, and mismatches surface as errored rows rather than low scores.

for a senior

Show how you keep the dataset alive from real traffic — promoting failing traces with hand-written references, tagging examples with metadata that lets you slice later, and checking what customer text you are allowed to retain in a hosted dataset.

for a principal

Own the question of what the dataset represents for the org: who is allowed to add rows, whether references are human-reviewed, and how you avoid a set that only tests what the current prompt already does well.

## What a dataset is in LangSmith A dataset is a named, server-side collection of **examples**. An example is a small record with an `inputs` dictionary, an optional `outputs` dictionary (the reference, sometimes called the expected answer or ground truth), and optional `metadata`. Nothing about the dataset is tied to a particular model, prompt, or framework — it is just the input/reference pairs. That separation is the whole point: the dataset is the fixed thing, and every prompt or model you try is run *against* it so the results are comparable. ## Creating the dataset ``` from langsmith import Client client = Client() dataset = client.create_dataset(dataset_name="support-qa", description="...") ``` `Client()` picks up credentials from the environment (`LANGSMITH_API_KEY`), so you rarely pass them in code. `create_dataset` returns a `Dataset` object; the field you actually need afterwards is `dataset.id`. Dataset names are unique within a workspace, so re-running the snippet against an existing name raises rather than silently reusing — scripts that must be idempotent typically try `client.read_dataset(dataset_name=...)` first and create only on failure. ## Adding examples ``` client.create_examples( dataset_id=dataset.id, examples=[{"inputs": {"question": "..."}, "outputs": {"answer": "..."}}], ) ``` `create_examples` takes a list of dicts and writes them in one bulk request, which matters when you are seeding hundreds of rows — a loop over `create_example` costs one HTTP round trip each. `create_example` is the single-row version and is what you reach for when promoting one interesting trace. ## The key-naming contract The SDK does not validate your key names, which is where beginners lose an afternoon. Three places have to agree: 1. **`inputs` keys** must match what the target function reads. If your example is `{"question": ...}` and your target does `inputs["query"]`, every row raises `KeyError` and the whole experiment shows up as errors rather than low scores. 2. **`outputs` keys** must match what your evaluators read as the reference. 3. **Your target's return keys** must match what evaluators read as the prediction. Pick the names once, write them in the dataset description, and keep them stable. Renaming an input key later means rewriting every example, because a dataset is not schema-migrated for you. ## Promoting real traces Synthetic examples written by the same person who wrote the prompt tend to test what the prompt already does well. The examples worth having are the ones your application actually got wrong. Given a run id from a trace, you can read it and turn it into an example: ``` run = client.read_run(run_id) client.create_example(inputs=run.inputs, outputs=run.outputs, dataset_id=dataset.id) ``` Note what this gives you: the run's *actual* output becomes the reference. That is only correct when the output was good. For a failure case you keep the inputs and hand-write the `outputs` to be what the system *should* have said — otherwise you have pinned the bug in place as ground truth. ## Examples without references `outputs` is optional. A dataset of inputs alone is perfectly usable when your evaluators are reference-free — a judge that scores tone, a checker that asserts the response parses as valid JSON, a rule that asserts no phone number appears. This matters because production traffic never comes with ground truth attached; a reference-requiring metric simply cannot be run over it. Decide which kind of dataset you are building before you seed a thousand rows. ## Metadata is worth filling in Each example can carry `metadata`. Putting the source (`"from_production_incident_4412"`), a customer segment, or a difficulty label there costs nothing at write time and later lets you filter and slice, and lets you find the row again when a result looks wrong. ## Common mistakes - Storing raw prompt strings as the input instead of the variables. If the input is the fully-rendered prompt, you can never evaluate a *new* prompt against the dataset — the prompt is baked into the data. Store the variables. - Putting personally identifiable customer text into a hosted dataset without checking whether you are allowed to retain it. A dataset is durable storage, unlike a sampled trace. - Treating the dataset as write-once. It should grow every time production surprises you.

  • When would you deliberately create examples with no outputs at all?
    When your evaluators are reference-free — a judge scoring tone or helpfulness, a structural check that the response is valid JSON, a rule that no email address leaks. Production traffic arrives without ground truth, so a dataset promoted straight from live traces often has inputs only. Reference-requiring checks like exact match simply cannot run on those rows, so decide which style of dataset you are building before seeding it.
  • You promote a traced run straight into a dataset with its own output as the reference. What is the risk?
    You pin the current behaviour as the definition of correct. That is fine for a run you verified was good, but for a failure case it is exactly backwards — the regression becomes the target. For failures, keep the run's inputs and hand-write the reference outputs to what the system should have produced, ideally with a human reviewing it.
  • Why bulk-write with create_examples instead of looping over create_example?
    create_examples sends the whole list in one request, while a loop pays an HTTP round trip per row. Seeding a few hundred examples one at a time is noticeably slow and much more likely to fail partway, leaving a half-populated dataset. Use the single-row call for the one-off case of promoting an individual trace.

saying these in an interview costs you the question

  • Thinking a dataset stores prompts or model settings
  • Storing the rendered prompt instead of the input variables
  • Assuming outputs is mandatory on every example
  • Expecting the SDK to validate input key names
  • Believing a dataset can only be seeded by hand

context

open as a page

In LangSmith's evaluate(), what must the target function accept and return?

level: middleimportance: must knowfreq 74%

basics

~20 s

The target takes one argument — a single example's inputs dict — and returns a dict of its outputs. LangSmith calls it once per example in the dataset and records each call as a run inside one named experiment.

open as a page

How do you run a LangSmith experiment over only one split of a dataset?

level: middleimportance: should knowfreq 40%

basics

~10 s

Splits are labels attached to examples inside a single dataset. Fetch just the labelled ones with client.list_examples(dataset_name=..., splits=["test"]) and pass that as the data argument of evaluate(), since data accepts any iterable of examples.

open as a page

Your LangSmith experiment over 2,000 examples keeps hitting provider 429s — how do you tune the run?

level: seniorimportance: should knowfreq 46%

basics

~20 s

Lower max_concurrency on evaluate(), which caps how many examples run in parallel. Count calls per example — one target call plus one per LLM evaluator — since that multiplier, not the example count, sets your request rate. Use aevaluate for async targets and shrink the run to a split while iterating.

open as a page

What does num_repetitions do in LangSmith's evaluate(), and what does it cost?

level: seniorimportance: should knowfreq 34%

basics

~20 s

num_repetitions runs every example N times inside one experiment instead of once, so you can see the spread of a non-deterministic system rather than a single sample. It multiplies both target calls and evaluator calls by N, so cost and wall clock scale linearly with it.

open as a page

Teammates added 40 examples to a LangSmith dataset between two experiments — are the runs comparable?

level: principalimportance: should knowfreq 30%

basics

~20 s

Not directly — the two experiments ran over different populations, so part of the score difference is the new examples, not the change under test. LangSmith versions datasets on every write, so pin gated runs to a tagged version and re-baseline the old configuration whenever the set grows.

open as a page