How do you create a LangSmith dataset and add examples with the Python SDK?
answer
- a named collection of examples
- two dicts per row
- inputs plus optional reference outputs
- create_dataset then create_examples
- dataset.id is what you pass on
basics
~10 sInstantiate langsmith.Client, call client.create_dataset(dataset_name=...) to get a Dataset, then client.create_examples(dataset_id=dataset.id, examples=[...]) where each example is a dict with an inputs dict and, usually, a reference outputs dict.
solid answer
~50 sA LangSmith dataset is a named collection of **examples**, and an example is essentially a pair of dicts: `inputs` (what your application will be called with) and optionally `outputs` (the reference answer you want to compare against). With the Python SDK you do `client = Client()`, then `dataset = client.create_dataset(dataset_name="support-qa", description="...")`, then `client.create_examples(dataset_id=dataset.id, examples=[{"inputs": {...}, "outputs": {...}}, ...])`. `create_example` adds one at a time. The keys inside `inputs` are yours to choose, but they must match what the target function you later evaluate expects to read, and the keys inside `outputs` must match what your evaluators expect as the reference. The most valuable examples usually don't come from your imagination — you promote them from real traces, reading a run with `client.read_run(run_id)` and writing its `inputs`/`outputs` into the dataset (the UI exposes the same thing as an "add to dataset" action on a trace).
code
python · 20 linesfrom langsmith import Client
client = Client()
dataset = client.create_dataset(
dataset_name="support-qa",
description="Curated support questions with reference answers",
)
client.create_examples(
dataset_id=dataset.id,
examples=[
{
"inputs": {"question": "How do I reset my password?"},
"outputs": {"answer": "Use the Forgot password link on the sign-in page."},
},
{
"inputs": {"question": "Where do I find invoices?"},
"outputs": {"answer": "Under Billing in account settings."},
},
],
)go deeper
Be able to name the two pieces of an example — an inputs dict and an optional reference outputs dict — and write the create_dataset then create_examples sequence from memory without looking it up.
Explain the key-naming contract: dataset input keys, what the target function reads, and what evaluators read as reference and prediction all have to line up, and mismatches surface as errored rows rather than low scores.
Show how you keep the dataset alive from real traffic — promoting failing traces with hand-written references, tagging examples with metadata that lets you slice later, and checking what customer text you are allowed to retain in a hosted dataset.
Own the question of what the dataset represents for the org: who is allowed to add rows, whether references are human-reviewed, and how you avoid a set that only tests what the current prompt already does well.
## What a dataset is in LangSmith A dataset is a named, server-side collection of **examples**. An example is a small record with an `inputs` dictionary, an optional `outputs` dictionary (the reference, sometimes called the expected answer or ground truth), and optional `metadata`. Nothing about the dataset is tied to a particular model, prompt, or framework — it is just the input/reference pairs. That separation is the whole point: the dataset is the fixed thing, and every prompt or model you try is run *against* it so the results are comparable. ## Creating the dataset ``` from langsmith import Client client = Client() dataset = client.create_dataset(dataset_name="support-qa", description="...") ``` `Client()` picks up credentials from the environment (`LANGSMITH_API_KEY`), so you rarely pass them in code. `create_dataset` returns a `Dataset` object; the field you actually need afterwards is `dataset.id`. Dataset names are unique within a workspace, so re-running the snippet against an existing name raises rather than silently reusing — scripts that must be idempotent typically try `client.read_dataset(dataset_name=...)` first and create only on failure. ## Adding examples ``` client.create_examples( dataset_id=dataset.id, examples=[{"inputs": {"question": "..."}, "outputs": {"answer": "..."}}], ) ``` `create_examples` takes a list of dicts and writes them in one bulk request, which matters when you are seeding hundreds of rows — a loop over `create_example` costs one HTTP round trip each. `create_example` is the single-row version and is what you reach for when promoting one interesting trace. ## The key-naming contract The SDK does not validate your key names, which is where beginners lose an afternoon. Three places have to agree: 1. **`inputs` keys** must match what the target function reads. If your example is `{"question": ...}` and your target does `inputs["query"]`, every row raises `KeyError` and the whole experiment shows up as errors rather than low scores. 2. **`outputs` keys** must match what your evaluators read as the reference. 3. **Your target's return keys** must match what evaluators read as the prediction. Pick the names once, write them in the dataset description, and keep them stable. Renaming an input key later means rewriting every example, because a dataset is not schema-migrated for you. ## Promoting real traces Synthetic examples written by the same person who wrote the prompt tend to test what the prompt already does well. The examples worth having are the ones your application actually got wrong. Given a run id from a trace, you can read it and turn it into an example: ``` run = client.read_run(run_id) client.create_example(inputs=run.inputs, outputs=run.outputs, dataset_id=dataset.id) ``` Note what this gives you: the run's *actual* output becomes the reference. That is only correct when the output was good. For a failure case you keep the inputs and hand-write the `outputs` to be what the system *should* have said — otherwise you have pinned the bug in place as ground truth. ## Examples without references `outputs` is optional. A dataset of inputs alone is perfectly usable when your evaluators are reference-free — a judge that scores tone, a checker that asserts the response parses as valid JSON, a rule that asserts no phone number appears. This matters because production traffic never comes with ground truth attached; a reference-requiring metric simply cannot be run over it. Decide which kind of dataset you are building before you seed a thousand rows. ## Metadata is worth filling in Each example can carry `metadata`. Putting the source (`"from_production_incident_4412"`), a customer segment, or a difficulty label there costs nothing at write time and later lets you filter and slice, and lets you find the row again when a result looks wrong. ## Common mistakes - Storing raw prompt strings as the input instead of the variables. If the input is the fully-rendered prompt, you can never evaluate a *new* prompt against the dataset — the prompt is baked into the data. Store the variables. - Putting personally identifiable customer text into a hosted dataset without checking whether you are allowed to retain it. A dataset is durable storage, unlike a sampled trace. - Treating the dataset as write-once. It should grow every time production surprises you.
- When would you deliberately create examples with no outputs at all?When your evaluators are reference-free — a judge scoring tone or helpfulness, a structural check that the response is valid JSON, a rule that no email address leaks. Production traffic arrives without ground truth, so a dataset promoted straight from live traces often has inputs only. Reference-requiring checks like exact match simply cannot run on those rows, so decide which style of dataset you are building before seeding it.
- You promote a traced run straight into a dataset with its own output as the reference. What is the risk?You pin the current behaviour as the definition of correct. That is fine for a run you verified was good, but for a failure case it is exactly backwards — the regression becomes the target. For failures, keep the run's inputs and hand-write the reference outputs to what the system should have produced, ideally with a human reviewing it.
- Why bulk-write with create_examples instead of looping over create_example?create_examples sends the whole list in one request, while a loop pays an HTTP round trip per row. Seeding a few hundred examples one at a time is noticeably slow and much more likely to fail partway, leaving a half-populated dataset. Use the single-row call for the one-off case of promoting an individual trace.
saying these in an interview costs you the question
- Thinking a dataset stores prompts or model settings
- Storing the rendered prompt instead of the input variables
- Assuming outputs is mandatory on every example
- Expecting the SDK to validate input key names
- Believing a dataset can only be seeded by hand