skip to content

How do you run a LangSmith experiment over only one split of a dataset?

level: middleimportance: should knowfreq 40%

answer

  1. labels on examples, one dataset
  2. no split argument on evaluate
  3. pass examples, not a dataset name
  4. list_examples takes splits
  5. two tiers: fast subset, full set

basics

~10 s

Splits are labels attached to examples inside a single dataset. Fetch just the labelled ones with client.list_examples(dataset_name=..., splits=["test"]) and pass that as the data argument of evaluate(), since data accepts any iterable of examples.

solid answer

~40 s

A split in LangSmith is not a separate dataset — it is a **label on examples within one dataset**, so a row can sit in a `smoke` split while still being part of the whole set. You assign splits from the dataset UI or with the SDK's example-update call (`client.update_examples(..., splits=[...])`). To run against one, you don't pass a special argument to `evaluate`: you pass **examples** instead of a dataset name, because `data` accepts any iterable of examples. So `evaluate(target, data=client.list_examples(dataset_name="support-qa", splits=["smoke"]), experiment_prefix="pr-smoke")` runs only the labelled rows. Because everything still lives in one dataset, experiments over different splits appear under that dataset and remain individually inspectable — but only experiments over the *same* set of examples are meaningfully comparable, so keep a split's membership stable if you are tracking it over time.

code

python · 14 lines
python
from langsmith import Client, evaluate

client = Client()


def target(inputs: dict) -> dict:
    return {"answer": "stub"}


evaluate(
    target,
    data=client.list_examples(dataset_name="support-qa", splits=["smoke"]),
    experiment_prefix="pr-1421-smoke",
)

go deeper

for a junior

Know that a split is a label on examples inside one dataset, and that you narrow a run by passing filtered examples as data rather than by a special argument.

for a middle

Be able to write the call: client.list_examples(dataset_name=..., splits=[...]) handed to evaluate() as data, and explain that data accepts any iterable of examples.

for a senior

Explain the two-tier pattern you actually run and its traps — cross-split averages are not comparable, and quietly changing split membership invalidates the trend you were watching.

for a principal

Own membership as a controlled artifact: who may add to a gated split, when the previous configuration is re-baselined after the set grows, and how a held-out split keeps the team from tuning to its own test set.

## Splits are labels, not datasets The mental model that trips people up is the machine-learning one, where train/validation/test are three separate files. In LangSmith a split is a **label on an example inside one dataset**. The dataset remains the unit of identity — the thing experiments attach to, the thing that gets versioned — and splits are a way to slice it. That has a pleasant consequence: an example can be in a split and still be part of the full set, so a small smoke split is a *view* on the big dataset, not a copy of some of its rows. When you fix a reference answer, you fix it once and both views see the fix. Copying rows into a second dataset gives you two references that drift apart. ## Assigning a split The usual path is the UI: select examples in the dataset view and put them in a split. In the SDK, splits are set through the example-update call, `client.update_examples(...)` with the `splits` argument, which is how you script the assignment from a rule — for instance, labelling every example whose metadata says it came from an incident. ## Running an experiment on one There is no `split=` parameter on `evaluate`. The mechanism is the one `data` already offers: it accepts a dataset name, a dataset id, **or an iterable of examples**. So you fetch the subset and hand it over: ``` evaluate( target, data=client.list_examples(dataset_name="support-qa", splits=["smoke"]), experiment_prefix="pr-1421-smoke", ) ``` `list_examples` returns an iterator over the matching examples, and the same trick generalises: any filtered listing you can express — by split, by metadata, by version — becomes the population of an experiment. ## Why teams reach for this The common shape is two tiers. A small, fast split runs often, because a suite that takes twenty minutes and eighty dollars will not survive contact with a pull-request workflow. A larger set runs less often, on a schedule or before a release. Splits let both tiers read from one curated source of truth rather than two datasets that slowly diverge. Other uses are just as ordinary: a split per failure category so you can see whether the change that helped summarisation hurt refusals; a split of examples with hand-verified references, separated from ones auto-promoted from traces; a split of the newest examples that no prompt has been tuned against yet, kept as a held-out set so you can tell tuning from overfitting. ## What splits do not do They do not make cross-split experiments comparable. An experiment over 40 smoke examples and an experiment over 800 full examples produce two averages that mean different things; putting them side by side is a category error even though both live under the same dataset. Compare like with like, and label the experiment prefix with the split so nobody misreads the list. They also do not protect you from membership churn. If someone adds twelve rows to the smoke split on Tuesday, Wednesday's number is not measuring the same thing as Monday's. If a split is a tracked gate, treat its membership as change-controlled, and when you do extend it, re-run the previous configuration so you have a fresh baseline on the new membership rather than comparing across a moved goalpost. Finally, splits do not by themselves reduce cost when you are careless with what you call: a 40-example smoke run with three LLM-judge evaluators is still 160 model calls. The split shrinks the row count; the per-row call count is your evaluator choice. ## Naming discipline Split names show up in filters and in nothing else — they carry no semantics for the tool. So they are worth naming for humans: `smoke`, `held-out`, `refusals`, `verified` all say what they are. Names like `split-2` mean that a year from now nobody can tell whether it is safe to run against.

  • Why keep splits inside one dataset instead of creating a second dataset for the fast tier?
    Because references stay single-sourced. With splits, correcting an expected answer fixes it for every view; with a copied dataset, the two copies drift and you eventually gate on a stale reference. One dataset also keeps every experiment under one roof and one version history, so you can still see the whole picture.
  • Can you compare an experiment on a smoke split against one on the full dataset?
    Not meaningfully. The two averages are computed over different populations, so a difference between them says more about which examples were included than about the change under test. Compare experiments that ran over the same examples, and put the split name in the experiment prefix so the list is not misread.
  • How would you build a held-out split, and what does it buy you?
    Label a portion of examples — ideally recent ones — as held out and never tune a prompt while looking at them. Run the frequent loop against the rest and check the held-out split occasionally. When the tuned split improves and the held-out one does not, you have been fitting the prompt to specific examples rather than improving the system.

saying these in an interview costs you the question

  • Believing a split is a separate dataset
  • Looking for a split= argument on evaluate()
  • Comparing a smoke-split average against a full-set average
  • Copying rows into a second dataset to make a subset
  • Changing split membership silently between gated runs

context