Teammates added 40 examples to a LangSmith dataset between two experiments — are the runs comparable?
answer
- averages over different populations
- every write makes a version
- tag a version, pin the run
- move the pin, re-measure the incumbent
- comparison lines up shared examples
basics
~20 sNot directly — the two experiments ran over different populations, so part of the score difference is the new examples, not the change under test. LangSmith versions datasets on every write, so pin gated runs to a tagged version and re-baseline the old configuration whenever the set grows.
solid answer
~50 sNo, and the mistake is common enough to be worth designing against. LangSmith **versions a dataset on every write**, so the two experiments are attached to different states of the set, and their averages are computed over different populations — a score that moved may simply reflect that the 40 new examples are harder or easier than the old ones. Two mechanisms fix it. First, pin the run: `client.update_dataset_tag(dataset_name=..., as_of=..., tag="gate")` names a version, and `data=client.list_examples(dataset_name=..., as_of="gate")` runs an experiment against exactly that snapshot, so a gate is stable even while people keep curating. Second, when you deliberately move the gate to a newer version, **re-run the previous configuration on it** so the baseline is measured on the same population as the challenger. Beyond that it is bookkeeping discipline: put what changed in `experiment_prefix` and record the model, prompt version and dataset version in `metadata`, because the comparison view can line experiments up example by example only when they share those examples.
code
python · 23 linesfrom datetime import datetime, timezone
from langsmith import Client, evaluate
client = Client()
def target(inputs: dict) -> dict:
return {"answer": "stub"}
client.update_dataset_tag(
dataset_name="support-qa",
as_of=datetime(2026, 8, 1, tzinfo=timezone.utc),
tag="gate",
)
evaluate(
target,
data=client.list_examples(dataset_name="support-qa", as_of="gate"),
experiment_prefix="prompt-v7-on-gate",
metadata={"dataset_version": "gate", "prompt": "v7"},
)go deeper
Know that an eval score is an average over whichever examples ran, so adding examples changes the number on its own — two runs over different sets are not a fair comparison.
Be able to name the mechanism: LangSmith versions a dataset on every write, tags name a version, and passing examples fetched as of that tag pins an experiment's population.
Demonstrate the operational rule — when the pinned version advances, re-run the shipping configuration to re-baseline — and record model, prompt version and dataset tag in metadata so the history stays interpretable.
Own it as policy: who may move a gate, on what cadence, whether references are human-reviewed before they can fail a build, and what user text you are permitted to retain in a durable dataset at all.
## Why this is not pedantry An eval score is an average over a population. Change the population and the average changes, entirely independently of the system under test. Adding 40 examples — very often failure cases somebody just harvested from production, so systematically *harder* than the existing set — will push the score down, and the prompt change being tested takes the blame. The reverse happens too: a batch of easy examples makes a mediocre change look like a win. Neither is a subtle statistical effect; it is arithmetic. ## What the platform gives you LangSmith treats a dataset as versioned: every write — new examples, edited references, deletions — produces a new version, and older versions remain addressable. Two SDK surfaces matter: - `client.update_dataset_tag(dataset_name=..., as_of=..., tag="gate")` puts a human-readable label on the version as of a point in time. - `client.list_examples(dataset_name=..., as_of="gate")` returns the examples as of that tagged version, and since `evaluate`'s `data` accepts an iterable of examples, that listing becomes the experiment's population. The result is a gate that does not move under you. People can keep adding examples all week; the gated experiment keeps measuring the snapshot you pinned. ## The re-baseline rule Pinning forever is not the goal — the dataset should grow, or it stops reflecting the product. The rule is: **when you move the pin, re-measure the incumbent.** Retag to the new version, re-run the currently-shipping configuration against it to establish the new baseline, and only then compare challengers. This costs one extra experiment and removes the single most common way an eval programme misleads its own team. Say that out loud in an interview, because it is the answer that distinguishes someone who has run this loop from someone who has read the docs. The failure mode is not "the tool did something odd"; it is a team confidently shipping a change on a comparison whose two sides were never comparable. ## Bookkeeping that makes the history readable Three fields do the work: - `experiment_prefix` — the name humans scan. It should say what varied: `prompt-v7-on-gate`, `gpt-4o-mini-on-gate`. A prefix that repeats across genuinely different configurations makes the list useless. - `metadata` — the machine-readable record: model, temperature, prompt version, git sha, dataset version or tag. This is what you filter and group by when someone asks in three months why the score moved. - the dataset version itself — implied by what you passed as `data`, which is precisely why the tag should also appear in the metadata rather than living only in someone's memory. ## The comparison view Comparing experiments is only meaningful over shared examples: the view lines runs up example by example, which is what lets you see *which rows* moved rather than only that the average did. Rows present in one experiment and absent from the other cannot participate in that alignment. This is the visible consequence of the population problem — if your two experiments barely overlap, the side-by-side view is mostly gaps. Example-level comparison is also the most valuable output of the whole exercise. An aggregate that moved by two points tells you nothing about what to do; five specific rows that flipped from pass to fail tell you exactly where to look. ## Organisational shape At scale this becomes policy, not technique: - **Who may write to a gated dataset?** If anyone can, the gate drifts continuously. Common answer: curation happens freely on the dataset, but the gate tag moves only through a deliberate action, ideally reviewed. - **How often does the pin move?** Tie it to a cadence or a release rather than to whoever last added examples. - **Are references reviewed?** A wrong reference is worse than a missing example — it makes correct behaviour score as a regression, permanently. - **What is retained?** Datasets are durable storage of real user text; whether you may keep it is a data-governance question that predates any of this tooling. ## What a weak answer sounds like "Just re-run both prompts now, on the current dataset." That is actually a legitimate fallback — re-running the incumbent on today's data restores comparability — but offered without naming the versioning mechanism it reads as luck rather than design, and it does not scale to a gate other people rely on.
- If you cannot pin a version, what is the cheapest way to restore comparability?Re-run the incumbent configuration on the current dataset so both sides of the comparison sit on the same population, then compare challenger against that fresh baseline. It costs one extra experiment and is exactly what the pin would have bought you. Record in metadata which dataset state both runs used, so the pair stays interpretable later.
- Why is a wrong reference answer worse than a missing example?A missing example only means you are not testing something. A wrong reference actively punishes correct behaviour — every configuration that gets that row right scores it as a failure, so the metric now rewards reproducing the mistake. That is why references promoted automatically from traces need human review before a dataset becomes a gate.
- What belongs in an experiment's metadata for the history to stay readable?Whatever varied and whatever might have varied: model name and version, temperature, prompt version, git sha, and the dataset version or tag the run used. The prefix is for scanning by eye; metadata is what you filter and group by months later when somebody asks why the score stepped down in the third week of the quarter.
- Who should be allowed to change a gated dataset, and how often?Curation should stay open — the set must keep absorbing what production teaches you — but the gate itself should move deliberately: a tagged version advanced on a cadence or at a release boundary, with the incumbent re-baselined at the same time. If the gate follows whoever last pasted in examples, nobody can interpret a week-over-week trend.
saying these in an interview costs you the question
- Comparing scores across two dataset versions as-is
- Assuming edits to a dataset leave prior experiments recomputed
- Treating a growing dataset as a stable gate
- Moving the pin without re-baselining the incumbent
- Recording nothing about which dataset version a run used