skip to content

In Langfuse, how do you build a dataset and run your app over it as a dataset run?

level: middleimportance: must knowfreq 58%

answer

  1. stored inputs you can replay
  2. expected_output only exists offline
  3. curate items straight from bad traces
  4. one trace per item, linked to a run
  5. task plus evaluators returning Evaluation

basics

~20 s

Create the dataset with create_dataset(name=...) and fill it using create_dataset_item(dataset_name=..., input=..., expected_output=...), optionally pointing source_trace_id at the production trace an item came from. Then fetch it with get_dataset(name) and execute a run — dataset.run_experiment(name=..., task=..., evaluators=[...]) — which traces every item and attaches its scores.

solid answer

~50 s

A Langfuse dataset is a stored list of items, each with an `input`, an optional `expected_output` and metadata. You create it with `langfuse.create_dataset(name=...)` and add items with `langfuse.create_dataset_item(dataset_name=..., input=..., expected_output=...)`; passing `source_trace_id` records which production trace an item was curated from, which is how a real failure becomes a permanent test case. To execute, fetch it with `langfuse.get_dataset(name)` and call `dataset.run_experiment(name=..., task=..., evaluators=[...])`. Your `task` receives each item and returns the output; each evaluator receives `input`, `output` and `expected_output` and returns `Evaluation(name=..., value=...)` objects, which are stored as scores. The call returns an `ExperimentResult` with `dataset_run_id`, `dataset_run_url` and per-item results, and `max_concurrency` bounds parallelism. The lower-level alternative is looping over `dataset.items` and using `with item.run(run_name=...) as root_span:`, which creates one trace per item linked to the named run so you can score it yourself with `root_span.score_trace(...)`.

code

python · 26 lines
python
from langfuse import get_client, Evaluation

langfuse = get_client()

langfuse.create_dataset(name="capitals")
langfuse.create_dataset_item(
    dataset_name="capitals",
    input={"country": "Italy"},
    expected_output="Rome",
)

def my_app(country: str) -> str:
    return "Rome"

def task(*, item, **kwargs):
    return my_app(item.input["country"])

def exact_match(*, input, output, expected_output=None, **kwargs):
    return Evaluation(
        name="exact_match",
        value=1.0 if output == expected_output else 0.0,
    )

dataset = langfuse.get_dataset("capitals")
result = dataset.run_experiment(name="baseline", task=task, evaluators=[exact_match])
print(result.format())

go deeper

for a junior

Know that a Langfuse dataset stores inputs (and expected outputs) you can replay, and that running your app over it produces one traced result per item.

for a middle

Be able to write it: create_dataset, create_dataset_item, get_dataset, then run_experiment with a task and evaluators — or the item.run() loop — and say where the scores end up.

for a senior

Show the discipline around a run: curate items from real traces, change one variable per run, record it in run metadata, and bound concurrency and judge cost.

for a principal

Own dataset size and refresh policy against budget — what runs on every change, what runs nightly, and how curated failures keep entering the set without breaking comparability.

## Datasets: the offline half Production traces tell you what happened once. A **dataset** lets you re-run the same inputs whenever you like, which is the only way to attribute a change in quality to a change you made. Each dataset item holds an `input` (any JSON — a string, a dict of arguments), an optional `expected_output` for the cases where ground truth exists, and free metadata. Items are created with `langfuse.create_dataset_item(dataset_name=..., input=..., expected_output=..., metadata=...)`, and the client also accepts `source_trace_id` and `source_observation_id`. Those two fields matter more than they look: they turn 'this real user request went wrong' into a dataset item that still links back to the original trace, so the reason it was added is never lost. Supplying your own `id` upserts, so a curation script can be re-run without duplicating items. ## Runs: executing the dataset Fetching a dataset gives you a client object with `.items`, and two ways to run it. The **experiment API** is the high-level path. `dataset.run_experiment(name=..., task=..., evaluators=[...], run_evaluators=[...], max_concurrency=50)` walks the items, calls your `task(*, item, **kwargs)` for each, runs the evaluators on the result, and creates the dataset run in Langfuse with everything linked. Item-level evaluators take `input`, `output` and `expected_output` and return one or more `Evaluation(name=..., value=..., comment=...)` objects; `run_evaluators` compute aggregates over the whole run, which is where a pass rate or a mean belongs. The returned `ExperimentResult` carries `dataset_run_id`, `dataset_run_url`, the per-item results and a `format()` helper for printing a summary in CI logs. `max_concurrency` defaults to 50 — that is the knob you turn down when a provider starts returning rate-limit errors. The **manual loop** is the lower-level path and is worth knowing because it composes with any application shape: 1. `for item in dataset.items:` 2. `with item.run(run_name='baseline-v2') as root_span:` — this opens a trace for that item and registers it as part of the named run. 3. Call your application inside the block. 4. Attach the trace-level input and output with `root_span.set_trace_io(...)` and score it with `root_span.score_trace(name=..., value=...)`. Note the v4 spelling: the v3 SDK's `span.update_trace(...)` no longer exists, and trace-level input/output is set with `set_trace_io`. `item.run()` also takes `run_metadata` and `run_description`, which is where you record what this run *is* — which prompt version, which model, which retriever — so a run is self-describing months later. ## What you get back Because every item's execution is a real trace, a dataset run is not just a table of numbers. Langfuse shows the run's aggregate scores next to other runs of the same dataset, and each cell drills into the full trace for that item: the prompt, the retrieval, the tool calls, the tokens and cost. Regressions are therefore diagnosable, not merely detectable. ## Practical cautions Runs are comparable only if the data is held still. If someone adds twenty items between two runs, the second run's average moved partly because the dataset moved; compare per item, or freeze the dataset while a comparison is in flight. Change one variable per run and record it in the run name and metadata. If your evaluators call a judge model, the run costs one or more model calls per item on top of the application's own calls, so a 2,000-item dataset with three judge metrics is a real bill and a long wall-clock time — which is why `max_concurrency` and dataset size are decisions, not defaults. And judgements that need `expected_output` only work here, offline; they have nothing to compare against on live traffic.

  • What does passing source_trace_id when creating a dataset item buy you?
    It records which production trace the item was curated from, so the item keeps a link back to the real failure that motivated it. Months later you can still see the original execution rather than guessing why an input is in the set. It also makes 'promote this bad trace into the regression set' a one-call workflow instead of copy-paste.
  • Two runs of the same Langfuse dataset differ by 4 points, but someone added items in between. Is the comparison valid?
    Not as an aggregate. The average moved partly because the population moved, so a run-level delta mixes your change with the data change. Compare item by item over the intersection, or freeze the dataset while a comparison is in flight and re-run the baseline after any curation. Recording dataset changes in run metadata makes this detectable rather than invisible.
  • Your dataset run keeps hitting provider rate limits. What do you change?
    Lower max_concurrency on run_experiment — it defaults to 50 parallel item executions, which is more than most provider quotas tolerate once each item also triggers judge calls. If that is not enough, split the dataset, run judge metrics in a second pass, or move the run off the per-pull-request path onto a nightly schedule.

saying these in an interview costs you the question

  • Thinks a dataset run is just a loop with no traces
  • Compares two runs after adding items to the dataset
  • Forgets judge evaluators cost a model call per item
  • Uses the removed v3 span.update_trace() to set trace input and output
  • Puts expected-output metrics on live production traffic

context