How do you run a DeepEval suite over a whole dataset as one pytest test per case?
answer
- one pytest node per case
- parametrize over the dataset's goldens
- generate the output inside the body
- a loop cannot be distributed
- readable ids make the CI log usable
basics
~20 sParametrize the test over the dataset's goldens — @pytest.mark.parametrize("golden", dataset.goldens) — and build one LLMTestCase inside the test body. Each golden then becomes its own pytest node, so failures are isolated and the cases can be distributed across processes.
solid answer
~50 sThe idiomatic DeepEval pattern is one test function parametrized over the dataset's goldens. Inside the body you call the system under test with `golden.input`, wrap the result in an `LLMTestCase`, and call `assert_test`. The alternative — a single test that loops over every case — technically works but is worse in three ways: the first failure aborts the loop so you never see the rest, the whole dataset collapses into one red node with no per-case identity, and pytest cannot distribute the cases across workers because there is only one node to schedule. With parametrization each case gets its own node id, so `deepeval test run -n 8` genuinely parallelises the judge calls, a rerun can target one failing case, and the end-of-run report lines up one row per golden. The mechanics of parametrize itself are ordinary pytest; what is DeepEval-specific is what you put in the parameter list and what you build inside the body.
code
python · 24 linesimport pytest
from deepeval import assert_test
from deepeval.dataset import EvaluationDataset, Golden
from deepeval.metrics import AnswerRelevancyMetric
from deepeval.test_case import LLMTestCase
dataset = EvaluationDataset(
goldens=[
Golden(input="What is DeepEval?"),
Golden(input="How do I set a metric threshold?"),
]
)
def my_app(question: str) -> str:
return f"An answer to: {question}"
@pytest.mark.parametrize(
"golden", dataset.goldens, ids=[g.input[:40] for g in dataset.goldens]
)
def test_relevancy(golden: Golden):
test_case = LLMTestCase(input=golden.input, actual_output=my_app(golden.input))
assert_test(test_case=test_case, metrics=[AnswerRelevancyMetric(threshold=0.7)])go deeper
Know that each dataset case should be its own test, and that pytest's parametrize decorator is how you get there. Be able to point at where the model call belongs: inside the test body.
Explain the three concrete wins over a loop — failure isolation, distribution across workers, and per-case reporting — and why the output must be generated at run time rather than at collection time.
Discuss how you keep the suite affordable at dataset scale: marked subsets per pull request, the full sweep on a schedule, readable ids so a red node maps straight to an input.
Own where the dataset lives and who curates it, and be ready to argue how many cases genuinely need to run on every change versus what belongs in a periodic sweep.
## The shape A DeepEval dataset holds *goldens*: the inputs (and, where you have them, expected outputs and context) you want to evaluate against, without the generated output. The output does not exist until you run your application, which is why the generation step belongs inside the test body rather than in the dataset. So the pattern is: 1. Load or build the dataset once at module level. 2. Parametrize the test function over `dataset.goldens`. 3. In the body, call your app with `golden.input` to get `actual_output`. 4. Assemble an `LLMTestCase` from the golden's fields plus that output. 5. `assert_test(test_case=..., metrics=[...])`. Because the parametrize argument list is evaluated at collection time, the dataset must be available when the module is imported — a pull from a remote store or a file read at module scope, not inside a fixture that only runs later. If the dataset genuinely has to be built lazily, generate the ids from something cheap and resolve the heavy object inside the test. ## Why one node per case matters **Failure isolation.** A loop stops at the first `AssertionError`. If case 3 of 200 fails, you learn nothing about cases 4 through 200 and you cannot tell whether you have one regression or forty. Separate nodes each fail independently, and the run report shows exactly which inputs broke. **Parallelism.** pytest-xdist distributes *nodes*. A single looping test is one node, so `-n 8` gives you nothing. Two hundred parametrized nodes spread across eight workers, and since eval tests are dominated by waiting on judge model calls, the speedup is close to linear until you hit provider rate limits. **Selectivity.** With node ids you can rerun a single case (`path::test_fn[golden3]`), and marks can split the dataset into a smoke subset for pull requests and the full set for a nightly run. **Reporting.** Each node's result carries the case's identity into the aggregated test run, which is what makes a run-to-run comparison meaningful. ## Readable ids Raw parametrization over objects produces ids like `golden0`, `golden1`, which are useless in a CI log. Give pytest something human: pass `ids=[g.input[:40] for g in dataset.goldens]`, or parametrize over the inputs and look the golden up. This is plain pytest, but it is the difference between a failure you can triage from the log and one that forces you to open the report. ## The trace-scoped variant DeepEval 4.x offers a second shape for agentic applications: parametrize over goldens as before, call your instrumented application inside the test, then `assert_test(golden=golden, metrics=[...])` with no test case at all. The pytest plugin wraps each test in an evaluation scope, so DeepEval scores the trace your app produced — spans, tool calls, retrievals — rather than a hand-assembled input/output pair. This shape requires running under `deepeval test run`, because that is what activates the scope; under bare pytest there is no trace to read. ## Cost is per node One parametrized suite of 500 goldens with three metrics is 1,500 metric evaluations, each of which is one or more judge calls. Parametrization makes that cost visible and schedulable rather than smaller. The usual mitigations are running a marked subset per pull request, enabling the result cache when the generated outputs have not changed, and keeping the full sweep on a schedule. ## Common mistakes Generating the outputs at collection time — calling the model in the parametrize expression — so that merely collecting the suite costs money and time, and a provider outage breaks collection rather than a test. Sharing mutable state between parametrized cases, which becomes a race the moment you add `-n`. Parametrizing over test cases you built ahead of time when the output should have been generated per run, which silently evaluates stale outputs. And forgetting that a dataset pulled at import time makes the whole module fail to collect when the store is unreachable.
- Why not build all the LLMTestCase objects up front and parametrize over those instead?Because building the test case requires the generated output, so doing it up front means calling your application at collection time. That makes collection slow and network-dependent, and it evaluates outputs produced before the test session's setup ran. Parametrize over goldens and generate inside the body, where a failure is a test failure rather than a collection error.
- How does this pattern interact with `deepeval test run -n 8`?Directly — xdist distributes pytest nodes, so parametrization is what creates work to distribute. Eight workers each run their share of the cases concurrently, which is a near-linear win for judge-bound tests until provider rate limits bite. A single looping test would leave seven workers idle.
- How would you split one dataset into a per-pull-request subset and a full nightly run?Apply pytest marks and select them with the runner's `-m` flag: mark a small representative slice as smoke, run `-m smoke` on pull requests, and run the whole file on a schedule. Alternatively parametrize over a filtered list of goldens chosen by a tag on the golden itself, so the split lives in the dataset rather than in the test code.
saying these in an interview costs you the question
- Loops over the dataset inside one test function
- Calls the model at collection time to build test cases
- Expects -n to speed up a single looping test
- Leaves ids as golden0, golden1 in CI logs
- Shares mutable state across parametrized cases