When do you use DeepEval's evaluate() instead of assert_test() in pytest?
answer
- one gates, the other measures
- returns results instead of raising
- EvaluationResult carries test_results
- owns its own batching and concurrency
- hyperparameters recorded with the run
basics
~20 sUse evaluate() when you want scores back rather than a pass/fail gate. It takes a list of test cases plus metrics, runs them with its own concurrency settings, and returns an EvaluationResult you can inspect — no pytest, no raised assertion.
solid answer
~40 s`assert_test` is a gate: one test case, metrics, and an `AssertionError` when a threshold is missed. `evaluate()` is a batch job: `evaluate(test_cases=[...], metrics=[...])` scores a whole list in one call and returns an `EvaluationResult` carrying `test_results`, a `confident_link`, and a `test_run_id`. Nothing raises, so you decide what the numbers mean. That makes it the right tool in notebooks, in a scoring script, in a scheduled sweep, or anywhere you want to compare two configurations rather than block a merge. It also exposes configuration that `assert_test` fixes internally — `async_config` for concurrency and throttling, `cache_config`, `error_config`, `display_config` — plus a `hyperparameters` dict recorded with the run. The rule of thumb: gate in CI with `assert_test`, measure everywhere else with `evaluate()`.
code
python · 20 linesfrom deepeval import evaluate
from deepeval.evaluate.configs import AsyncConfig, ErrorConfig
from deepeval.metrics import AnswerRelevancyMetric
from deepeval.test_case import LLMTestCase
test_cases = [
LLMTestCase(input="What is DeepEval?", actual_output="An LLM eval framework."),
LLMTestCase(input="What is a threshold?", actual_output="A pass mark."),
]
result = evaluate(
test_cases=test_cases,
metrics=[AnswerRelevancyMetric(threshold=0.7)],
hyperparameters={"model": "gpt-4o-mini", "chunk_size": 500},
async_config=AsyncConfig(max_concurrent=5),
error_config=ErrorConfig(ignore_errors=True),
)
pass_rate = sum(r.success for r in result.test_results) / len(result.test_results)
print(pass_rate, result.confident_link)go deeper
Remember the split: assert_test raises and belongs in a test, evaluate() returns results and belongs in a script or notebook. Do not expect evaluate() to fail anything.
Explain the returned EvaluationResult and the four config dataclasses, and note that evaluate() batches cases itself while assert_test relies on the test runner for cross-case parallelism.
Show judgment about layering: a small strict gate in CI, a broad evaluate() sweep on a schedule, concurrency tuned to the provider's limits rather than left at the default.
Own the policy question of what a score is allowed to block. Argue for aggregate gates over per-case strictness where the metric is noisy, and for keeping breadth off the merge path.
## Two different jobs Both functions run metrics over test cases. The difference is what they do with the result. `assert_test(test_case=..., metrics=[...])` evaluates one case and raises `AssertionError` if any metric misses its threshold. It is designed to sit inside a pytest test, where raising is the contract. `evaluate(test_cases=[...], metrics=[...])` evaluates a list and returns an `EvaluationResult`. In deepeval 4.1.9 that object holds `test_results` (one `TestResult` per case, each with `success` and `metrics_data`), `confident_link` (a URL when the run was uploaded), and `test_run_id`. It never raises for a low score. ## The configuration surface `evaluate()` accepts four dataclass configs that `assert_test` sets for you: - `AsyncConfig(run_async=True, throttle_value=0, max_concurrent=20)` — whether to run asynchronously, how much to throttle, and how many evaluations may be in flight. `max_concurrent` is the knob you turn down when the judge provider starts returning rate-limit errors. - `DisplayConfig` — the progress indicator, whether results are printed, verbose metric output, which results to display, and an optional results folder for a timestamped JSON export of the run. - `CacheConfig(write_cache=True, use_cache=False)` — whether to write and whether to reuse cached metric results. - `ErrorConfig(ignore_errors=False, skip_on_missing_params=False)` — whether an errored metric or a case missing a required field aborts or is tolerated. It also takes `hyperparameters`, a dict of strings, numbers, or `Prompt` objects, recorded alongside the run so two runs can be compared by what changed, and `identifier` to label the run. ## Batching versus per-node parallelism This is the subtler distinction. `assert_test` evaluates one test case at a time; concurrency *across* cases comes from pytest running several nodes at once under `-n`. `evaluate()` owns the batching itself: give it 500 cases and it schedules them against `max_concurrent` without any test runner involved. So the two express parallelism at different layers, and mixing them — calling `evaluate()` on a large list inside a heavily parallel pytest session — multiplies concurrency in a way that reliably trips provider rate limits. ## Where each belongs Use `assert_test` when a human should be blocked from merging. Use `evaluate()` when you want a number: exploring in a notebook, sweeping a parameter, scoring a sample of production traffic on a schedule, or producing a report that a person reads and decides on. A useful hybrid is to run `evaluate()` in a scheduled job and gate CI on a much smaller `assert_test` suite, so the expensive breadth is decoupled from the merge path. You can, of course, use `evaluate()` inside pytest and assert on the aggregate yourself — for instance requiring that at least 90% of cases pass rather than every single one. That is a legitimate pattern when per-case strictness is too brittle, and it is only possible because `evaluate()` hands the results back instead of raising. ## Common mistakes Expecting `evaluate()` to fail a build: it will not, and a pipeline step that calls it and ignores the return value is green no matter how bad the scores are. Passing an enormous list without lowering `max_concurrent` and then blaming the provider for rate limiting. Assuming the run appears in a hosted dashboard without an API key configured — `confident_link` is `None` in that case. And treating the two functions as interchangeable spellings of the same thing when their failure semantics are opposites.
- Can you use evaluate() inside pytest and still fail the build?Yes, but you write the assertion. Call `evaluate()`, then assert on the aggregate — for example that the share of `test_results` with `success` true clears a rate you chose. That is a deliberately different gate from `assert_test`'s per-case strictness, and it is often steadier for a noisy suite.
- Which evaluate() setting do you reach for first when the judge provider starts rate-limiting you?`async_config=AsyncConfig(max_concurrent=...)`, lowered from its default of 20, with `throttle_value` raised if bursts are still too aggressive. Setting `run_async=False` serialises everything, which fixes the errors but is usually slower than necessary.
- What is in an EvaluationResult besides the per-case results?A `confident_link` — the URL of the uploaded run, or `None` when no Confident AI key is configured — and a `test_run_id`. Each `TestResult` inside `test_results` carries the case's name, a `success` flag, and `metrics_data` with per-metric score, threshold, reason, and any error.
saying these in an interview costs you the question
- Expects evaluate() to fail a build on low scores
- Thinks evaluate() is just assert_test for a list
- Ignores the returned EvaluationResult entirely
- Runs a large evaluate() inside a parallel pytest session
- Assumes a confident_link exists without an API key