In DeepEval, what does assert_test() do inside a pytest test function?
answer
- the pytest-shaped entry point
- score compared against a threshold
- raises on a low score
- AssertionError message carries score and reason
- metrics run concurrently by default
basics
~10 sassert_test(test_case=..., metrics=[...]) scores the test case with every metric and raises AssertionError if any metric lands below its threshold. That turns an LLM evaluation into an ordinary pytest failure.
solid answer
~40 s`assert_test` is DeepEval's bridge between an evaluation and a test runner. You build an `LLMTestCase` (input, `actual_output`, and whatever else the metrics need, such as `retrieval_context`), pass it with a list of metric objects, and DeepEval runs each metric against that case. At least one metric must carry a threshold. If every metric passes, the function returns quietly; if any metric scores below its threshold — or errors — it raises an `AssertionError` whose message lists the metric name, score, threshold, and the metric's reason. In deepeval 4.x it runs the metrics concurrently by default (`run_async=True`), and there is a second shape, `assert_test(golden=..., metrics=[...])`, that scores the trace your instrumented app produced during the test instead of a hand-built test case.
code
python · 13 linesfrom deepeval import assert_test
from deepeval.metrics import AnswerRelevancyMetric
from deepeval.test_case import LLMTestCase
def my_app(question: str) -> str:
return "DeepEval is an open-source LLM evaluation framework."
def test_answer_is_relevant():
question = "What is DeepEval?"
test_case = LLMTestCase(input=question, actual_output=my_app(question))
assert_test(test_case=test_case, metrics=[AnswerRelevancyMetric(threshold=0.7)])go deeper
Know the shape: build an LLMTestCase, pass it with a list of metrics, and a score under the threshold raises AssertionError. Say plainly that this is what makes an eval run as a pytest test.
Explain that at least one metric must have a threshold, that metrics run concurrently by default, and that an errored metric fails the assertion unless errors are explicitly ignored.
Show that you have read a real failure message — name, score, threshold, reason — and talk about how you use the reason field to tell a genuine regression from a mis-specified metric in a CI log.
Own the framing that a non-deterministic assertion belongs in a gate only when the threshold and the blast radius are chosen deliberately, and be ready to argue what an eval assertion should be allowed to block.
## The problem it solves An LLM evaluation produces a score between 0 and 1, not a boolean. A test runner wants a boolean. `assert_test` is the adapter: it runs metrics, compares each score to that metric's threshold, and converts the outcome into a pass or a raised `AssertionError`. That single conversion is what lets an eval suite live in the same place as your unit tests and block a merge the same way they do. ## The signature In deepeval 4.1.9: ``` assert_test(test_case=None, metrics=None, golden=None, run_async=True) ``` Two valid shapes exist. - **Test-case shape** — `assert_test(test_case=<LLMTestCase or ConversationalTestCase>, metrics=[...])`. You supply the inputs and the output you already generated. Passing a test case without metrics raises `ValueError`; metric types must match the case type (`BaseMetric` for `LLMTestCase`, `BaseConversationalMetric` for `ConversationalTestCase`). - **Trace shape** — `assert_test(golden=<Golden>, metrics=[...])`. Here you call your own instrumented application inside the test body, and DeepEval scores the trace the pytest plugin wrapped around that test. This shape only works under the `deepeval test run` CLI, because that is what activates the plugin's evaluation scope. DeepEval also enforces `check_at_least_one_metric_has_threshold` on the test-case shape: a list of metrics where none has a threshold has nothing to assert on, so it is rejected rather than silently passing. ## What a failure looks like When at least one metric fails, DeepEval collects the failing metric data and raises one `AssertionError` whose message is a comma-joined list of `name (score: …, threshold: …, strict: …, error: …, reason: …)`. Two things are worth internalising here. First, an *errored* metric — a judge call that timed out or returned unparsable JSON — counts as a failure unless you asked DeepEval to ignore errors. Second, `reason` is the judge model's own natural-language justification, and it is the single most useful field in a CI log, because it tells you whether the output was genuinely bad or the metric was mis-specified. A per-case escape hatch exists: `LLMTestCase(..., flaky=True)`. For a flaky case, a failing metric emits a Python warning (which shows up in pytest's warnings summary) instead of raising, so a known-unstable case can stay visible without blocking the pipeline. ## Concurrency and cost `run_async` defaults to `True`, which means the metrics attached to that one test case are measured concurrently rather than one after another. This matters because almost every DeepEval metric is itself an LLM call — often several per metric. Three metrics on one test case is not "one assertion"; it is a handful of judge requests with real latency and a real bill. Concurrency across *different* test cases is a separate concern, handled by running the suite with multiple processes. Because metrics call a model, an `assert_test` is not deterministic in the way `assert x == y` is. The same test case can score 0.71 on one run and 0.68 on the next. That is a property of the judge, not a bug in the assertion, and it is why thresholds, repeats, and the `flaky` flag exist. ## Where it does and does not run `assert_test` is a plain Python function, so it executes under bare `pytest` as well as under `deepeval test run`. Under bare pytest the metrics still run and a failure still fails the test — what you lose is the aggregated test-run report, the result cache, and the upload to Confident AI, all of which the CLI turns on. ## A minimal example A test body typically does three things: call the system under test, wrap the input and output in an `LLMTestCase`, then assert. Any per-case setup that is not evaluation-specific — fixtures, temporary directories, fake clients — is ordinary pytest and behaves exactly as it always does; `assert_test` neither knows nor cares about it. ## Common mistakes Treating the return value as a score is the first one: `assert_test` returns `None`, and the scores live in the raised message or in the run report. Building the test case with a field the metric needs but you did not populate is the second — a metric that requires `retrieval_context` on a case that has none will error rather than quietly score zero. And forgetting that each assertion costs money is the third; a thousand-case suite on every push is a budget decision, not a testing decision.
- What does assert_test return when everything passes, and where do the actual scores go?It returns `None`. On success nothing is printed as a failure; the per-metric scores and reasons live in the test run DeepEval assembles, which the `deepeval test run` CLI prints at the end and, when a Confident AI key is configured, uploads. If you want the scores in hand as Python objects, use `evaluate()` instead, which returns an `EvaluationResult`.
- What happens if you pass a list of metrics where none of them has a threshold?DeepEval rejects it. `assert_test` calls an internal check that at least one metric carries a threshold, because otherwise there is no criterion to assert against and the test would pass unconditionally — the worst failure mode for a gate. Set a threshold on the metrics you intend to gate on.
- How does the assert_test(golden=..., metrics=...) shape differ from passing a test case?With `golden=`, you do not build the output yourself. You invoke your instrumented application inside the test, and DeepEval scores the trace the pytest plugin captured around that test. It only works under `deepeval test run`, since that is what turns the plugin's evaluation scope on, and it is how you evaluate an agent end to end rather than a single prompt-response pair.
saying these in an interview costs you the question
- Thinks assert_test returns a score object you inspect
- Assumes it is deterministic like a normal equality assertion
- Believes a metric error silently passes rather than failing
- Forgets every metric in the list is a paid LLM call
- Thinks assert_test only works under the deepeval CLI