Why run `deepeval test run` instead of plain pytest for a DeepEval suite?
answer
- a wrapper, not a separate runner
- it calls pytest.main under the hood
- flags pytest does not have
- the run object is the added value
- pytest's exit code is propagated
basics
~20 sdeepeval test run invokes pytest with DeepEval's plugin active and its own flags on top: parallel processes, repeats, result caching, a run identifier, and an aggregated end-of-run report that can be pushed to Confident AI. Plain pytest still runs the assertions but produces none of that.
solid answer
~50 sThe CLI is a thin wrapper around `pytest.main()`. It marks the process as a DeepEval run, forwards the file or directory you named, and adds flags pytest does not have: `-n` for parallel processes, `-r` to repeat each test, `-c` to reuse cached metric results, `-i` to ignore metric errors, `-m` for pytest marks, `-d` to control which results are displayed, and `-id` to label the run. Unrecognised arguments are passed straight through to pytest. At the end it assembles all test results into one test run, prints the summary, and — when a Confident AI key is set — uploads it. Under bare pytest the metrics still execute and failures still fail, but there is no aggregated run, no cache writing, and no upload. Failures propagate as pytest's exit code, so CI goes red.
code
bash · 5 lines# four worker processes, reuse cached metric results, label the run
deepeval test run tests/evals -n 4 -c -id "pr-1421"
# probe metric variance: three passes per test (caching is disabled by -r)
deepeval test run tests/evals/test_rag.py -r 3 -d failinggo deeper
Remember the command shape deepeval test run <path> and that it runs your pytest tests with DeepEval's reporting turned on. Knowing -n for parallelism is enough at this level.
Explain that the CLI calls pytest.main with the plugin active, name what the extra flags map to, and say what you lose under bare pytest — the aggregated run, caching, and the upload.
Talk about how you invoke it in a pipeline: which flags you set, how you read the exit code, and how you use a run identifier so a red build maps to a specific reviewable run.
Own the decision of what the pipeline gates on versus merely records, and whether the team's eval runs should live in a hosted run store at all given what the prompts and outputs contain.
## What the command actually is `deepeval test run <file_or_dir>` is a Typer command that ends up calling `pytest.main()` in the same process. It is not a separate test engine. That is the first thing to say in an interview, because it explains everything else: fixtures, conftest files, markers, assertion rewriting and plugins all behave exactly as they do under pytest, and the DeepEval-specific behaviour is a plugin plus a handful of extra flags. DeepEval registers its plugin through the `pytest11` entry point, so it auto-loads in any environment where deepeval is installed. The CLI additionally sets an internal "running deepeval" flag before starting pytest. The plugin checks that flag, and only when it is set does it create a test run, keep results on disk, and wrap each test in an evaluation scope. ## The flags, and what each one maps to - `-n / --num-processes` — forwarded to pytest as `-n`, i.e. pytest-xdist, which ships as a deepeval dependency. This is the main lever for wall-clock time, since eval tests are almost entirely waiting on judge model calls. - `-r / --repeat` — forwarded as `--count` (pytest-repeat). Each test is executed N times, which is how you probe a non-deterministic metric's variance. - `-c / --use-cache` — reuse previously computed metric results instead of re-calling the judge. Note the interaction: caching is disabled whenever `-r` is set, because repeating a test that returns a cached score measures nothing. - `-i / --ignore-errors` — a metric that *errored* (timeout, unparsable judge response) no longer fails the assertion. A metric that scored below threshold still does. - `-s / --skip-on-missing-params` — skip cases missing a field a metric requires rather than erroring on them. - `-m / --mark` — forwarded as pytest's `-m`, so `-m smoke` runs a subset. - `-d / --display` — controls whether the final report shows all cases or only the passing/failing ones. - `-id / --identifier` — labels the run so you can find it later among many. - `-o / --official` — marks the run as the baseline on Confident AI; it is ignored with a warning if `CONFIDENT_API_KEY` is not set. - `-x` stops on first failure, `-v` turns on verbose metric output, `--pdb` drops into the debugger. Anything the command does not recognise is appended to the pytest argument list, so `--maxfail=3` or a `::test_name` selector still work. ## The aggregated test run The most valuable thing the CLI adds is the *run* as a first-class object. Individual assertions know only about themselves; the run knows every test case, every metric score, the durations, and the hyperparameters you logged. That is what produces the end-of-run table, and it is what gets uploaded to Confident AI when a key is present, giving you a shareable link and a run-to-run comparison. Under xdist the plugin takes care to account for the session once rather than once per worker, so a `-n 8` run still reports as one run. ## Exit codes The command propagates pytest's return code by exiting with it, which is exactly what a pipeline needs. Remember what pytest's codes mean: 1 means tests failed, 2 means the session was interrupted (which is also how a halted xdist run reports), 3 an internal error, 4 a usage error, and 5 means *no tests were collected*. That last one is the quiet trap — a bad path or a mark that matched nothing exits non-zero, so a misconfigured invocation shows up as a red build rather than a falsely green one. ## When plain pytest is the right call If your evals are a handful of assertions that you want to run alongside unit tests, `pytest` is fine: the assertions work and a failure is a failure. You reach for the CLI when you want the run-level artefacts — parallelism, caching, repeats, a labelled run, the uploaded report. Many teams run both: bare `pytest` locally for a fast single test, `deepeval test run -n 4` in the pipeline. ## Common mistakes Assuming the CLI replaces pytest and therefore that fixtures or conftest will not work. Assuming `-c` and `-r` compose, when caching is switched off in the presence of repeats. Expecting the run to appear on Confident AI without a key configured. And expecting `-i` to make failing metrics pass — it only neutralises metrics that errored.
- If the plugin auto-loads via the pytest11 entry point, what is actually different when the CLI starts the session?The CLI sets an internal flag before pytest starts. The plugin reads it and only then creates a test run, saves results to disk, and wraps each test in an evaluation scope so trace-scoped assertions work. Without the flag the plugin stays quiet, which is deliberate: deepeval being installed should not make unrelated pytest suites emit evaluation runs.
- Your pipeline runs `deepeval test run tests/evals -m nightly` and the job goes red with no failures listed. What is the likely cause?The mark probably matched nothing, so pytest collected zero tests and exited with code 5, which the command propagates. It looks like a failure but is a selection problem. Check the mark is registered and spelled as expected, and treat exit code 5 as a distinct diagnosis from exit code 1.
- Can you pass ordinary pytest arguments through the CLI?Yes. The command allows extra and unknown options and appends them to the pytest argument list, so things like `--maxfail=2`, a `file.py::test_name` selector, or an xdist distribution mode reach pytest unchanged. Only the DeepEval-specific short flags are intercepted first.
saying these in an interview costs you the question
- Thinks the CLI is a separate test runner that replaces pytest
- Says fixtures and conftest do not work under the CLI
- Believes -c caching still applies when -r repeat is set
- Assumes results reach Confident AI without an API key
- Reads exit code 5 (nothing collected) as a test failure