skip to content

DeepEval

An evaluation framework that deliberately looks like pytest: test cases, metric objects, and assertions you can run in CI. It covers relevancy, faithfulness, and hallucination alongside the RAG-specific metrics, plus generating the dataset you score against.

on this pageshow

explore

questions

24

In DeepEval, what is the difference between a Golden and an LLMTestCase?

level: juniorimportance: must knowfreq 72%

answer

  1. one is stored, one is produced
  2. which field only exists after a run
  3. datasets outlive models
  4. actual_output is the dividing line
  5. goldens in, test cases out

basics

~20 s

A Golden holds only the static half of a test — input, expected_output, context — with no actual_output. An LLMTestCase adds the actual_output your application produced on this run, and that is what metrics score.

solid answer

~50 s

A `Golden` is the stored, reusable half of a test: `input`, and optionally `expected_output`, `context` and `retrieval_context`. It deliberately has no meaningful `actual_output`, because that only exists once you run your app. An `EvaluationDataset` holds a list of goldens (`dataset.goldens`) — that list is model-agnostic and survives every model, prompt and retriever change. At evaluation time you iterate the goldens, call your application with `golden.input`, and build an `LLMTestCase` carrying `input`, the freshly produced `actual_output`, and whichever golden fields the metrics need. The split is what lets one curated dataset be replayed against a new model tomorrow: if actual outputs were baked into the dataset, every run would need a new dataset. `Golden` does expose an optional `actual_output` field for pre-computed outputs, but the normal flow leaves it empty and fills it at run time.

code

python · 30 lines
python
from deepeval.dataset import EvaluationDataset, Golden
from deepeval.test_case import LLMTestCase

dataset = EvaluationDataset(
    goldens=[
        Golden(
            input="What is the refund window?",
            expected_output="Thirty days from delivery.",
            context=["Refunds are accepted within 30 days of delivery."],
        )
    ]
)


def my_llm_app(question: str):
    return "Thirty days.", ["Refunds are accepted within 30 days of delivery."]


test_cases = []
for golden in dataset.goldens:
    answer, chunks = my_llm_app(golden.input)
    test_cases.append(
        LLMTestCase(
            input=golden.input,
            actual_output=answer,
            expected_output=golden.expected_output,
            context=golden.context,
            retrieval_context=chunks,
        )
    )

go deeper

for a junior

Be able to say plainly that a Golden is the saved input and reference, and an LLMTestCase adds the actual output your app produced. Know that you loop over dataset.goldens and build test cases at run time.

for a middle

Explain the field-by-field mapping, including which fields a metric requires and why context and retrieval_context are not the same thing. Be ready to write the golden-to-test-case loop from memory.

for a senior

Show the operational payoff: one curated golden set replayed across model, prompt and retriever changes, with actual outputs never written back. Talk about who owns the goldens and how references get vouched for.

for a principal

Own the dataset as a long-lived asset separate from any build. Argue for what belongs in a golden versus in run metadata, and how that choice determines whether year-old results are still comparable to today's.

## Two objects, two lifetimes DeepEval separates the part of a test that you *maintain* from the part that is *produced*. The maintained part is a `Golden`. The produced part is an `LLMTestCase`. A `Golden` is imported from `deepeval.dataset` and carries: - `input` — the user query or task. The only genuinely required field. - `expected_output` — the reference answer, when you have one. - `context` — the ground-truth facts a correct answer should rest on. - `retrieval_context` — chunks a retriever returned, when you are storing them. - `additional_metadata` and `comments` — free-form annotations (source document, owner, why this case exists). - `actual_output` — present on the class but normally left empty. An `LLMTestCase` is imported from `deepeval.test_case` and is what metrics actually consume. Its distinguishing field is `actual_output`: the string your application produced for this input, on this run, with this model and this prompt. ## Why the split exists The value of an evaluation dataset is that it outlives any particular build. You curate fifty hard support questions once; you then replay them against GPT-class model A, model B, a new system prompt, a re-chunked index, and a cheaper retriever. Everything that changes between those runs is the output. Everything that stays is the golden. If actual outputs were stored inside the dataset, the dataset would be a snapshot of one run rather than a test suite, and "re-run the suite on the new model" would mean rebuilding it. The split also makes the dataset shareable with non-engineers: a domain expert can write inputs and expected outputs without ever running the app. A second consequence is that goldens are what the `Synthesizer` produces. `generate_goldens_from_docs` cannot produce an `actual_output` — it has no idea what your application would say — so `Golden` is the only object it *can* return. ## The run loop The idiomatic flow is a loop over `dataset.goldens` that invokes your application and constructs test cases: 1. Pull or build an `EvaluationDataset`. 2. For each golden, call your app with `golden.input`. 3. Construct `LLMTestCase(input=golden.input, actual_output=..., expected_output=golden.expected_output, retrieval_context=...)`. 4. Hand the resulting test cases to your metrics. Note step 3 copies `retrieval_context` from the *live* run, not from the golden, when you are evaluating retrieval quality — the chunks your retriever fetched this time are the thing under test. `context`, by contrast, is ground truth and comes from the golden. ## Which fields a metric needs Different metrics require different fields, and a missing one raises an error rather than silently scoring zero. Reference-free metrics typically need only `input` and `actual_output`; reference-requiring ones also need `expected_output` or `context`. This matters when you plan the dataset: if you intend to run a metric that compares against a reference, the golden must carry that reference, and you must decide at curation time — not at evaluation time — where it comes from. ## Where teams get this wrong The most common mistake is treating the dataset as a store of past outputs — pushing actual outputs back into goldens after a run, which turns the suite into a regression-against-yesterday's-model rather than a test of correctness. The second is filling `context` with whatever the retriever returned, which quietly makes retrieval look perfect: `context` is supposed to be the ground truth a human vouches for, and `retrieval_context` is what the system found. Conflating them removes the very gap the evaluation is meant to measure. The third is assuming a golden must have an `expected_output`. Many useful goldens do not — an open-ended question with no single right answer is still a perfectly good test case for a reference-free metric, and inventing a reference answer just to fill the field creates a false standard that penalizes correct outputs.

  • If a Golden can hold an actual_output field, when would you ever populate it?
    When outputs were produced elsewhere and you are only scoring them — for example a batch you exported from production or from an offline inference job, where the golden is really a record of a completed run. It is also handy for a fixture that pins a known-bad output so you can assert a metric catches it. In the normal flow you leave it empty and fill actual_output on the LLMTestCase at run time.
  • What is the difference between a golden's context and the retrieval_context you attach to the test case?
    context is ground truth: the facts a human says a correct answer should rest on, stored on the golden and stable across runs. retrieval_context is what your retriever actually returned on this run, so it changes whenever the index, embedding model or chunking changes. Retrieval-quality metrics compare the two; filling context from the retriever collapses that comparison and makes retrieval look flawless.
  • Your dataset has 200 goldens and no expected_output on any of them. What can you still evaluate?
    Plenty. Reference-free scoring only needs input and actual_output, so relevancy-style and safety-style checks still run. What you lose is anything that compares against a reference answer or against ground-truth facts. If those matter, the fix is a curation pass to add references to the subset where a single correct answer genuinely exists, rather than generating references and treating them as authoritative.

saying these in an interview costs you the question

  • Says Golden and LLMTestCase are just aliases for each other
  • Stores each run's outputs back into the dataset as goldens
  • Fills the golden's context from the retriever's own output
  • Claims every golden must carry an expected_output
  • Thinks a dataset must be regenerated for each new model

context

open as a page

What does evaluation_params control in a DeepEval GEval metric?

level: juniorimportance: must knowfreq 65%

basics

~20 s

evaluation_params is the list of LLMTestCaseParams members — INPUT, ACTUAL_OUTPUT, EXPECTED_OUTPUT, RETRIEVAL_CONTEXT and so on — naming which fields of the LLMTestCase the judge model is shown. Fields you leave out are invisible to the judge.

open as a page

In DeepEval, how do you score one LLMTestCase with a built-in metric?

level: juniorimportance: must knowfreq 78%

basics

~10 s

Build an LLMTestCase with at least input and actual_output, construct a metric such as AnswerRelevancyMetric(threshold=0.7), then call metric.measure(test_case). DeepEval then fills metric.score, metric.reason and metric.is_successful() for that one case.

open as a page

In DeepEval, what does assert_test() do inside a pytest test function?

level: juniorimportance: must knowfreq 70%

basics

~10 s

assert_test(test_case=..., metrics=[...]) scores the test case with every metric and raises AssertionError if any metric lands below its threshold. That turns an LLM evaluation into an ordinary pytest failure.

open as a page

What does DeepEval's Synthesizer.generate_goldens_from_docs actually do?

level: middleimportance: must knowfreq 62%

basics

~20 s

It reads the documents you point it at, splits and groups them into contexts, then asks an LLM to write synthetic inputs grounded in each context, rewrites them to be harder, and returns Golden objects — optionally with a generated expected_output.

open as a page

In DeepEval's GEval, when do you pass evaluation_steps instead of criteria?

level: middleimportance: must knowfreq 70%

basics

~20 s

criteria is a sentence GEval expands into judging steps with an extra model call each time the metric is built. evaluation_steps hands GEval that list yourself, so every run applies the identical rubric. Pass one or the other, never both.

open as a page

Which LLMTestCase fields does each built-in DeepEval metric require?

level: middleimportance: must knowfreq 72%

basics

~10 s

All metrics need input and actual_output. FaithfulnessMetric and ContextualRelevancyMetric add retrieval_context; ContextualPrecisionMetric and ContextualRecallMetric add both retrieval_context and expected_output; HallucinationMetric needs context. Missing a required field raises an error rather than scoring zero.

open as a page

Why run `deepeval test run` instead of plain pytest for a DeepEval suite?

level: middleimportance: must knowfreq 62%

basics

~20 s

deepeval test run invokes pytest with DeepEval's plugin active and its own flags on top: parallel processes, repeats, result caching, a run identifier, and an aggregated end-of-run report that can be pushed to Confident AI. Plain pytest still runs the assertions but produces none of that.

open as a page

Your DeepEval suite in CI goes red at random. How do you stabilise it?

level: seniorimportance: must knowfreq 58%

basics

~20 s

Separate metric errors from genuine low scores with the runner's ignore-errors flag, measure each case's score variance with repeated runs, move thresholds off the noise floor, mark genuinely unstable cases flaky so they warn instead of raising, and gate on an aggregate rather than every case.

open as a page

In DeepEval, how do you build a multi-turn ConversationalTestCase?

level: middleimportance: should knowfreq 38%

basics

~20 s

Construct ConversationalTestCase with a turns list of Turn objects, each carrying a role of user or assistant and its content. Optional scenario and expected_outcome describe what the conversation was supposed to achieve, and conversational metrics score the exchange as a whole.

open as a page

In DeepEval's Synthesizer, what do evolutions do to a generated input?

level: middleimportance: should knowfreq 48%

basics

~20 s

Evolutions are rewrite passes that make a freshly generated question harder — adding reasoning steps, forcing several contexts to be combined, adding constraints or hypotheticals. Each pass is another LLM call, and too many push the question away from its source context.

open as a page

How does the rubric argument change scoring in DeepEval's GEval metric?

level: middleimportance: should knowfreq 40%

basics

~20 s

rubric takes a list of Rubric objects, each pairing a score_range with an expected_outcome description. It tells the judge what each band of scores means, so scores land on defined levels instead of on the judge's private sense of what 0.7 is worth.

open as a page

What does strict_mode=True do to a DeepEval GEval metric's score?

level: middleimportance: should knowfreq 45%

basics

~20 s

strict_mode turns the metric into a binary one: the score becomes 1 for a perfect judgement and 0 otherwise, and the threshold is overridden to 1. Partial credit disappears, so anything less than flawless fails.

open as a page

In DeepEval, how do LLMTestCase context and retrieval_context differ?

level: middleimportance: should knowfreq 55%

basics

~20 s

context holds human-supplied ground-truth background for the input; retrieval_context holds the chunks your retriever actually returned at runtime. HallucinationMetric reads context, while FaithfulnessMetric and the contextual metrics read retrieval_context — swapping them is a common wiring bug.

open as a page

In DeepEval, is a high HallucinationMetric score good or bad?

level: middleimportance: should knowfreq 45%

basics

~20 s

Bad. HallucinationMetric and ToxicityMetric are lower-is-better: the score measures how much hallucination or toxicity was found, and the case passes when the score is at or below the threshold — the inverse of AnswerRelevancyMetric or FaithfulnessMetric.

open as a page

When do you use DeepEval's evaluate() instead of assert_test() in pytest?

level: middleimportance: should knowfreq 46%

basics

~20 s

Use evaluate() when you want scores back rather than a pass/fail gate. It takes a list of test cases plus metrics, runs them with its own concurrency settings, and returns an EvaluationResult you can inspect — no pytest, no raised assertion.

open as a page

How do you run a DeepEval suite over a whole dataset as one pytest test per case?

level: middleimportance: should knowfreq 52%

basics

~20 s

Parametrize the test over the dataset's goldens — @pytest.mark.parametrize("golden", dataset.goldens) — and build one LLMTestCase inside the test body. Each golden then becomes its own pytest node, so failures are isolated and the cases can be distributed across processes.

open as a page

How do you keep DeepEval's synthetic goldens from being too easy?

level: seniorimportance: should knowfreq 55%

basics

~20 s

Generate from the chunks your production retriever really returns, weight evolutions toward reasoning and multi-context, use StylingConfig so inputs read like real user messages, set a strong critic model in FiltrationConfig, and sample-read the output before it becomes the standard.

open as a page

How do you write a custom DeepEval metric by subclassing BaseMetric?

level: seniorimportance: should knowfreq 30%

basics

~10 s

Subclass BaseMetric and implement measure(test_case), the async a_measure, and is_successful(), setting self.score and self.success inside measure (plus self.reason if you want an explanation). Expose a name property so the metric is labelled in reports.

open as a page

When would you build a DeepEval DAGMetric instead of a single GEval metric?

level: seniorimportance: should knowfreq 32%

basics

~20 s

Use DAGMetric when the judgement is a sequence of decisions rather than one holistic score — a decision tree of judgement nodes ending in verdicts. It makes the reasoning path explicit and auditable, where GEval collapses everything into one opaque judged number.

open as a page

How many judge LLM calls does a DeepEval metric suite make per test case?

level: seniorimportance: should knowfreq 42%

basics

~20 s

More than one per metric. Each built-in metric runs its own chain of judge prompts — extracting statements, judging them, and with include_reason=True writing an explanation — so cost scales as cases times metrics times several calls, not as one call per case.

open as a page

Which DeepEval built-in metrics work on production traces with no expected_output?

level: seniorimportance: should knowfreq 50%

basics

~10 s

AnswerRelevancyMetric, FaithfulnessMetric, ContextualRelevancyMetric, ToxicityMetric and BiasMetric all run without a reference. ContextualPrecisionMetric and ContextualRecallMetric need expected_output, and HallucinationMetric needs human-supplied context, so none of those three can score raw traffic.

open as a page

Should DeepEval goldens live in your repo or in Confident AI?

level: principalimportance: should knowfreq 40%

basics

~20 s

It is a tradeoff between reproducibility and collaboration. Files in the repo are pinned by commit and reviewable in a pull request; a hosted dataset pulled with dataset.pull(alias=...) lets domain experts curate it, but the suite can change under CI without any code change.

open as a page

How would you keep a DeepEval suite's judge cost and runtime viable on every pull request?

level: principalimportance: should knowfreq 40%

basics

~20 s

Decide what the merge path must prove, then run only that: a small marked subset per pull request with parallel processes and the result cache on, the full sweep on a schedule, cheaper judges for breadth and expensive ones for the gate, and hyperparameters logged so runs stay comparable.

open as a page