In DeepEval, what is the difference between a Golden and an LLMTestCase?
answer
- one is stored, one is produced
- which field only exists after a run
- datasets outlive models
- actual_output is the dividing line
- goldens in, test cases out
basics
~20 sA Golden holds only the static half of a test — input, expected_output, context — with no actual_output. An LLMTestCase adds the actual_output your application produced on this run, and that is what metrics score.
solid answer
~50 sA `Golden` is the stored, reusable half of a test: `input`, and optionally `expected_output`, `context` and `retrieval_context`. It deliberately has no meaningful `actual_output`, because that only exists once you run your app. An `EvaluationDataset` holds a list of goldens (`dataset.goldens`) — that list is model-agnostic and survives every model, prompt and retriever change. At evaluation time you iterate the goldens, call your application with `golden.input`, and build an `LLMTestCase` carrying `input`, the freshly produced `actual_output`, and whichever golden fields the metrics need. The split is what lets one curated dataset be replayed against a new model tomorrow: if actual outputs were baked into the dataset, every run would need a new dataset. `Golden` does expose an optional `actual_output` field for pre-computed outputs, but the normal flow leaves it empty and fills it at run time.
code
python · 30 linesfrom deepeval.dataset import EvaluationDataset, Golden
from deepeval.test_case import LLMTestCase
dataset = EvaluationDataset(
goldens=[
Golden(
input="What is the refund window?",
expected_output="Thirty days from delivery.",
context=["Refunds are accepted within 30 days of delivery."],
)
]
)
def my_llm_app(question: str):
return "Thirty days.", ["Refunds are accepted within 30 days of delivery."]
test_cases = []
for golden in dataset.goldens:
answer, chunks = my_llm_app(golden.input)
test_cases.append(
LLMTestCase(
input=golden.input,
actual_output=answer,
expected_output=golden.expected_output,
context=golden.context,
retrieval_context=chunks,
)
)go deeper
Be able to say plainly that a Golden is the saved input and reference, and an LLMTestCase adds the actual output your app produced. Know that you loop over dataset.goldens and build test cases at run time.
Explain the field-by-field mapping, including which fields a metric requires and why context and retrieval_context are not the same thing. Be ready to write the golden-to-test-case loop from memory.
Show the operational payoff: one curated golden set replayed across model, prompt and retriever changes, with actual outputs never written back. Talk about who owns the goldens and how references get vouched for.
Own the dataset as a long-lived asset separate from any build. Argue for what belongs in a golden versus in run metadata, and how that choice determines whether year-old results are still comparable to today's.
## Two objects, two lifetimes DeepEval separates the part of a test that you *maintain* from the part that is *produced*. The maintained part is a `Golden`. The produced part is an `LLMTestCase`. A `Golden` is imported from `deepeval.dataset` and carries: - `input` — the user query or task. The only genuinely required field. - `expected_output` — the reference answer, when you have one. - `context` — the ground-truth facts a correct answer should rest on. - `retrieval_context` — chunks a retriever returned, when you are storing them. - `additional_metadata` and `comments` — free-form annotations (source document, owner, why this case exists). - `actual_output` — present on the class but normally left empty. An `LLMTestCase` is imported from `deepeval.test_case` and is what metrics actually consume. Its distinguishing field is `actual_output`: the string your application produced for this input, on this run, with this model and this prompt. ## Why the split exists The value of an evaluation dataset is that it outlives any particular build. You curate fifty hard support questions once; you then replay them against GPT-class model A, model B, a new system prompt, a re-chunked index, and a cheaper retriever. Everything that changes between those runs is the output. Everything that stays is the golden. If actual outputs were stored inside the dataset, the dataset would be a snapshot of one run rather than a test suite, and "re-run the suite on the new model" would mean rebuilding it. The split also makes the dataset shareable with non-engineers: a domain expert can write inputs and expected outputs without ever running the app. A second consequence is that goldens are what the `Synthesizer` produces. `generate_goldens_from_docs` cannot produce an `actual_output` — it has no idea what your application would say — so `Golden` is the only object it *can* return. ## The run loop The idiomatic flow is a loop over `dataset.goldens` that invokes your application and constructs test cases: 1. Pull or build an `EvaluationDataset`. 2. For each golden, call your app with `golden.input`. 3. Construct `LLMTestCase(input=golden.input, actual_output=..., expected_output=golden.expected_output, retrieval_context=...)`. 4. Hand the resulting test cases to your metrics. Note step 3 copies `retrieval_context` from the *live* run, not from the golden, when you are evaluating retrieval quality — the chunks your retriever fetched this time are the thing under test. `context`, by contrast, is ground truth and comes from the golden. ## Which fields a metric needs Different metrics require different fields, and a missing one raises an error rather than silently scoring zero. Reference-free metrics typically need only `input` and `actual_output`; reference-requiring ones also need `expected_output` or `context`. This matters when you plan the dataset: if you intend to run a metric that compares against a reference, the golden must carry that reference, and you must decide at curation time — not at evaluation time — where it comes from. ## Where teams get this wrong The most common mistake is treating the dataset as a store of past outputs — pushing actual outputs back into goldens after a run, which turns the suite into a regression-against-yesterday's-model rather than a test of correctness. The second is filling `context` with whatever the retriever returned, which quietly makes retrieval look perfect: `context` is supposed to be the ground truth a human vouches for, and `retrieval_context` is what the system found. Conflating them removes the very gap the evaluation is meant to measure. The third is assuming a golden must have an `expected_output`. Many useful goldens do not — an open-ended question with no single right answer is still a perfectly good test case for a reference-free metric, and inventing a reference answer just to fill the field creates a false standard that penalizes correct outputs.
- If a Golden can hold an actual_output field, when would you ever populate it?When outputs were produced elsewhere and you are only scoring them — for example a batch you exported from production or from an offline inference job, where the golden is really a record of a completed run. It is also handy for a fixture that pins a known-bad output so you can assert a metric catches it. In the normal flow you leave it empty and fill actual_output on the LLMTestCase at run time.
- What is the difference between a golden's context and the retrieval_context you attach to the test case?context is ground truth: the facts a human says a correct answer should rest on, stored on the golden and stable across runs. retrieval_context is what your retriever actually returned on this run, so it changes whenever the index, embedding model or chunking changes. Retrieval-quality metrics compare the two; filling context from the retriever collapses that comparison and makes retrieval look flawless.
- Your dataset has 200 goldens and no expected_output on any of them. What can you still evaluate?Plenty. Reference-free scoring only needs input and actual_output, so relevancy-style and safety-style checks still run. What you lose is anything that compares against a reference answer or against ground-truth facts. If those matter, the fix is a curation pass to add references to the subset where a single correct answer genuinely exists, rather than generating references and treating them as authoritative.
saying these in an interview costs you the question
- Says Golden and LLMTestCase are just aliases for each other
- Stores each run's outputs back into the dataset as goldens
- Fills the golden's context from the retriever's own output
- Claims every golden must carry an expected_output
- Thinks a dataset must be regenerated for each new model