Which LLMTestCase fields does each built-in DeepEval metric require?
answer
- everything needs input and actual_output
- two metrics also want a reference answer
- precision and recall are the reference pair
- hallucination reads the other context field
- missing field raises, never scores zero
basics
~10 sAll metrics need input and actual_output. FaithfulnessMetric and ContextualRelevancyMetric add retrieval_context; ContextualPrecisionMetric and ContextualRecallMetric add both retrieval_context and expected_output; HallucinationMetric needs context. Missing a required field raises an error rather than scoring zero.
solid answer
~40 sIn deepeval 4.x every metric wants `input` and `actual_output`, and then differs: - `AnswerRelevancyMetric`, `ToxicityMetric`, `BiasMetric` — nothing else. - `FaithfulnessMetric`, `ContextualRelevancyMetric` — plus `retrieval_context`. - `ContextualPrecisionMetric`, `ContextualRecallMetric` — plus `retrieval_context` **and** `expected_output`, because both judge retrieval against a human reference answer. - `HallucinationMetric` — plus `context`, not `retrieval_context`. The practical consequence is that the two contextual metrics that need `expected_output` cannot run on raw production traffic, which has no ground truth, while the rest can. DeepEval validates the required fields at measure time and raises rather than returning 0.0, so a wiring mistake shows up as an error instead of masquerading as a quality regression.
code
python · 17 linesfrom deepeval.metrics import (
AnswerRelevancyMetric,
FaithfulnessMetric,
ContextualRecallMetric,
)
from deepeval.test_case import LLMTestCase
case = LLMTestCase(
input="Do you ship to Canada?",
actual_output="Yes, standard shipping to Canada takes 5-7 business days.",
expected_output="Yes, Canadian orders arrive in 5-7 business days.",
retrieval_context=["Canada: standard shipping, 5-7 business days."],
)
AnswerRelevancyMetric().measure(case) # input + actual_output
FaithfulnessMetric().measure(case) # + retrieval_context
ContextualRecallMetric().measure(case) # + retrieval_context and expected_outputgo deeper
Memorise that input and actual_output are always required, and that anything with the word Contextual or Faithfulness in its name also needs retrieval_context on the test case.
Explain the full map from memory, including that ContextualPrecisionMetric and ContextualRecallMetric additionally need expected_output while ContextualRelevancyMetric does not, and that HallucinationMetric reads context.
Show that the field map drives suite design: reference-requiring metrics on a small curated set, reference-free metrics on sampled traffic, and a loud failure rather than a silent zero when the builder is wrong.
Own the consequence for the labelling budget — every metric that needs expected_output turns into recurring human annotation, so the metric list you standardise on is really a staffing decision.
## Why this is the question that separates readers from users Every DeepEval tutorial shows the same happy-path snippet, so reciting `LLMTestCase(input=..., actual_output=...)` proves nothing. What proves you have run a suite is knowing which metric refuses to run on which case, because that is the error you hit within your first hour. ## The field map `LLMTestCase` (from `deepeval.test_case`) carries, among others: `input`, `actual_output`, `expected_output`, `context`, `retrieval_context`, plus tool-related fields for agent cases. Each metric declares the subset it needs. | Metric | Required fields | |---|---| | `AnswerRelevancyMetric` | `input`, `actual_output` | | `ToxicityMetric` | `input`, `actual_output` | | `BiasMetric` | `input`, `actual_output` | | `FaithfulnessMetric` | `input`, `actual_output`, `retrieval_context` | | `ContextualRelevancyMetric` | `input`, `actual_output`, `retrieval_context` | | `ContextualPrecisionMetric` | `input`, `actual_output`, `expected_output`, `retrieval_context` | | `ContextualRecallMetric` | `input`, `actual_output`, `expected_output`, `retrieval_context` | | `HallucinationMetric` | `input`, `actual_output`, `context` | ## Why each metric needs what it needs The map is not arbitrary — each requirement follows from what the judge is being asked to compare. **Output-only metrics.** Answer relevancy asks whether the answer addresses the question, and toxicity and bias ask whether the text itself is acceptable. None of those comparisons involve retrieval or a reference, so they need only the question and the answer. **Metrics that compare the answer to what was retrieved.** Faithfulness asks whether the answer's claims are supported by the chunks the retriever returned, so it needs `retrieval_context`. Contextual relevancy asks how much of what was retrieved was actually pertinent to the question, so it needs `retrieval_context` too — but it never looks at a reference answer, because "was this chunk on topic?" is answerable from the question alone. **Metrics that need a human reference.** Contextual precision asks whether the *relevant* retrieved chunks were ranked above the irrelevant ones, and contextual recall asks whether everything the reference answer says could have been found in what was retrieved. Both need a definition of "relevant", and DeepEval takes that definition from `expected_output`. That is why these two are the reference-requiring pair. **The odd one out.** `HallucinationMetric` uses `context`, the human-supplied ground-truth background, not `retrieval_context`, the chunks your live retriever produced. Handing it a case that populates only `retrieval_context` is the single most frequent field error in DeepEval. ## What happens when a field is missing DeepEval checks the required parameters before it starts judging and raises an error naming the missing field. This is a deliberate design choice worth calling out in an interview: a framework that scored 0.0 instead would turn a plumbing bug into what looks like a catastrophic quality drop, and a team would spend a day debugging their prompt instead of their test-case builder. The runner also offers a way to skip cases with missing parameters when you are scoring a heterogeneous batch, but the default is to fail loudly. ## The consequence for real suites Because different metrics require different fields, your test-case builder — not your metric list — is what determines which metrics you can run. Teams usually end up with two shapes of case: 1. **Curated cases**, hand-written or generated, that carry `expected_output` and often `context`. These support the full metric set including contextual precision/recall and hallucination. 2. **Production-derived cases**, reconstructed from traces, that carry `input`, `actual_output` and `retrieval_context` but no reference. These support answer relevancy, faithfulness, contextual relevancy, toxicity and bias — and nothing that needs a reference, unless a human labels them first. A good answer names that split, because it is exactly the constraint that shapes how a team's evaluation strategy is built: reference-requiring metrics on a small curated set, reference-free metrics on a sample of live traffic. ## A note on scope What faithfulness or context recall *mean* as measurement concepts is RAG-evaluation theory. What is specific to DeepEval is the field names, the requirement map, and the fact that violating it raises rather than scores zero. Interviewers use this question precisely because the mapping is unmemorable unless you have hit the errors.
- You built cases from production traces with input, actual_output and retrieval_context — which built-in metrics can you still run?Answer relevancy, faithfulness, contextual relevancy, toxicity and bias all run, because none of them needs a reference. Contextual precision and contextual recall cannot, since both require `expected_output`, and hallucination cannot either, since it wants human-supplied `context`. To reach those three you must have someone label the traces first.
- Why does DeepEval raise instead of returning a score of 0.0 when a required field is missing?Because a zero is indistinguishable from a genuine quality collapse. Raising turns a plumbing bug into an obviously-wrong build rather than a plausible-looking regression that sends the team to debug prompts. The runner offers a skip-on-missing-params path for heterogeneous batches, but the loud failure is the default for good reason.
- Which two contextual metrics differ only by needing expected_output, and why?`ContextualRelevancyMetric` needs only `retrieval_context`, because judging whether a chunk is on topic is answerable from the question alone. `ContextualPrecisionMetric` and `ContextualRecallMetric` add `expected_output` because they judge ranking and coverage against a definition of the right answer, and DeepEval takes that definition from the reference.
saying these in an interview costs you the question
- Assuming every metric only needs input and actual_output
- Passing retrieval_context to HallucinationMetric and expecting it to work
- Thinking a missing field yields a score of zero
- Believing contextual relevancy needs expected_output
- Claiming reference-requiring metrics can score raw production traffic