skip to content

In DeepEval, how do LLMTestCase context and retrieval_context differ?

level: middleimportance: should knowfreq 55%

answer

  1. ideal versus actual background
  2. one is authored, one is recorded
  3. hallucination reads the human-written field
  4. the other four grade the pipeline
  5. swapping them fails silently, never errors

basics

~20 s

context holds human-supplied ground-truth background for the input; retrieval_context holds the chunks your retriever actually returned at runtime. HallucinationMetric reads context, while FaithfulnessMetric and the contextual metrics read retrieval_context — swapping them is a common wiring bug.

solid answer

~40 s

They look interchangeable and are not. `context` is the **ideal** background: text a human asserts is true for this input, independent of your system. `retrieval_context` is the **actual** background: whatever your retriever pulled for this query on this run, correct or not. That distinction decides which metric can use which. `HallucinationMetric` compares `actual_output` against `context`, so it is asking "did the model contradict known truth?". `FaithfulnessMetric`, `ContextualPrecisionMetric`, `ContextualRecallMetric` and `ContextualRelevancyMetric` all read `retrieval_context`, so they are asking about your pipeline: was the answer grounded in what was fetched, and was what was fetched any good? Populating `retrieval_context` from a hand-written gold document, or `context` from your live retriever, produces scores that are not measuring what you think — the usual symptom is a hallucination score that only ever reflects retrieval quality.

code

python · 12 lines
python
from deepeval.metrics import FaithfulnessMetric, HallucinationMetric
from deepeval.test_case import LLMTestCase

case = LLMTestCase(
    input="When does the warranty expire?",
    actual_output="The warranty runs for 24 months from purchase.",
    context=["Warranty period is 12 months from date of purchase."],
    retrieval_context=["Extended plan: 24 months of coverage available."],
)

HallucinationMetric().measure(case)  # contradicts context -> flagged
FaithfulnessMetric().measure(case)   # grounded in retrieval_context -> looks fine

go deeper

for a junior

Remember the one-line difference: context is written by a human as the truth, retrieval_context is what your search step actually returned at runtime.

for a middle

Name which metric reads which field, and explain that both are lists of strings so a swap produces plausible scores rather than an error.

for a senior

Demonstrate the failure mode: a confidently wrong retriever plus a faithful model scores clean on faithfulness and is only caught against human context, which is why the two data sources must stay separate.

for a principal

Own the cost implication — context can only come from human labelling, so how many cases carry it is a budget decision that determines whether factual regressions are detectable at all.

## Two fields, two different questions `LLMTestCase` carries both `context` and `retrieval_context`, both are lists of strings, and nothing in the type signature stops you from putting the same value in both. The difference is entirely semantic, and it maps onto two genuinely different evaluation questions. **`context` — what is true.** This is the ground-truth background a human attaches to the test case: the paragraph from the policy document that actually answers the question, the facts about the account, the source of truth. It exists independently of your application. If you swapped out your retriever tomorrow, `context` would not change. **`retrieval_context` — what your system saw.** This is the output of your retrieval step for this input on this run: the chunks a vector search returned, in the order it returned them, including the irrelevant ones. It is a *record of system behaviour*. If you change your embedding model, chunk size or top-k, `retrieval_context` changes with it. ## Which metric reads which - `HallucinationMetric` → `context`. It asks whether the output contradicts what is known to be true. - `FaithfulnessMetric` → `retrieval_context`. It asks whether the output's claims are supported by what was actually fetched. - `ContextualPrecisionMetric`, `ContextualRecallMetric`, `ContextualRelevancyMetric` → `retrieval_context`. These grade the retrieval step itself. So `context` supports a question about the *model*, and `retrieval_context` supports questions about the *pipeline*. That is the cleanest one-line version of the distinction, and it is what an interviewer is listening for. ## Why swapping them silently corrupts a suite Suppose you populate `context` from your live retriever because it was the convenient variable to hand. `HallucinationMetric` now grades the model against its own retrieved chunks — which is roughly what faithfulness already does. Your hallucination score stops being able to detect the case that matters most: a retriever that returns confident, wrong text, which the model then faithfully repeats. The output contradicts reality but agrees perfectly with what was retrieved, so it scores clean. You have built a metric that cannot fail on your worst failure mode. The reverse mistake is milder but still wrong. Populating `retrieval_context` from a curated gold passage makes faithfulness and the contextual metrics grade a retrieval step that never ran. Contextual relevancy will look excellent — of course the hand-picked passage is relevant — and you will ship a retriever nobody measured. Both failures share a signature: the scores are stable, high and useless. Nothing errors, because both fields are just lists of strings. ## How to keep them straight in practice The rule of thumb that survives contact with a real codebase: **the field you can only fill by hand is `context`; the field you can only fill by running your app is `retrieval_context`.** If your test-case builder is filling both from the same variable, one of them is wrong. This also explains the asymmetry in where cases come from. Cases reconstructed from production traces naturally carry `retrieval_context` — the trace recorded it — but never `context`, because nobody wrote the ground truth for a live query. Curated cases carry `context` because a human authored them, and gain `retrieval_context` only when you replay them through the application. A mature suite ends up with cases in both shapes and simply runs different metric sets over each. ## A related trap Because `HallucinationMetric` is the only stock metric using `context`, teams sometimes drop `context` entirely and conclude they have no way to catch factual errors. The honest answer is that without ground truth you cannot check facts against reality with a stock metric — you can only check grounding against retrieved text with faithfulness. Recognising that limit, rather than pretending faithfulness covers it, is the senior-sounding version of this answer. ## What is out of scope here What "faithfulness" means as a measurement concept, and whether entailment-based grounding checks are trustworthy, is evaluation theory that lives outside the tool. The DeepEval-specific content is exactly this: two similarly named list-of-string fields, a fixed mapping from metric to field, and no runtime protection against putting the wrong data in either.

  • Your retriever returns confident but wrong text and the model repeats it — which metric catches that, and which does not?
    Faithfulness will not: the answer is fully supported by `retrieval_context`, so it scores well. `HallucinationMetric` catches it, because it compares the output against human-supplied `context` — reality, not what was fetched. This is the concrete reason the two fields must not be filled from the same source.
  • Can you populate context from production traces?
    Not honestly. A trace records what your retriever returned, which is `retrieval_context` by definition; nobody wrote ground truth for a live query. To get `context` on trace-derived cases someone has to label them. That labelling cost is why hallucination scoring is usually confined to a small curated set.
  • If both fields hold the same list, what do the scores tell you?
    That faithfulness and hallucination are now asking nearly the same question, so you have redundancy where you thought you had coverage. Nothing errors — both fields are just lists of strings — which is what makes the mistake durable. The tell is suspiciously stable, high scores on a suite that never catches anything.

saying these in an interview costs you the question

  • Treating context and retrieval_context as synonyms
  • Filling both fields from the same variable in the builder
  • Believing FaithfulnessMetric checks facts against reality
  • Expecting DeepEval to error when the fields are swapped
  • Claiming production traces can supply ground-truth context

context