In DeepEval, how do you score one LLMTestCase with a built-in metric?
answer
- two objects: the case and the judge
- the judge holds the result, not the case
- one call fills three attributes
- threshold decides pass, not score
- measure(test_case) after constructing the metric
basics
~10 sBuild an LLMTestCase with at least input and actual_output, construct a metric such as AnswerRelevancyMetric(threshold=0.7), then call metric.measure(test_case). DeepEval then fills metric.score, metric.reason and metric.is_successful() for that one case.
solid answer
~40 sDeepEval separates the **data** from the **judge**. The data is an `LLMTestCase`, a plain object whose fields include `input`, `actual_output`, `expected_output`, `context` and `retrieval_context`. The judge is a metric object such as `AnswerRelevancyMetric(threshold=0.7, model="gpt-4o-mini", include_reason=True)`. Calling `metric.measure(test_case)` runs the metric's internal judge prompts against that case and populates three things on the metric instance: `metric.score` (a float, normally 0.0-1.0), `metric.reason` (a sentence explaining the score, only when `include_reason=True`), and pass/fail via `metric.is_successful()`, which compares the score to the `threshold` you passed. Note the state lives on the *metric*, not the test case, so re-measuring overwrites it. Each metric constructs its own judge calls, so measuring one case with four metrics means four independent evaluations of that case.
code
python · 14 linesfrom deepeval.metrics import AnswerRelevancyMetric
from deepeval.test_case import LLMTestCase
test_case = LLMTestCase(
input="What is your refund window?",
actual_output="You can request a refund within 30 days of purchase.",
)
metric = AnswerRelevancyMetric(threshold=0.7, model="gpt-4o-mini", include_reason=True)
metric.measure(test_case)
print(metric.score)
print(metric.reason)
print(metric.is_successful())go deeper
Be able to write the five lines from memory: import the metric and LLMTestCase, build the case with input and actual_output, construct the metric with a threshold, call measure, then read score and is_successful().
Explain that results live on the metric instance and are overwritten each call, and that a single measure is several judge requests rather than a local computation.
Show that you plan around measure being slow, paid and slightly non-deterministic: choose the judge model deliberately, decide whether reasons are worth their extra generation, and set thresholds with margin.
Frame the metric object as the unit of policy: which judge model, what threshold and whether reasons are stored are decisions the whole org inherits, and changing any of them silently reprices and re-baselines every suite that uses it.
## The two objects you need DeepEval is built around two kinds of object, and almost every mistake beginners make comes from confusing them. **`LLMTestCase`** is the data — one interaction you want judged. It is imported from `deepeval.test_case` and its commonly used fields are: - `input` — what the user asked. - `actual_output` — what your application actually produced. - `expected_output` — the human-written reference answer, when you have one. - `context` — ground-truth background supplied by a human. - `retrieval_context` — the chunks your retriever actually fetched at runtime. Only `input` and `actual_output` are always required; the rest are filled in when a metric needs them. **A metric** is the judge — a class imported from `deepeval.metrics`, such as `AnswerRelevancyMetric`, `FaithfulnessMetric`, `ContextualPrecisionMetric`, `ContextualRecallMetric`, `ContextualRelevancyMetric`, `HallucinationMetric` or `ToxicityMetric`. You construct it once with its configuration and then apply it to as many cases as you like. ## The constructor arguments you will actually use - `threshold` — the pass mark. Defaults to 0.5 on the stock metrics. It does not change the score; it only decides whether that score counts as a success. - `model` — which model acts as the judge. You can pass a model name string, or a custom judge object if you have wrapped your own model. Because nearly every stock metric is LLM-judged, this argument is where your evaluation bill is decided. - `include_reason` — when true (the default on the LLM-judged metrics), the metric spends an extra generation writing a human-readable justification into `metric.reason`. - `async_mode` — lets the metric's internal judge calls run concurrently rather than one after another, which matters because a single metric typically issues several calls per case. ## The measure call `metric.measure(test_case)` is synchronous and blocking; `metric.a_measure(test_case)` is the awaitable form for use inside async code. After it returns, read the results **from the metric instance**: - `metric.score` — the float the judge produced. - `metric.reason` — the explanation string, or `None` if you disabled reasons. - `metric.is_successful()` — the boolean verdict against your threshold. This is the single most common trip-up: the score is not attached to the test case, and a metric instance holds only the results of its most recent measurement. If you loop over 100 cases with one metric object and forget to collect `metric.score` inside the loop, you end up with the last case's score only. ## What actually happens during a measure The stock metrics are not string comparisons. Each one runs a small chain of judge prompts. `AnswerRelevancyMetric`, for example, has the judge break the `actual_output` into individual statements, classify each one as relevant or irrelevant to the `input`, and derive a score from that ratio; the reason is then a further generation summarising the verdict. So one `measure()` call is several model requests, it takes seconds not milliseconds, and it costs money. That is why the level of parallelism, the judge model and `include_reason` are the knobs people reach for first when a suite gets slow or expensive. It also means a measure can be non-deterministic. Two runs of the same case can differ slightly, which is why thresholds are usually set with a margin rather than at the exact score you observed once. ## What a missing field does If a metric needs a field the test case does not carry — say you hand `FaithfulnessMetric` a case with no `retrieval_context` — DeepEval raises an error rather than scoring zero. That is deliberate: a silent zero would look like a quality regression when it is really a wiring bug. Knowing which metric needs which field is the practical half of using the built-ins. ## Interview framing The answer an interviewer wants is short and concrete: construct the case, construct the metric with a threshold, call `measure`, read `score` / `reason` / `is_successful()`. If you can then add that each metric is itself an LLM call chain with a configurable judge model, you have shown you have run it rather than skimmed the README.
- You loop one metric instance over 200 test cases and only inspect metric.score at the end — what do you get?Only the last case's score. The metric instance is mutable state: every `measure()` overwrites `score`, `reason` and the success flag. Collect the values inside the loop (or use a fresh metric per case) if you need all 200. This is also why sharing one metric object across threads is unsafe.
- What does the threshold argument change about the score itself?Nothing. The judge produces the score independently; `threshold` is only the pass mark that `is_successful()` compares against. Raising a threshold makes more cases fail, it does not make the model score differently. That separation is what lets you re-run the same recorded scores against a stricter bar later.
- Why is measure() slow enough that people notice?Because it is not arithmetic — it is a chain of LLM judge calls. A single metric typically extracts statements, judges each one, and (with `include_reason=True`) writes an explanation, so one case costs several requests plus their latency. `async_mode` overlaps a metric's internal calls, which is why a suite of a few hundred cases is a background job rather than a unit test.
saying these in an interview costs you the question
- Thinking measure() returns the score instead of setting metric.score
- Expecting the score to be stored on the LLMTestCase
- Believing threshold changes the computed score
- Assuming a built-in metric is a string comparison, not an LLM call
- Reusing one metric instance and reading the score only after the loop