skip to content

Built-in Metrics

The stock metric classes and how you feed them a test case. Knowing which metric needs retrieval_context versus expected_output is the fastest way to show you have actually run DeepEval rather than read its README.

on this pageshow

questions

6

In DeepEval, how do you score one LLMTestCase with a built-in metric?

level: juniorimportance: must knowfreq 78%

answer

  1. two objects: the case and the judge
  2. the judge holds the result, not the case
  3. one call fills three attributes
  4. threshold decides pass, not score
  5. measure(test_case) after constructing the metric

basics

~10 s

Build an LLMTestCase with at least input and actual_output, construct a metric such as AnswerRelevancyMetric(threshold=0.7), then call metric.measure(test_case). DeepEval then fills metric.score, metric.reason and metric.is_successful() for that one case.

solid answer

~40 s

DeepEval separates the **data** from the **judge**. The data is an `LLMTestCase`, a plain object whose fields include `input`, `actual_output`, `expected_output`, `context` and `retrieval_context`. The judge is a metric object such as `AnswerRelevancyMetric(threshold=0.7, model="gpt-4o-mini", include_reason=True)`. Calling `metric.measure(test_case)` runs the metric's internal judge prompts against that case and populates three things on the metric instance: `metric.score` (a float, normally 0.0-1.0), `metric.reason` (a sentence explaining the score, only when `include_reason=True`), and pass/fail via `metric.is_successful()`, which compares the score to the `threshold` you passed. Note the state lives on the *metric*, not the test case, so re-measuring overwrites it. Each metric constructs its own judge calls, so measuring one case with four metrics means four independent evaluations of that case.

code

python · 14 lines
python
from deepeval.metrics import AnswerRelevancyMetric
from deepeval.test_case import LLMTestCase

test_case = LLMTestCase(
    input="What is your refund window?",
    actual_output="You can request a refund within 30 days of purchase.",
)

metric = AnswerRelevancyMetric(threshold=0.7, model="gpt-4o-mini", include_reason=True)
metric.measure(test_case)

print(metric.score)
print(metric.reason)
print(metric.is_successful())

go deeper

for a junior

Be able to write the five lines from memory: import the metric and LLMTestCase, build the case with input and actual_output, construct the metric with a threshold, call measure, then read score and is_successful().

for a middle

Explain that results live on the metric instance and are overwritten each call, and that a single measure is several judge requests rather than a local computation.

for a senior

Show that you plan around measure being slow, paid and slightly non-deterministic: choose the judge model deliberately, decide whether reasons are worth their extra generation, and set thresholds with margin.

for a principal

Frame the metric object as the unit of policy: which judge model, what threshold and whether reasons are stored are decisions the whole org inherits, and changing any of them silently reprices and re-baselines every suite that uses it.

## The two objects you need DeepEval is built around two kinds of object, and almost every mistake beginners make comes from confusing them. **`LLMTestCase`** is the data — one interaction you want judged. It is imported from `deepeval.test_case` and its commonly used fields are: - `input` — what the user asked. - `actual_output` — what your application actually produced. - `expected_output` — the human-written reference answer, when you have one. - `context` — ground-truth background supplied by a human. - `retrieval_context` — the chunks your retriever actually fetched at runtime. Only `input` and `actual_output` are always required; the rest are filled in when a metric needs them. **A metric** is the judge — a class imported from `deepeval.metrics`, such as `AnswerRelevancyMetric`, `FaithfulnessMetric`, `ContextualPrecisionMetric`, `ContextualRecallMetric`, `ContextualRelevancyMetric`, `HallucinationMetric` or `ToxicityMetric`. You construct it once with its configuration and then apply it to as many cases as you like. ## The constructor arguments you will actually use - `threshold` — the pass mark. Defaults to 0.5 on the stock metrics. It does not change the score; it only decides whether that score counts as a success. - `model` — which model acts as the judge. You can pass a model name string, or a custom judge object if you have wrapped your own model. Because nearly every stock metric is LLM-judged, this argument is where your evaluation bill is decided. - `include_reason` — when true (the default on the LLM-judged metrics), the metric spends an extra generation writing a human-readable justification into `metric.reason`. - `async_mode` — lets the metric's internal judge calls run concurrently rather than one after another, which matters because a single metric typically issues several calls per case. ## The measure call `metric.measure(test_case)` is synchronous and blocking; `metric.a_measure(test_case)` is the awaitable form for use inside async code. After it returns, read the results **from the metric instance**: - `metric.score` — the float the judge produced. - `metric.reason` — the explanation string, or `None` if you disabled reasons. - `metric.is_successful()` — the boolean verdict against your threshold. This is the single most common trip-up: the score is not attached to the test case, and a metric instance holds only the results of its most recent measurement. If you loop over 100 cases with one metric object and forget to collect `metric.score` inside the loop, you end up with the last case's score only. ## What actually happens during a measure The stock metrics are not string comparisons. Each one runs a small chain of judge prompts. `AnswerRelevancyMetric`, for example, has the judge break the `actual_output` into individual statements, classify each one as relevant or irrelevant to the `input`, and derive a score from that ratio; the reason is then a further generation summarising the verdict. So one `measure()` call is several model requests, it takes seconds not milliseconds, and it costs money. That is why the level of parallelism, the judge model and `include_reason` are the knobs people reach for first when a suite gets slow or expensive. It also means a measure can be non-deterministic. Two runs of the same case can differ slightly, which is why thresholds are usually set with a margin rather than at the exact score you observed once. ## What a missing field does If a metric needs a field the test case does not carry — say you hand `FaithfulnessMetric` a case with no `retrieval_context` — DeepEval raises an error rather than scoring zero. That is deliberate: a silent zero would look like a quality regression when it is really a wiring bug. Knowing which metric needs which field is the practical half of using the built-ins. ## Interview framing The answer an interviewer wants is short and concrete: construct the case, construct the metric with a threshold, call `measure`, read `score` / `reason` / `is_successful()`. If you can then add that each metric is itself an LLM call chain with a configurable judge model, you have shown you have run it rather than skimmed the README.

  • You loop one metric instance over 200 test cases and only inspect metric.score at the end — what do you get?
    Only the last case's score. The metric instance is mutable state: every `measure()` overwrites `score`, `reason` and the success flag. Collect the values inside the loop (or use a fresh metric per case) if you need all 200. This is also why sharing one metric object across threads is unsafe.
  • What does the threshold argument change about the score itself?
    Nothing. The judge produces the score independently; `threshold` is only the pass mark that `is_successful()` compares against. Raising a threshold makes more cases fail, it does not make the model score differently. That separation is what lets you re-run the same recorded scores against a stricter bar later.
  • Why is measure() slow enough that people notice?
    Because it is not arithmetic — it is a chain of LLM judge calls. A single metric typically extracts statements, judges each one, and (with `include_reason=True`) writes an explanation, so one case costs several requests plus their latency. `async_mode` overlaps a metric's internal calls, which is why a suite of a few hundred cases is a background job rather than a unit test.

saying these in an interview costs you the question

  • Thinking measure() returns the score instead of setting metric.score
  • Expecting the score to be stored on the LLMTestCase
  • Believing threshold changes the computed score
  • Assuming a built-in metric is a string comparison, not an LLM call
  • Reusing one metric instance and reading the score only after the loop

context

open as a page

Which LLMTestCase fields does each built-in DeepEval metric require?

level: middleimportance: must knowfreq 72%

basics

~10 s

All metrics need input and actual_output. FaithfulnessMetric and ContextualRelevancyMetric add retrieval_context; ContextualPrecisionMetric and ContextualRecallMetric add both retrieval_context and expected_output; HallucinationMetric needs context. Missing a required field raises an error rather than scoring zero.

open as a page

In DeepEval, how do LLMTestCase context and retrieval_context differ?

level: middleimportance: should knowfreq 55%

basics

~20 s

context holds human-supplied ground-truth background for the input; retrieval_context holds the chunks your retriever actually returned at runtime. HallucinationMetric reads context, while FaithfulnessMetric and the contextual metrics read retrieval_context — swapping them is a common wiring bug.

open as a page

In DeepEval, is a high HallucinationMetric score good or bad?

level: middleimportance: should knowfreq 45%

basics

~20 s

Bad. HallucinationMetric and ToxicityMetric are lower-is-better: the score measures how much hallucination or toxicity was found, and the case passes when the score is at or below the threshold — the inverse of AnswerRelevancyMetric or FaithfulnessMetric.

open as a page

How many judge LLM calls does a DeepEval metric suite make per test case?

level: seniorimportance: should knowfreq 42%

basics

~20 s

More than one per metric. Each built-in metric runs its own chain of judge prompts — extracting statements, judging them, and with include_reason=True writing an explanation — so cost scales as cases times metrics times several calls, not as one call per case.

open as a page

Which DeepEval built-in metrics work on production traces with no expected_output?

level: seniorimportance: should knowfreq 50%

basics

~10 s

AnswerRelevancyMetric, FaithfulnessMetric, ContextualRelevancyMetric, ToxicityMetric and BiasMetric all run without a reference. ContextualPrecisionMetric and ContextualRecallMetric need expected_output, and HallucinationMetric needs human-supplied context, so none of those three can score raw traffic.

open as a page