skip to content

G-Eval & Custom Metrics

How you score something no built-in metric covers, by describing the criteria in natural language and letting a judge model apply it. The interesting part is making that judgement repeatable enough to gate a pipeline on.

on this pageshow

questions

6

What does evaluation_params control in a DeepEval GEval metric?

level: juniorimportance: must knowfreq 65%

answer

  1. the judge only sees what you list
  2. enum members, not raw strings
  3. a step about context needs the context param
  4. listed means required on the test case
  5. EXPECTED_OUTPUT makes it dataset-only

basics

~20 s

evaluation_params is the list of LLMTestCaseParams members — INPUT, ACTUAL_OUTPUT, EXPECTED_OUTPUT, RETRIEVAL_CONTEXT and so on — naming which fields of the LLMTestCase the judge model is shown. Fields you leave out are invisible to the judge.

solid answer

~40 s

`GEval` does not automatically see the whole test case. `evaluation_params` takes a list of `LLMTestCaseParams` enum members — `INPUT`, `ACTUAL_OUTPUT`, `EXPECTED_OUTPUT`, `CONTEXT`, `RETRIEVAL_CONTEXT`, and the tool-related ones — and only those fields of the `LLMTestCase` are put in front of the judge model. That makes it a real correctness knob, not boilerplate. A rubric step that says "check the answer is grounded in the retrieved passages" is meaningless unless `RETRIEVAL_CONTEXT` is in the list; the judge will invent an assessment from nothing. Every parameter you list must also be populated on the test case, or measuring raises an error. It is also where reference-dependence is decided: the moment `EXPECTED_OUTPUT` appears, the metric needs ground truth for every case, which a curated dataset has and live production traffic does not.

code

python · 25 lines
python
from deepeval.metrics import GEval
from deepeval.test_case import LLMTestCase, LLMTestCaseParams

# Reference-free: runnable over sampled production traffic.
groundedness = GEval(
    name="Groundedness",
    evaluation_steps=[
        "List every factual claim made in 'actual output'.",
        "For each claim, check whether 'retrieval context' supports it.",
        "Score down in proportion to the number of unsupported claims.",
    ],
    evaluation_params=[
        LLMTestCaseParams.INPUT,
        LLMTestCaseParams.ACTUAL_OUTPUT,
        LLMTestCaseParams.RETRIEVAL_CONTEXT,
    ],
)

case = LLMTestCase(
    input="What is our refund window?",
    actual_output="You can request a refund within 30 days of delivery.",
    retrieval_context=["Refunds are accepted up to 30 days after delivery."],
)
groundedness.measure(case)
print(groundedness.score, groundedness.reason)

go deeper

for a junior

Recall that evaluation_params is a list of LLMTestCaseParams members and that it decides which test-case fields the judge model actually sees. Name INPUT, ACTUAL_OUTPUT, EXPECTED_OUTPUT and RETRIEVAL_CONTEXT.

for a middle

Explain the coupling between rubric text and parameters: a step referring to a field you did not pass produces a confident but baseless score. Also say that every listed parameter must be populated or measuring raises.

for a senior

Show the reference-based versus reference-free split as a design decision: EXPECTED_OUTPUT confines a metric to curated datasets, so production-traffic metrics must be grounded in INPUT and RETRIEVAL_CONTEXT instead. Mention dataset validation before a paid run.

for a principal

Own the metric portfolio: which custom metrics are dataset-only and which are safe to point at sampled live traffic, and how the token cost of large context fields scales across the whole suite at your sample volumes.

## The mechanic An `LLMTestCase` in DeepEval carries several fields — `input`, `actual_output`, `expected_output`, `context`, `retrieval_context`, `tools_called`, `expected_tools`. A `GEval` metric renders a prompt for the judge model, and `evaluation_params` decides which of those fields go into that prompt. You express the selection with members of the `LLMTestCaseParams` enum: ```python from deepeval.test_case import LLMTestCase, LLMTestCaseParams evaluation_params=[ LLMTestCaseParams.INPUT, LLMTestCaseParams.ACTUAL_OUTPUT, LLMTestCaseParams.RETRIEVAL_CONTEXT, ] ``` Anything not listed simply is not shown. There is no implicit "the judge can see everything" default. ## Why it is a correctness knob The classic failure is a mismatch between the rubric and the parameters. Someone writes evaluation steps about groundedness — "penalise any claim not supported by the retrieved passages" — and lists only `INPUT` and `ACTUAL_OUTPUT`. The judge cannot refuse; it produces a score anyway, based on plausibility rather than on the passages. The metric looks like it works, reports numbers with confident-sounding reasons, and measures something other than what its name says. The reverse failure is over-inclusion. Handing the judge `EXPECTED_OUTPUT` on a metric that is meant to score tone or format invites it to mark down outputs that differ from the reference in ways the rubric never asked about. More context in the prompt also means more tokens, and every sample pays for them. So the rule of thumb: list precisely the fields your evaluation steps refer to, and no more. ## The population requirement Every parameter you list must be populated on the `LLMTestCase` you measure. If `EXPECTED_OUTPUT` is in `evaluation_params` and a case leaves `expected_output` as `None`, measuring that case errors rather than silently scoring it. This is deliberate — a half-populated judge prompt would produce a number that means nothing — but it does mean a dataset with patchy fields will break a run partway through. Validate the dataset before the suite, not during it. ## Reference-dependence, and where it bites Deciding whether `EXPECTED_OUTPUT` is in the list is the same decision as deciding whether the metric is reference-based: - **With `EXPECTED_OUTPUT`** the metric compares against ground truth. It is sharp and it catches factual drift, but it only runs where ground truth exists: a curated dataset with human-written answers. - **Without it** the metric judges the output on its own terms — is it relevant, is it grounded in the supplied context, does it follow the policy. This kind can run over anything, including outputs captured from production, because it needs no answer key. Teams hit this when they try to point their offline suite at real traffic. Half the metrics run and half of them error or would need a reference nobody wrote. The practical answer is to build two families of custom metrics: reference-based ones for the regression dataset, and reference-free ones (grounded in `RETRIEVAL_CONTEXT` and `INPUT`) for sampled live traffic. ## Writing steps and params together A useful discipline is to phrase evaluation steps using the field names themselves — "compare 'actual output' with 'expected output'" — because it makes the dependency obvious at review time. If a step mentions a field that is not in `evaluation_params`, that is a bug you can spot by reading, without running anything. ## Quick checklist - Do the steps mention a field the judge cannot see? Add it or reword the step. - Is a field listed that no step uses? Remove it — tokens and unwanted influence. - Is `EXPECTED_OUTPUT` listed? Then this metric is dataset-only. - Does every case in the dataset populate all listed fields? If not, the run errors.

  • What happens if a listed parameter is missing on a test case at measure time?
    DeepEval raises rather than scoring it. A parameter in evaluation_params is a contract: the judge prompt has a slot for that field, and filling it with nothing would produce a score with no basis. Practically this means dataset validation belongs before the run — sweep the goldens for empty expected_output or retrieval_context first, so you fail fast on the dataset instead of halfway through a paid evaluation.
  • Which parameters would you pick for a metric that scores whether an answer stays grounded in retrieved passages?
    INPUT, ACTUAL_OUTPUT and RETRIEVAL_CONTEXT. The judge needs the question to know what was asked, the answer to inspect, and the retrieved passages to check each claim against. Deliberately leave EXPECTED_OUTPUT out: groundedness is about support from the passages, not agreement with a reference answer, and omitting it keeps the metric runnable on production traffic where no reference exists.
  • Is there a cost to listing more parameters than the rubric needs?
    Two costs. Every listed field is rendered into the judge prompt, so you pay input tokens for it on every single sample — noticeable across thousands of cases. And extra context influences the judge: showing a reference answer to a tone metric invites it to penalise wording differences the rubric never mentioned. List what the steps actually reference.

saying these in an interview costs you the question

  • Assumes the judge sees the whole LLMTestCase regardless of evaluation_params
  • Writes rubric steps about context that was never passed as a parameter
  • Thinks a missing listed field is silently skipped instead of raising
  • Treats evaluation_params as required boilerplate with no effect on the score
  • Plans to run an EXPECTED_OUTPUT metric over production traffic with no ground truth

context

open as a page

In DeepEval's GEval, when do you pass evaluation_steps instead of criteria?

level: middleimportance: must knowfreq 70%

basics

~20 s

criteria is a sentence GEval expands into judging steps with an extra model call each time the metric is built. evaluation_steps hands GEval that list yourself, so every run applies the identical rubric. Pass one or the other, never both.

open as a page

How does the rubric argument change scoring in DeepEval's GEval metric?

level: middleimportance: should knowfreq 40%

basics

~20 s

rubric takes a list of Rubric objects, each pairing a score_range with an expected_outcome description. It tells the judge what each band of scores means, so scores land on defined levels instead of on the judge's private sense of what 0.7 is worth.

open as a page

What does strict_mode=True do to a DeepEval GEval metric's score?

level: middleimportance: should knowfreq 45%

basics

~20 s

strict_mode turns the metric into a binary one: the score becomes 1 for a perfect judgement and 0 otherwise, and the threshold is overridden to 1. Partial credit disappears, so anything less than flawless fails.

open as a page

How do you write a custom DeepEval metric by subclassing BaseMetric?

level: seniorimportance: should knowfreq 30%

basics

~10 s

Subclass BaseMetric and implement measure(test_case), the async a_measure, and is_successful(), setting self.score and self.success inside measure (plus self.reason if you want an explanation). Expose a name property so the metric is labelled in reports.

open as a page

When would you build a DeepEval DAGMetric instead of a single GEval metric?

level: seniorimportance: should knowfreq 32%

basics

~20 s

Use DAGMetric when the judgement is a sequence of decisions rather than one holistic score — a decision tree of judgement nodes ending in verdicts. It makes the reasoning path explicit and auditable, where GEval collapses everything into one opaque judged number.

open as a page