What does evaluation_params control in a DeepEval GEval metric?
answer
- the judge only sees what you list
- enum members, not raw strings
- a step about context needs the context param
- listed means required on the test case
- EXPECTED_OUTPUT makes it dataset-only
basics
~20 sevaluation_params is the list of LLMTestCaseParams members — INPUT, ACTUAL_OUTPUT, EXPECTED_OUTPUT, RETRIEVAL_CONTEXT and so on — naming which fields of the LLMTestCase the judge model is shown. Fields you leave out are invisible to the judge.
solid answer
~40 s`GEval` does not automatically see the whole test case. `evaluation_params` takes a list of `LLMTestCaseParams` enum members — `INPUT`, `ACTUAL_OUTPUT`, `EXPECTED_OUTPUT`, `CONTEXT`, `RETRIEVAL_CONTEXT`, and the tool-related ones — and only those fields of the `LLMTestCase` are put in front of the judge model. That makes it a real correctness knob, not boilerplate. A rubric step that says "check the answer is grounded in the retrieved passages" is meaningless unless `RETRIEVAL_CONTEXT` is in the list; the judge will invent an assessment from nothing. Every parameter you list must also be populated on the test case, or measuring raises an error. It is also where reference-dependence is decided: the moment `EXPECTED_OUTPUT` appears, the metric needs ground truth for every case, which a curated dataset has and live production traffic does not.
code
python · 25 linesfrom deepeval.metrics import GEval
from deepeval.test_case import LLMTestCase, LLMTestCaseParams
# Reference-free: runnable over sampled production traffic.
groundedness = GEval(
name="Groundedness",
evaluation_steps=[
"List every factual claim made in 'actual output'.",
"For each claim, check whether 'retrieval context' supports it.",
"Score down in proportion to the number of unsupported claims.",
],
evaluation_params=[
LLMTestCaseParams.INPUT,
LLMTestCaseParams.ACTUAL_OUTPUT,
LLMTestCaseParams.RETRIEVAL_CONTEXT,
],
)
case = LLMTestCase(
input="What is our refund window?",
actual_output="You can request a refund within 30 days of delivery.",
retrieval_context=["Refunds are accepted up to 30 days after delivery."],
)
groundedness.measure(case)
print(groundedness.score, groundedness.reason)go deeper
Recall that evaluation_params is a list of LLMTestCaseParams members and that it decides which test-case fields the judge model actually sees. Name INPUT, ACTUAL_OUTPUT, EXPECTED_OUTPUT and RETRIEVAL_CONTEXT.
Explain the coupling between rubric text and parameters: a step referring to a field you did not pass produces a confident but baseless score. Also say that every listed parameter must be populated or measuring raises.
Show the reference-based versus reference-free split as a design decision: EXPECTED_OUTPUT confines a metric to curated datasets, so production-traffic metrics must be grounded in INPUT and RETRIEVAL_CONTEXT instead. Mention dataset validation before a paid run.
Own the metric portfolio: which custom metrics are dataset-only and which are safe to point at sampled live traffic, and how the token cost of large context fields scales across the whole suite at your sample volumes.
## The mechanic An `LLMTestCase` in DeepEval carries several fields — `input`, `actual_output`, `expected_output`, `context`, `retrieval_context`, `tools_called`, `expected_tools`. A `GEval` metric renders a prompt for the judge model, and `evaluation_params` decides which of those fields go into that prompt. You express the selection with members of the `LLMTestCaseParams` enum: ```python from deepeval.test_case import LLMTestCase, LLMTestCaseParams evaluation_params=[ LLMTestCaseParams.INPUT, LLMTestCaseParams.ACTUAL_OUTPUT, LLMTestCaseParams.RETRIEVAL_CONTEXT, ] ``` Anything not listed simply is not shown. There is no implicit "the judge can see everything" default. ## Why it is a correctness knob The classic failure is a mismatch between the rubric and the parameters. Someone writes evaluation steps about groundedness — "penalise any claim not supported by the retrieved passages" — and lists only `INPUT` and `ACTUAL_OUTPUT`. The judge cannot refuse; it produces a score anyway, based on plausibility rather than on the passages. The metric looks like it works, reports numbers with confident-sounding reasons, and measures something other than what its name says. The reverse failure is over-inclusion. Handing the judge `EXPECTED_OUTPUT` on a metric that is meant to score tone or format invites it to mark down outputs that differ from the reference in ways the rubric never asked about. More context in the prompt also means more tokens, and every sample pays for them. So the rule of thumb: list precisely the fields your evaluation steps refer to, and no more. ## The population requirement Every parameter you list must be populated on the `LLMTestCase` you measure. If `EXPECTED_OUTPUT` is in `evaluation_params` and a case leaves `expected_output` as `None`, measuring that case errors rather than silently scoring it. This is deliberate — a half-populated judge prompt would produce a number that means nothing — but it does mean a dataset with patchy fields will break a run partway through. Validate the dataset before the suite, not during it. ## Reference-dependence, and where it bites Deciding whether `EXPECTED_OUTPUT` is in the list is the same decision as deciding whether the metric is reference-based: - **With `EXPECTED_OUTPUT`** the metric compares against ground truth. It is sharp and it catches factual drift, but it only runs where ground truth exists: a curated dataset with human-written answers. - **Without it** the metric judges the output on its own terms — is it relevant, is it grounded in the supplied context, does it follow the policy. This kind can run over anything, including outputs captured from production, because it needs no answer key. Teams hit this when they try to point their offline suite at real traffic. Half the metrics run and half of them error or would need a reference nobody wrote. The practical answer is to build two families of custom metrics: reference-based ones for the regression dataset, and reference-free ones (grounded in `RETRIEVAL_CONTEXT` and `INPUT`) for sampled live traffic. ## Writing steps and params together A useful discipline is to phrase evaluation steps using the field names themselves — "compare 'actual output' with 'expected output'" — because it makes the dependency obvious at review time. If a step mentions a field that is not in `evaluation_params`, that is a bug you can spot by reading, without running anything. ## Quick checklist - Do the steps mention a field the judge cannot see? Add it or reword the step. - Is a field listed that no step uses? Remove it — tokens and unwanted influence. - Is `EXPECTED_OUTPUT` listed? Then this metric is dataset-only. - Does every case in the dataset populate all listed fields? If not, the run errors.
- What happens if a listed parameter is missing on a test case at measure time?DeepEval raises rather than scoring it. A parameter in evaluation_params is a contract: the judge prompt has a slot for that field, and filling it with nothing would produce a score with no basis. Practically this means dataset validation belongs before the run — sweep the goldens for empty expected_output or retrieval_context first, so you fail fast on the dataset instead of halfway through a paid evaluation.
- Which parameters would you pick for a metric that scores whether an answer stays grounded in retrieved passages?INPUT, ACTUAL_OUTPUT and RETRIEVAL_CONTEXT. The judge needs the question to know what was asked, the answer to inspect, and the retrieved passages to check each claim against. Deliberately leave EXPECTED_OUTPUT out: groundedness is about support from the passages, not agreement with a reference answer, and omitting it keeps the metric runnable on production traffic where no reference exists.
- Is there a cost to listing more parameters than the rubric needs?Two costs. Every listed field is rendered into the judge prompt, so you pay input tokens for it on every single sample — noticeable across thousands of cases. And extra context influences the judge: showing a reference answer to a tone metric invites it to penalise wording differences the rubric never mentioned. List what the steps actually reference.
saying these in an interview costs you the question
- Assumes the judge sees the whole LLMTestCase regardless of evaluation_params
- Writes rubric steps about context that was never passed as a parameter
- Thinks a missing listed field is silently skipped instead of raising
- Treats evaluation_params as required boilerplate with no effect on the score
- Plans to run an EXPECTED_OUTPUT metric over production traffic with no ground truth