skip to content

In DeepEval's GEval, when do you pass evaluation_steps instead of criteria?

level: middleimportance: must knowfreq 70%

answer

  1. one describes, the other prescribes
  2. an extra model call at construction
  3. rubric in your repo vs rubric invented
  4. supply one, not both
  5. pinning removes rubric drift, not judge variance

basics

~20 s

criteria is a sentence GEval expands into judging steps with an extra model call each time the metric is built. evaluation_steps hands GEval that list yourself, so every run applies the identical rubric. Pass one or the other, never both.

solid answer

~50 s

`GEval` scores an output by having a judge model follow a list of evaluation steps. If you give it `criteria` — a plain-sentence description like "determine whether the actual output is factually correct given the expected output" — DeepEval first asks the judge to expand that sentence into concrete steps, then scores against them. If you give it `evaluation_steps` — an explicit list of strings — that generation call is skipped and the judge follows exactly what you wrote. DeepEval expects one of the two, not both. I reach for `criteria` while exploring, because writing one sentence is fast and the generated steps are a good first draft. I switch to `evaluation_steps` the moment the metric is something a pipeline depends on: the rubric stops being regenerated, one judge call per sample disappears, and score movements become attributable to the model under test rather than to a rubric that quietly rewrote itself.

code

python · 29 lines
python
from deepeval.metrics import GEval
from deepeval.test_case import LLMTestCase, LLMTestCaseParams

# Exploration: one sentence, expanded into steps by the judge model.
draft = GEval(
    name="Correctness",
    criteria="Determine whether the actual output is factually correct given the expected output.",
    evaluation_params=[LLMTestCaseParams.INPUT, LLMTestCaseParams.ACTUAL_OUTPUT, LLMTestCaseParams.EXPECTED_OUTPUT],
)

# Production: the same rubric, pinned in the repo.
pinned = GEval(
    name="Correctness",
    evaluation_steps=[
        "Check whether every factual claim in 'actual output' is supported by 'expected output'.",
        "Penalise contradictions heavily; penalise omitted details lightly.",
        "Ignore differences in wording, ordering and formatting.",
    ],
    evaluation_params=[LLMTestCaseParams.INPUT, LLMTestCaseParams.ACTUAL_OUTPUT, LLMTestCaseParams.EXPECTED_OUTPUT],
    threshold=0.7,
)

test_case = LLMTestCase(
    input="When was the Eiffel Tower completed?",
    actual_output="It was finished in 1889.",
    expected_output="The Eiffel Tower was completed in 1889.",
)
pinned.measure(test_case)
print(pinned.score, pinned.reason)

go deeper

for a junior

Know that GEval judges an output using a rubric written in plain English, and that you supply that rubric either as a one-line criteria string or as an explicit list of evaluation_steps.

for a middle

Explain that criteria is expanded into steps by an extra judge call while evaluation_steps skips it, and that DeepEval takes one or the other. Say why a rubric that regenerates is a problem for comparing runs.

for a senior

Show the workflow: prototype with criteria and verbose_mode, read the generated steps against hand-labelled cases, then pin the edited list. Be clear that pinning removes rubric drift but not judge sampling variance, and that editing steps re-baselines the metric.

for a principal

Own the policy: which metrics are allowed to use free-form criteria, which must ship pinned steps reviewed like code, and what happens to historical score comparisons when a rubric is edited. Decide where step lists live and who signs off on changes.

## What GEval is `GEval` is DeepEval's general-purpose custom metric. Instead of a fixed algorithm, you describe in natural language what "good" means for your task, and a judge model applies that description to each `LLMTestCase`, returning a score between 0 and 1 plus a written reason. It is the escape hatch you use when no built-in metric covers the thing you actually care about — "does this reply follow our refund policy", "is the tone appropriate for a bereavement email", "did it cite a statute number when it made a legal claim". ```python from deepeval.metrics import GEval from deepeval.test_case import LLMTestCase, LLMTestCaseParams ``` ## The two ways to describe the rubric `GEval` takes either `criteria` or `evaluation_steps`, and DeepEval expects exactly one of them. **`criteria`** is a single string: "Determine whether the actual output is factually correct based on the expected output." When you construct the metric with only `criteria`, DeepEval performs a chain-of-thought expansion — it asks the judge model to turn that sentence into an ordered list of evaluation steps, and those generated steps are what the judge then applies when scoring each test case. **`evaluation_steps`** is a list of strings you author: ```python evaluation_steps=[ "Check whether every factual claim in 'actual output' appears in 'expected output'.", "Penalise contradictions heavily; penalise omissions lightly.", "Ignore differences in wording, ordering and formatting.", ] ``` Supplying them skips the expansion entirely. ## Why the difference matters **Repeatability.** A generated step list is itself model output. Regenerate it — new process, new SDK version, a judge model that was silently upgraded behind the same alias — and you may get a subtly different rubric. Your scores then move for a reason that has nothing to do with the system you are testing. Pinned `evaluation_steps` are text in your repo: reviewable in a diff, versioned with the prompt they judge, and identical on every machine. **Cost and latency.** The expansion is an extra model call. It happens per metric instance, not per test case, so on a 500-case suite it is noise — but in a service that constructs metrics per request, or a parametrized suite that builds a fresh metric for every case, it multiplies. **Precision of the rubric.** A one-line criterion leaves the judge to invent the details, including how to weigh partial credit. When you find that your metric is punishing something you do not care about — say, marking a correct answer down for being terse — the fix is to write the step that says so. You cannot edit a step list you never wrote. **Debuggability.** When someone asks why a case scored 0.4, "here are the seven steps the judge followed" is an answer. "The model wrote some steps from this sentence" is not. ## The practical workflow 1. Start with `criteria` and `verbose_mode=True` so DeepEval prints the intermediate reasoning to the console. 2. Read the steps the judge is actually following, on a handful of cases you have hand-labelled. 3. Copy those steps out, edit them where they disagree with your judgement, and pass them as `evaluation_steps`. 4. From then on, treat the step list like any other prompt asset — changing it invalidates comparisons with earlier runs, so re-baseline when you edit it. That last point is the one candidates miss. Pinned steps do not make a G-Eval score deterministic; the judge call is still sampled, and the same output can score 0.7 one run and 0.6 the next. Pinning removes *one* source of drift — the rubric — so that the remaining variance is the judge's, which you can attack separately with a rubric, a stricter step list, or a more capable judge model. Whether the resulting number is trustworthy at all, and how much of a delta counts as a real regression, is a question about evaluation methodology rather than about configuring the metric. ## Things that trip people up - Passing both arguments is a configuration error, not a merge. - `name` is required and is only a label — it does not influence scoring. - `evaluation_params` is separate and mandatory: steps that talk about retrieval context are meaningless if you never showed the judge the retrieval context. - Rewriting the steps changes the metric. Old scores are not comparable to new ones.

  • If I pin evaluation_steps, are my GEval scores now deterministic?
    No. Pinning removes one source of drift — the rubric no longer changes between runs — but the scoring call itself is still a sampled model generation, so the same output can score differently on repeat runs. To narrow that spread you add a rubric that anchors what each band means, write sharper steps, or use a stronger judge model. Determinism is not on offer; reduced variance is.
  • How would you get from a rough criteria string to a step list you trust?
    Run the metric with criteria and verbose_mode=True over a small set of cases you have already labelled by hand. The console output shows the steps and the reasoning behind each score. Where the judge disagrees with your labels, look at which step produced the disagreement, rewrite it, and promote the edited list to evaluation_steps. That is a normal prompt-iteration loop, done against known answers.
  • Does the expansion happen once per metric or once per test case?
    Once per metric instance, when it expands the criteria — not once per test case. So a suite that builds one metric and measures 500 cases pays it once. A suite that constructs a fresh GEval inside a parametrized test, or a service that builds one per request, pays it every time, which is a good reason to hoist the metric out or pin the steps.

saying these in an interview costs you the question

  • Thinks criteria and evaluation_steps can be passed together and are merged
  • Believes pinned evaluation_steps make GEval scores deterministic
  • Assumes the criteria expansion runs once per test case
  • Thinks the name argument affects how the judge scores
  • Cannot say that criteria triggers an extra model call at all

context