skip to content

How does the rubric argument change scoring in DeepEval's GEval metric?

level: middleimportance: should knowfreq 40%

answer

  1. anchor the scale, don't leave it implicit
  2. bands plus prose describing each band
  3. written 0-10, reported 0-1
  4. describe observable properties, not adjectives
  5. legible boundary cases, still not deterministic

basics

~20 s

rubric takes a list of Rubric objects, each pairing a score_range with an expected_outcome description. It tells the judge what each band of scores means, so scores land on defined levels instead of on the judge's private sense of what 0.7 is worth.

solid answer

~50 s

Without a rubric, a `GEval` judge picks a number using whatever internal calibration it has, and two similar outputs can land a band apart for no articulable reason. `rubric` fixes that by handing the judge an explicit scale: a list of `Rubric` entries, each with a `score_range` and an `expected_outcome` string describing what an output in that band looks like. ```python from deepeval.metrics.g_eval import Rubric rubric=[ Rubric(score_range=(0, 2), expected_outcome="Contradicts the reference on a material fact."), Rubric(score_range=(3, 6), expected_outcome="Broadly correct but omits a required detail."), Rubric(score_range=(7, 9), expected_outcome="Correct, with only cosmetic differences."), Rubric(score_range=(10, 10), expected_outcome="Fully correct and complete."), ] ``` The bands are written on a 0-10 scale; the metric's final score is still normalized into 0-1 and compared with `threshold`. The practical payoff is that a score becomes interpretable — 0.3 now means "omitted a required detail", not "the judge felt lukewarm".

code

python · 19 lines
python
from deepeval.metrics import GEval
from deepeval.metrics.g_eval import Rubric
from deepeval.test_case import LLMTestCaseParams

correctness = GEval(
    name="Correctness",
    evaluation_steps=[
        "Compare each factual claim in 'actual output' with 'expected output'.",
        "Note any contradiction, and separately note any required detail that is missing.",
    ],
    rubric=[
        Rubric(score_range=(0, 2), expected_outcome="Contradicts the reference on a material fact."),
        Rubric(score_range=(3, 6), expected_outcome="No contradiction, but a required detail is missing."),
        Rubric(score_range=(7, 9), expected_outcome="Complete and correct, with cosmetic wording differences."),
        Rubric(score_range=(10, 10), expected_outcome="Fully correct, complete and precisely phrased."),
    ],
    evaluation_params=[LLMTestCaseParams.INPUT, LLMTestCaseParams.ACTUAL_OUTPUT, LLMTestCaseParams.EXPECTED_OUTPUT],
    threshold=0.7,
)

go deeper

for a junior

Know that rubric takes Rubric objects pairing a score_range with an expected_outcome description, and that it tells the judge what each score band should mean instead of leaving the scale implicit.

for a middle

Explain that bands are written on a 0-10 scale while the metric still reports 0-1 against threshold, and that band text must describe observable properties rather than restating the number as an adjective.

for a senior

Show how you tune a rubric in practice: find the cases that move between runs, recognise them as boundary cases, and rewrite the band description that failed to decide them. Pair the threshold with a band boundary so passing means something describable.

for a principal

Treat rubric text as a reviewable artefact that domain experts sign off on, since it encodes what the organisation considers acceptable output. Own how rubric edits invalidate historical score comparisons and when a re-baseline is required.

## The problem rubrics solve Ask a model to "score this answer from 0 to 1" and it will comply, but the number is anchored to nothing. Its notion of 0.6 is not stable across cases, across prompt phrasings, or across model versions. You see this as scores that cluster oddly, as a suite whose average shifts when you change nothing important, and — worst — as an inability to explain to a colleague what a given score means. `rubric` is the fix inside the tool: instead of leaving the scale implicit, you define the bands and describe each one. ## The shape of it `GEval` accepts `rubric` as a list of `Rubric` objects, imported from `deepeval.metrics.g_eval`. Each carries: - **`score_range`** — an inclusive integer band on a 0-10 scale, for example `(3, 6)`. - **`expected_outcome`** — prose describing the output that belongs in that band. The bands should tile the whole 0-10 range without gaps or overlaps, or the judge has to decide where an in-between output goes, which reintroduces exactly the ambiguity you were removing. The metric's reported `score` remains a value between 0 and 1, as with every DeepEval metric, and it is that normalized value that is compared with `threshold`. So a rubric whose passing band starts at 7 out of 10 pairs naturally with `threshold=0.7`. ## Writing bands that work The descriptions carry the weight. Two rules make them useful: **Describe observable properties, not degrees of goodness.** "Excellent", "good", "fair" restates the number and teaches the judge nothing. "Cites the correct statute but paraphrases the clause" is a test the judge can apply. **Make the boundaries mechanical.** The most valuable sentence in a rubric is usually the one that says what pushes an output *down* a band: "any contradiction of the reference caps the score at 2, regardless of other quality". That gives the judge a rule rather than an impression, and it means a reviewer can check the judge's work. Four or five bands is usually right. Ten bands invites the judge to split hairs it cannot split reliably; two bands is `strict_mode` with extra ceremony. ## What it does and does not buy you It buys **interpretability** — a score maps to a description a human wrote — and **tighter clustering**, because the judge is choosing among described levels rather than picking a real number freehand. It also buys **reviewability**: rubric bands are text in your repo, so a domain expert who does not read Python can audit what the metric rewards. It does not buy determinism. The judge call is still sampled, and a case that sits on a band boundary can land either side of it on a rerun. Bands make that failure mode *legible* — you can look at the boundary case and decide whether your band description is ambiguous — but they do not eliminate it. Whether the resulting scores agree with human judgement at all is a separate question, answered by checking the metric against labelled examples rather than by any argument about configuration. ## How it composes with the other knobs - With **`evaluation_steps`**: steps say *how to look*, the rubric says *how to score what you saw*. They complement each other, and a metric that a pipeline depends on usually has both. - With **`threshold`**: choose the threshold to match a band boundary, so "passing" means a describable level of quality rather than an arbitrary cut. - With **`strict_mode`**: the two pull in opposite directions. Strict mode collapses everything to 1 or 0, which discards the graded scale a rubric exists to create. Use one or the other. ## A worked read Suppose your correctness metric with the rubric above returns 0.6 on a case. Because the bands are defined, you know immediately that the judge placed it in the 3-6 band — broadly correct, missing a required detail — and the `reason` should name the missing detail. That is a bug report. Without the rubric, 0.6 is a mood, and the next conversation is about whether the metric is any good rather than about the output that failed.

  • If the rubric bands are written on a 0-10 scale, what does the metric actually report?
    A score between 0 and 1, like every DeepEval metric, and that normalized value is what gets compared with threshold and drives is_successful(). The 0-10 range exists inside the rubric definition because integer bands are easier to write and to reason about than fractions. In practice you pick a threshold that lines up with a band boundary — a passing band starting at 7 pairs with threshold=0.7.
  • How would you write bands that reduce disagreement between reruns?
    Make each band a test rather than an adjective, and state explicitly what caps a score. "Any contradiction of the reference caps at 2" gives the judge a rule; "fair quality" gives it a feeling. Then look at the cases that moved between runs — a case that flips band is usually sitting on a boundary your description left ambiguous, and rewriting that one sentence does more than changing judge models.
  • Would you use rubric and strict_mode together?
    No, they work against each other. A rubric exists to create a graded, interpretable scale; strict_mode collapses the result to 1 or 0 and forces the threshold to 1, throwing that grading away. Pick the shape that matches the criterion: graded rubric for quality dimensions, strict binary for hard rules that have no middle.
  • How many bands should a rubric have?
    Usually four or five. Too few and you are approximating a binary metric with extra ceremony; too many and you ask the judge to draw distinctions it cannot draw reliably, so adjacent bands blur and reruns wander. The right count is the number of genuinely distinguishable levels a human reviewer could also agree on for your task.

saying these in an interview costs you the question

  • Uses adjectives like good or fair as band descriptions
  • Thinks a rubric makes GEval scores deterministic
  • Expects the final metric score to be reported on the 0-10 band scale
  • Leaves gaps or overlaps between score ranges
  • Combines rubric with strict_mode and expects graded output

context