skip to content

How do you write a custom DeepEval metric by subclassing BaseMetric?

level: seniorimportance: should knowfreq 30%

answer

  1. subclass, don't judge, when code can decide
  2. measure, a_measure, is_successful
  3. results live on self.score and self.success
  4. a property gives the metric its label
  5. normalise into the 0-1 scale everyone expects

basics

~10 s

Subclass BaseMetric and implement measure(test_case), the async a_measure, and is_successful(), setting self.score and self.success inside measure (plus self.reason if you want an explanation). Expose a name property so the metric is labelled in reports.

solid answer

~50 s

`GEval` covers criteria you can only describe in prose. When the check is mechanical — latency under a budget, valid JSON against a schema, a required disclaimer present, a regulated term absent — you want no judge at all, and DeepEval lets you supply one by subclassing `BaseMetric` from `deepeval.metrics`. The contract is small. Implement `measure(self, test_case)` to compute the score, set `self.score` and `self.success` on the instance, optionally set `self.reason`, and return the score. Implement `a_measure` as the async counterpart — when there is no real async work it can just call `measure`. Implement `is_successful()` to report the verdict. Add a `__name__` property so the metric shows a readable label in the report. Once that exists it is a first-class metric: it goes in the same metric list as built-ins and judged metrics, and it costs nothing per sample and never flakes.

code

python · 32 lines
python
import json
from deepeval.metrics import BaseMetric
from deepeval.test_case import LLMTestCase


class ValidJsonMetric(BaseMetric):
    def __init__(self, required_keys: set[str], threshold: float = 1.0):
        self.required_keys = required_keys
        self.threshold = threshold

    def measure(self, test_case: LLMTestCase) -> float:
        try:
            payload = json.loads(test_case.actual_output)
        except json.JSONDecodeError as exc:
            self.score = 0.0
            self.reason = f"Output is not valid JSON: {exc.msg}"
        else:
            missing = self.required_keys - set(payload)
            self.score = 0.0 if missing else 1.0
            self.reason = f"Missing keys: {sorted(missing)}" if missing else "All required keys present."
        self.success = self.score >= self.threshold
        return self.score

    async def a_measure(self, test_case: LLMTestCase, *args, **kwargs) -> float:
        return self.measure(test_case)

    def is_successful(self) -> bool:
        return self.success

    @property
    def __name__(self) -> str:
        return "Valid JSON"

go deeper

for a junior

Know that DeepEval lets you write your own metric by subclassing BaseMetric, and that the class must implement measure, its async a_measure counterpart, and is_successful.

for a middle

Explain that measure sets self.score and self.success on the instance and optionally self.reason, that name is a property giving the report label, and that the score belongs on the 0-1 scale threshold uses.

for a senior

Make the case for coding a check rather than judging it: no per-sample cost, no sampling variance, so the suite can run on every change. Cover error handling — fail with a score and reason, never an exception — and instance state under concurrency.

for a principal

Own where the line sits between coded and judged metrics across the portfolio, since it directly sets the evaluation budget and the flakiness floor of the gate. Push mechanically decidable criteria into code so judged calls are spent only on genuine judgement.

## Why bother Every judged metric costs a model call per sample and carries sampling variance. For anything decidable in code, that is a bad trade: you pay money and flakiness for a question Python answers exactly. Typical candidates are structural or policy checks — schema validity, presence of a required citation or disclaimer, absence of a blocklisted term, output length inside a bound, latency or cost inside a budget, an exact-match or regex check against expected output. DeepEval's answer is `BaseMetric`. Subclassing it gets you into the same machinery as everything else: your metric can sit in the same metrics list as an `AnswerRelevancyMetric` and a `GEval`, and its result is reported alongside theirs. ## The contract ```python from deepeval.metrics import BaseMetric from deepeval.test_case import LLMTestCase class ValidJsonMetric(BaseMetric): def __init__(self, threshold: float = 1.0): self.threshold = threshold def measure(self, test_case: LLMTestCase) -> float: try: json.loads(test_case.actual_output) self.score = 1.0 self.reason = "Output parsed as JSON." except json.JSONDecodeError as exc: self.score = 0.0 self.reason = f"Not valid JSON: {exc.msg}" self.success = self.score >= self.threshold return self.score async def a_measure(self, test_case: LLMTestCase, *args, **kwargs) -> float: return self.measure(test_case) def is_successful(self) -> bool: return self.success @property def __name__(self) -> str: return "Valid JSON" ``` The pieces: - **`measure`** does the work. It must set `self.score` and `self.success` on the instance — the framework reads those attributes afterwards — and returning the score is the convention. - **`a_measure`** is the async entry point. DeepEval can evaluate metrics concurrently, and it needs an awaitable path. If your metric does no I/O, delegating to `measure` is fine; if it calls something over the network, do the real async work here so concurrency actually helps. - **`is_successful`** reports the verdict. Keep it a pure read of state computed in `measure`, not a second computation that could disagree. - **`__name__`** as a property gives the metric its label in reports. Without it you get a class-shaped name that is harder to read in results. - **`self.reason`** is optional but worth setting. A failing metric with no explanation forces whoever reads the report to reproduce the check by hand. ## Conventions worth honouring **Score in 0-1.** Every other metric reports on that scale and `threshold` is compared against it, so a metric returning milliseconds or a token count breaks every downstream assumption. Normalise: for a latency budget, score 1 when inside and 0 outside, or map the ratio into 0-1 if you want a gradient. **Never raise from `measure` for an expected failure.** A malformed output is a score of 0 with a reason, not an exception — an exception aborts a case rather than recording it as failing, which is the opposite of what a test suite wants. Reserve raising for genuine programming errors. **Keep state per-instance and be careful about reuse.** Because results live on `self`, a metric instance measured concurrently over several cases can have its attributes overwritten. Construct fresh instances per case, or make the metric stateless enough that this cannot bite. **Take configuration in `__init__`.** Thresholds, blocklists, schemas and budgets belong as constructor arguments so the same class serves several checks. ## Where a custom metric wins outright Against a judged metric, a coded one is free per sample, instant, and identical on every run. On a suite of a few thousand cases that is the difference between an evaluation you run on every change and one you run nightly because of the bill. The discipline that follows: for every criterion, ask whether code can decide it. If yes, write a `BaseMetric`. If only prose can express it, use `GEval`. If it is a mix of gates and judgement, a DAG lets you compose both. A coded metric is not automatically a *good* metric — it can encode the wrong rule perfectly, and it can only test what is mechanically checkable. But its failure modes are ordinary software failure modes, which your team already knows how to find and fix.

  • Why does BaseMetric require an async a_measure as well as measure?
    DeepEval can evaluate metrics concurrently, which needs an awaitable entry point per metric. For a purely computational check there is no real async work, so a_measure delegating to measure satisfies the contract without pretending. If your metric calls something over the network — a classifier service, a moderation API — implement the genuine async path there, otherwise concurrency gains nothing because you block the loop.
  • Should measure() raise when the output is malformed?
    No. A malformed output is precisely the failure the metric exists to catch, so it should record score 0 with a reason explaining what was wrong. Raising aborts that test case rather than reporting it as failed, so the run either dies or reports an error that is harder to triage than a clean failure. Reserve exceptions for real programming errors, like a missing configuration.
  • What scale should a custom metric's score use?
    0 to 1, because that is what threshold is compared against and what every other DeepEval metric reports. A metric returning raw milliseconds or a token count silently breaks threshold semantics and any aggregate someone computes across metrics. Normalise deliberately — binary 1/0 for a pass-fail rule, or a mapped ratio if a gradient is meaningful for your check.
  • When would you still prefer GEval over writing a coded metric?
    When the criterion cannot be reduced to a decidable rule: tone, helpfulness, whether an explanation is understandable to a non-expert, whether a refusal was appropriately worded. Code cannot check those without becoming a bad proxy, and a bad proxy that is perfectly stable is worse than a judged score that is roughly right. The rule of thumb is: code what is decidable, judge what is not.

saying these in an interview costs you the question

  • Only implements measure and omits a_measure and is_successful
  • Returns raw latency or token counts instead of a 0-1 score
  • Raises an exception for an output the metric is meant to fail
  • Recomputes the verdict inside is_successful instead of reading state
  • Reuses one metric instance across concurrent cases and reads stale attributes

context