skip to content

Which DeepEval built-in metrics work on production traces with no expected_output?

level: seniorimportance: should knowfreq 50%

answer

  1. live traffic has no reference answer
  2. two lists, split by one field
  3. precision and recall are the blocked pair
  4. hallucination is blocked too, for context
  5. label a sample to promote cases

basics

~10 s

AnswerRelevancyMetric, FaithfulnessMetric, ContextualRelevancyMetric, ToxicityMetric and BiasMetric all run without a reference. ContextualPrecisionMetric and ContextualRecallMetric need expected_output, and HallucinationMetric needs human-supplied context, so none of those three can score raw traffic.

solid answer

~40 s

Split the built-ins by what they compare against. Reference-free metrics compare the output to the input or to what was retrieved, so a case rebuilt from a trace — `input`, `actual_output`, `retrieval_context` — is enough for `AnswerRelevancyMetric`, `FaithfulnessMetric`, `ContextualRelevancyMetric`, `ToxicityMetric` and `BiasMetric`. Reference-requiring metrics need something a human wrote: `ContextualPrecisionMetric` and `ContextualRecallMetric` need `expected_output`, and `HallucinationMetric` needs `context`. In practice that produces two suites, not one. A small curated dataset carries references and runs the full metric set as a pre-merge gate; a sampled slice of production traffic runs only the reference-free metrics, continuously. Trying to force the reference-requiring metrics onto live traffic means either fabricating an `expected_output` — which makes the metric grade agreement with a guess — or paying for human labelling on every case.

code

python · 24 lines
python
from deepeval.metrics import (
    AnswerRelevancyMetric,
    FaithfulnessMetric,
    ContextualRelevancyMetric,
    ToxicityMetric,
)
from deepeval.test_case import LLMTestCase

trace = {"question": "...", "answer": "...", "chunks": ["..."]}

case = LLMTestCase(
    input=trace["question"],
    actual_output=trace["answer"],
    retrieval_context=trace["chunks"],
)

for metric in (
    AnswerRelevancyMetric(model="gpt-4o-mini", include_reason=False),
    FaithfulnessMetric(model="gpt-4o-mini", include_reason=False),
    ContextualRelevancyMetric(model="gpt-4o-mini", include_reason=False),
    ToxicityMetric(threshold=0.0),
):
    metric.measure(case)
    print(type(metric).__name__, metric.score)

go deeper

for a junior

Know that some metrics need a human-written reference answer and therefore cannot run on live traffic, and be able to name answer relevancy and toxicity as ones that can.

for a middle

Sort the built-ins into reference-free and reference-requiring from memory, and explain that the split follows from the required LLMTestCase fields rather than from metric quality.

for a senior

Describe the two-suite setup — full metrics on a small curated set as a gate, reference-free metrics on sampled traffic continuously — and reject the fabricated-reference shortcut or scope it explicitly to screening.

for a principal

Own the economics and the data policy: sampling rate times metrics times judge calls is a recurring bill, human labelling is the only way to grow reference coverage, and in a regulated setting whether traces may leave the system decides the whole design.

## The constraint Production traffic has no ground truth. Nobody wrote down what the right answer to a live user's question was, and by the time you know, the conversation is over. So the question of which DeepEval metrics you can point at real traffic is answered entirely by which ones need a human-authored field. ## The two lists **Runs on a trace-derived case** (`input`, `actual_output`, and `retrieval_context` if you traced the retriever): - `AnswerRelevancyMetric` — compares the answer to the question. - `FaithfulnessMetric` — compares the answer's claims to what was retrieved. - `ContextualRelevancyMetric` — asks how much of what was retrieved was on topic. - `ToxicityMetric`, `BiasMetric` — judge the output text on its own. **Needs a human-authored field**: - `ContextualPrecisionMetric` — needs `expected_output` to define which chunks count as relevant for ranking. - `ContextualRecallMetric` — needs `expected_output` to check whether its content was recoverable from retrieval. - `HallucinationMetric` — needs `context`, the ground-truth background. ## Why you cannot fake the reference Two shortcuts get proposed and both are traps. **Using the model's own output as the reference.** This scores perfectly by construction and measures nothing. **Generating `expected_output` with a stronger model.** This is more seductive because it produces plausible numbers. But the metric is then measuring agreement between your production model and the reference model, on a task where the reference model's answer has never been checked. Where they disagree, you cannot tell which one was wrong, and any systematic bias in the reference model becomes an invisible bias in your score. If you do this, be honest about it in the interview: it is a *screening* signal for finding candidate failures to look at, not a quality measurement, and you should not gate a deploy on it. The defensible path is labelling: sample traces, have a human write `expected_output` (or `context`) for the sampled subset, and promote those into a curated dataset. That converts a cheap continuous signal into an expensive durable one, and it is why the curated set stays small. ## The shape a real setup ends up with 1. **Curated set, full metrics, pre-merge.** A few dozen to a few hundred cases carrying `expected_output` and `context`. Every metric applies, thresholds gate the change, and the cost is bounded because the set is small and fixed. 2. **Sampled production, reference-free metrics, continuous.** A percentage of live traces rebuilt into `LLMTestCase` objects. Faithfulness and contextual relevancy tell you whether retrieval and grounding are degrading; toxicity and bias tell you whether anything unacceptable is going out. No gate — these feed dashboards and alerts. 3. **A promotion path.** Cases that score badly on the reference-free metrics are the best candidates for human labelling, because they are where the system is already suspected of being wrong. That is how the curated set grows without anyone labelling randomly. ## The cost dimension Every one of these metrics is an LLM judge, so pointing five of them at a meaningful sample of production is a real recurring bill: number of sampled traces times metrics times several judge calls each. Sampling rate is therefore a budget decision, and it is normal to sample a low single-digit percentage, use a cheaper judge model on the production path than on the pre-merge gate, and turn off reason generation where nobody reads the reasons. Reserving the expensive configuration for the small curated set is the same instinct. ## The privacy dimension Rebuilding test cases from traces means copying prompts and completions — often user content — into your evaluation store, and then sending them to a judge model. In a regulated environment that is exactly the data you may not be allowed to move. This frequently ends up being the binding constraint rather than cost, and mentioning it unprompted is what marks the answer as coming from someone who has shipped this. ## What good sounds like Name the two lists, explain that the split is caused by the reference field and not by metric quality, refuse the fabricated-reference shortcut or scope it honestly to screening, and describe the two-suite shape with a labelling promotion path between them.

  • A teammate proposes generating expected_output with a stronger model so ContextualRecallMetric can run on live traffic. What do you say?
    That it measures agreement with an unverified reference, not correctness. Where the two models disagree you cannot tell which is wrong, and the reference model's blind spots become invisible in your score. It is usable as a screening signal to surface cases worth a human look, but I would not gate a deploy on it or report it as a quality number.
  • How do you choose what fraction of production traffic to score?
    By budget and by what you need to detect. Each sampled case costs several judge calls per metric, so cost scales linearly with the rate. I would sample a low single-digit percentage continuously for trend detection, stratify so rare intents are not lost, and use a cheaper judge with reasons disabled on that path while keeping the expensive configuration for the small curated set.
  • What stops some teams from evaluating production traces at all?
    Data controls. Building test cases from traces copies user prompts and completions into an evaluation store and then sends them to a judge model, frequently a third-party API. In regulated settings that movement is not permitted, so the options become self-hosted judges, redaction before the case is built, or restricting evaluation to synthetic and curated data.

saying these in an interview costs you the question

  • Claiming every DeepEval metric runs on production traffic
  • Using the model's own output as expected_output
  • Treating a model-generated reference as ground truth
  • Forgetting HallucinationMetric also needs a human field
  • Ignoring judge cost when scaling to sampled live traffic

context