How many judge LLM calls does a DeepEval metric suite make per test case?
answer
- not one call per case
- decompose, verdict, explain
- metrics do not share work
- cases times metrics times calls
- cheap judge on volume, strong judge on the gate
basics
~20 sMore than one per metric. Each built-in metric runs its own chain of judge prompts — extracting statements, judging them, and with include_reason=True writing an explanation — so cost scales as cases times metrics times several calls, not as one call per case.
solid answer
~50 sThe mental model that gets teams into trouble is "one case, one call". In reality each stock metric is an LLM pipeline: `AnswerRelevancyMetric`, for example, has the judge extract statements from `actual_output`, classify each against the `input`, and then — if `include_reason` is on — generate a written justification. Attach five metrics to a case and you have five independent pipelines, each with several requests and none sharing work with the others. So the levers are: how many metrics you attach, the `model` you give each metric (a cheaper judge on the high-volume path, a stronger one on the gate), whether `include_reason` is worth its extra generation for metrics nobody reads reasons from, `async_mode` to overlap a metric's internal calls so wall-clock stops tracking cost, and the size of the dataset. Estimate before you scale: a few thousand cases times five metrics is a bill worth approving deliberately rather than discovering.
code
python · 13 linesfrom deepeval.metrics import AnswerRelevancyMetric, FaithfulnessMetric
# High-volume path: cheap judge, no reason generation.
volume = [
AnswerRelevancyMetric(model="gpt-4o-mini", include_reason=False, async_mode=True),
FaithfulnessMetric(model="gpt-4o-mini", include_reason=False, async_mode=True),
]
# Pre-merge gate: stronger judge, reasons kept so a red build explains itself.
gate = [
AnswerRelevancyMetric(model="gpt-4o", include_reason=True, threshold=0.8),
FaithfulnessMetric(model="gpt-4o", include_reason=True, threshold=0.9),
]go deeper
Know that DeepEval's built-in metrics call an LLM to score, so running a suite costs money and takes seconds per case rather than milliseconds.
Explain the decompose-verdict-explain chain inside a metric, and that cost scales as cases times metrics times several calls, with include_reason adding a generation each.
Demonstrate the levers you would actually pull: drop metrics that answer a question nobody asked this run, pick judge models per path, disable reasons on high-volume paths, and pilot on ten cases before scaling to ten thousand.
Own evaluation as a standing line item — judge choice, metric list and sampling rate together set a recurring bill and the team's feedback latency, and both need an owner and a budget rather than accumulating by default.
## Where the calls come from A DeepEval metric is not a scoring function, it is a small prompt chain. The pattern across the LLM-judged built-ins is consistent: 1. **Decompose.** Ask the judge to break something into units — statements in the `actual_output`, claims to verify, truths present in the `retrieval_context`. 2. **Verdict.** Ask the judge to rule on each unit: is this statement relevant to the input, is this claim supported, does this contradict the context. 3. **Explain.** If `include_reason=True`, spend one more generation writing the sentence that ends up in `metric.reason`. The score is then arithmetic over the verdicts. Nothing here is a single request, and nothing is shared between metrics: `FaithfulnessMetric` and `ContextualRelevancyMetric` both read the same `retrieval_context` and both re-tokenise it in their own prompts. That gives the cost shape: cases times metrics times several calls each times tokens per call. The token term matters as much as the call count for RAG cases, because `retrieval_context` is often thousands of tokens and gets resent in every prompt of every metric that reads it. ## The levers, roughly in order of effect **Number of metrics.** The cheapest saving is the metric you did not attach. Teams routinely run all five RAG-ish metrics on every case out of completeness, when contextual precision and recall answer a retriever question that only needs to be asked when retrieval changes. Splitting a suite into "generation metrics on every run" and "retrieval metrics when the retriever or index changes" often halves the bill without losing coverage. **Judge model.** `model` on each metric constructor decides the per-call price, and the gap between a frontier judge and a small one is large. Nothing forces one judge across the suite: it is entirely reasonable to run a cheap judge over sampled production traffic for trend detection and a stronger judge over the small curated gate, since the gate is where a wrong verdict blocks a deploy. **`include_reason`.** Reasons are the first thing to turn off on high-volume paths and the last thing to turn off on a gate. On a gate, the reason is how an engineer understands a red build without re-running anything; on a dashboard aggregating thousands of scores, nobody opens them. It costs a generation per metric per case. **Dataset size and sampling.** Cost is linear in cases, so the sampling rate on production traffic is a direct budget dial. Stratify the sample rather than taking it uniformly, so rare intents are still represented at a low rate. **`async_mode`.** This one buys latency rather than money: it lets a metric's internal judge calls overlap instead of running strictly in sequence. Worth knowing because a suite that is merely slow is often assumed to be expensive, and the two are separate problems with separate fixes. ## Estimating before you scale The useful discipline is to measure one case and multiply. Run a single test case with your intended metric set against your intended judge model, look at what it cost and how long it took, then multiply by the dataset size. Doing this on a ten-case pilot before pointing the suite at ten thousand cases takes minutes and routinely changes the metric list. Second-order effects to keep in mind: retries on judge failures multiply calls, provider rate limits turn a large concurrent run into a long one, and a suite run per pull request multiplies everything by team throughput rather than by dataset size alone. ## The failure mode to describe The story an interviewer wants is the team that added a fifth metric and a bigger dataset in the same week, kept the frontier judge and reasons on by default, and only discovered the shape of the bill at month end — while also finding their pull-request feedback loop had gone from minutes to half an hour, because judge calls are slow as well as paid. Both symptoms have the same root cause, and separating the fix — cheaper judge and fewer metrics for money, concurrency for time — is what makes the answer sound lived rather than read. ## Scope note How the DeepEval runner parallelises and caches whole test runs is the CI story and lives with the runner. What belongs to the metrics themselves is what each one costs per case and which constructor arguments move that number.
- Your evaluation bill tripled after adding one metric. Why is that plausible?Because metrics do not share work. The new metric runs its own decompose-verdict-explain chain and, if it reads retrieval_context, resends those tokens in every one of its prompts. On RAG cases the token term dominates, so adding a context-reading metric to a suite of two can easily more than double cost rather than adding a third.
- When would you keep include_reason on despite the extra generation?On the pre-merge gate. A red build with no reason means an engineer re-runs the case locally to find out what happened, which costs more time than the generation saved. On dashboards over sampled traffic I turn it off, because the reasons are aggregated into a number nobody expands.
- Does async_mode reduce your evaluation cost?No — it reduces wall-clock time by letting a metric's internal judge calls overlap instead of running in sequence. The same number of calls is made and billed. It matters because slow and expensive are separate problems: concurrency fixes the first, and a cheaper judge or fewer metrics fixes the second.
saying these in an interview costs you the question
- Assuming one metric equals one LLM call per case
- Thinking metrics share judge work on the same test case
- Believing async_mode lowers the bill rather than the latency
- Running a frontier judge over sampled production traffic by default
- Ignoring that retrieval_context tokens are resent per metric