Why does each Ragas metric cost more than one judge LLM call per sample?
answer
- metrics are pipelines, not prompts
- fan-out per chunk, per claim
- context tokens dominate the bill
- strictness is a direct multiplier
- cut the metric list first
basics
~20 sRagas metrics are multi-step judge pipelines, not single prompts. Faithfulness extracts statements then verifies each against context; the context precision metrics issue a verdict per retrieved chunk; FactualCorrectness decomposes two texts and compares claim sets. Cost scales with metrics times samples times chunks.
solid answer
~50 sAlmost every Ragas metric is a small pipeline. `Faithfulness` first asks the judge to break the response into statements, then asks it to verify those statements against the retrieved context — at least two rounds, and more as the answer gets longer. The `LLMContextPrecision*` classes judge usefulness chunk by chunk, so a sample retrieved at k=10 costs roughly ten times a sample retrieved at k=1. `LLMContextRecall` decomposes the reference and attributes each piece. `FactualCorrectness` decomposes two texts and runs entailment across both claim sets. `ResponseRelevancy` adds embedding calls on top of a generation call, one per generated question controlled by `strictness`, and `AspectCritic` multiplies by its own `strictness`. So per-run cost is metrics x samples x per-sample fan-out, and the fan-out is driven by answer length and retrieval depth. The levers that actually matter are cutting the metric list to what you will act on, and sizing the dataset deliberately rather than pointing the suite at everything you have.
go deeper
Know that every Ragas metric calls an LLM to produce its score, so an evaluation run costs real money and time proportional to how many samples and metrics you use.
Be ready to explain why a metric is multi-step — statement extraction then verification, or a verdict per retrieved chunk — and to name strictness as a direct call multiplier.
Show the arithmetic before the run: calls per sample per metric, times samples, with retrieved context dominating input tokens; then cut the metric list to what drives a decision and size the dataset deliberately.
Own the evaluation budget as a standing cost line. Decide which metrics are worth paying for continuously versus per release, set the judge-model policy, and make adding a metric a decision with a price attached rather than a pull request.
## The mental model people arrive with Most engineers assume an eval metric is one prompt: send the question, the context and the answer, get back a number. Under that model, four metrics on two thousand samples is eight thousand calls, which sounds affordable. The bill then arrives an order of magnitude higher and nobody can explain it. The reason is that Ragas metrics are compositions of judge calls, and several of them fan out with properties of your data rather than being constant per sample. ## Where the fan-out comes from **Decomposition-then-verification.** `Faithfulness` cannot check grounding in one shot, because "is this paragraph supported" is not a well-posed question — a paragraph is many claims, some supported, some not. So it extracts statements from the response first, then judges them against the retrieved context. Two stages minimum, and the verification stage grows with how many statements the answer contained. A one-line answer and a six-paragraph answer are not the same cost. **Per-chunk judging.** The `LLMContextPrecision*` classes ask, for each retrieved chunk in order, whether that chunk was useful. That is inherently per-chunk work, so per-sample cost tracks your retrieval depth directly. Raising k from 5 to 20 to improve recall roughly quadruples what your context precision metric costs — a coupling between a retrieval-tuning decision and an evaluation bill that surprises people. **Two-sided comparison.** `FactualCorrectness` decomposes both the response and the reference, then runs entailment across both claim sets. Both sides scale with text length. **Generate-then-embed.** `ResponseRelevancy` has the judge write candidate questions from the response, then embeds them and the original question to compare. `strictness` controls how many questions, so it is a direct multiplier. Embedding calls are far cheaper than judge calls, but they are a second provider dependency and a second thing that can rate-limit. **Explicit self-consistency.** `AspectCritic` exposes `strictness` for majority voting; strictness 3 is three times the calls for that metric. ## Doing the arithmetic before the run The estimate worth doing on a whiteboard: for each metric, per-sample calls; times the number of samples; times a per-call token cost dominated by the retrieved context, which you are pasting into nearly every judge prompt. That last point is the one people miss. The dominant token driver is usually not the answer, it is the context. Every grounding and context metric sends chunks to the judge, so a system retrieving 10 chunks of 800 tokens is putting roughly 8k tokens into many of those calls. Multiply by fan-out and by samples and the input tokens, not the output tokens, are the bill. ## Levers that actually work **Cut the metric list.** The single biggest saving. Four metrics where two would drive the same decision is double the spend for no extra decision quality. Ask of each metric: if this moved, what would we do? If there is no answer, it is a dashboard ornament. **Size the dataset deliberately.** A curated few hundred samples that covers your query clusters detects real regressions; ten thousand samples costs twenty times more and mostly adds duplicates of the easy cases. Evaluation datasets should be designed, not accumulated. **Match judge to metric.** Not every metric needs your most capable model. Binary critics with a crisp definition are far more tolerant of a cheaper judge than a five-band rubric or a claim-decomposition metric, where subtle judgement is the whole point. Mixing judges is legitimate — but pin which metric uses which, because changing a judge changes the baseline just as much as changing your application. **Watch retrieval depth.** If you evaluate at the same k you serve at, remember that raising k for quality raises evaluation cost superlinearly across the suite. ## The scaling failure to name in an interview The common production incident is a suite that was fine on 50 development samples and becomes unaffordable — or starts hitting provider rate limits — when someone points it at the full dataset in CI. Nothing broke; the fan-out was always there and only became visible at volume. The senior answer is that you estimate calls per sample per metric before scaling the dataset, and you treat the metric list itself as a budget with a fixed size rather than something that only ever grows.
- Which Ragas metric's per-sample cost scales with your retrieval depth k?The LLMContextPrecision classes, both WithReference and WithoutReference, because they judge usefulness chunk by chunk in rank order. Evaluating a system that retrieves twenty chunks costs roughly four times one retrieving five, on that metric alone. It couples a retrieval-tuning decision to your evaluation bill, which is worth flagging before someone raises k to chase recall.
- Why do input tokens rather than output tokens dominate a Ragas run's cost?Because nearly every grounding and context metric pastes the retrieved chunks into the judge prompt, and chunks are the largest text in the sample. Judge outputs are short — verdicts, small claim lists — while inputs carry thousands of context tokens repeated across each stage of a multi-step metric. Estimating cost from the answer length badly underestimates it.
- Is using a cheaper judge model for some metrics defensible?Yes, if you match it to the difficulty of the judgement and pin it. A crisply worded binary critic tolerates a smaller model well; claim decomposition and multi-band rubrics do not, because fine discrimination is the whole task. The hard rule is that the judge is part of the metric definition, so record which model scored which metric and treat a swap as a baseline break.
saying these in an interview costs you the question
- Assuming one metric equals one LLM call per sample
- Estimating cost from output tokens only
- Adding metrics nobody would act on
- Scaling the dataset before estimating calls per sample
- Swapping judge models without rebaselining