How do you measure judge-model token usage and cost for a Ragas evaluation run?
answer
- measurement is opt-in, not automatic
- a parser reads usage off the responses
- you supply the per-token rates
- samples times metrics times calls each
- estimate on fifty, then extrapolate
basics
~10 sPass a token usage parser to evaluate() — for OpenAI-style judges, token_usage_parser=get_token_usage_for_openai from ragas.cost. The result then answers result.total_tokens() and result.total_cost(cost_per_input_token=..., cost_per_output_token=...). Without a parser, no usage is collected.
solid answer
~50 sRagas collects judge-model token usage only if you ask it to. You pass a parser via `evaluate(..., token_usage_parser=get_token_usage_for_openai)` — the parser's job is to pull input and output token counts out of the LLM response objects your evaluator model returns. Afterwards, `result.total_tokens()` gives the aggregate input and output token counts for the run, and `result.total_cost(cost_per_input_token=..., cost_per_output_token=...)` converts that into money using rates you supply, since the library does not carry a price list. The reason this matters is the multiplier: cost scales with samples × metrics × calls-per-metric, and several metrics issue more than one judge call per sample (claim extraction, then verdicts). Retries add on top. The professional habit is to run fifty samples with the parser attached, read the actual per-sample cost, and extrapolate before pointing a run at a few thousand rows.
code
python · 17 linesfrom ragas import evaluate
from ragas.cost import get_token_usage_for_openai
result = evaluate(
dataset=dataset,
metrics=metrics,
llm=evaluator_llm,
token_usage_parser=get_token_usage_for_openai,
)
print(result.total_tokens())
print(
result.total_cost(
cost_per_input_token=0.15 / 1_000_000,
cost_per_output_token=0.60 / 1_000_000,
)
)go deeper
Know that every LLM-based Ragas metric is a paid judge call, and that usage tracking has to be switched on for the run to report tokens at all.
Explain the opt-in parser, the two result accessors for tokens and cost, and why you supply the per-token prices yourself rather than the library knowing them.
Show the estimation ritual — measure on a small stratified subset, extrapolate, and identify which multiplier dominates (calls per metric, context size, or retries) before choosing a lever.
Own evaluation as a budget line: record cost as run metadata, decide which metrics justify running on every change versus nightly, and argue judge choice with measured figures rather than list-price intuition.
## Why this is a real question and not trivia Every LLM-based metric in Ragas is a judge model call. That is easy to forget while iterating on ten samples, and expensive to remember at three thousand. Teams routinely discover that the evaluation of a change costs more than serving the feature did for the same period, because evaluation runs at full metric fan-out over every row while production serves one call per user request. ## Turning measurement on Usage tracking is opt-in. `evaluate()` accepts `token_usage_parser`, a callable that knows how to read token counts out of whatever LLM response object your evaluator produces. Ragas ships `get_token_usage_for_openai` in `ragas.cost` for OpenAI-shaped responses. Attach it and the run accumulates usage as it goes. With the parser attached, the returned `EvaluationResult` answers two questions: - `result.total_tokens()` — the run's aggregate token usage, split into input and output. - `result.total_cost(cost_per_input_token=..., cost_per_output_token=...)` — that usage multiplied by the per-token rates you pass in. Ragas deliberately does not embed a price table, because provider pricing changes and would go stale in a pinned library version. You supply the rates, which also means you can supply your own negotiated or internal chargeback rates. Without the parser, the result has no usage to report. There is no retroactive accounting — you cannot ask a finished run what it cost if you did not measure while it ran. ## Where the cost actually comes from The naive model is samples × metrics. The real model has three more multipliers: **Calls per metric.** Several metrics decompose the work. A faithfulness-style metric typically extracts claims from the answer in one call, then verifies claims against the context in further calls. One "metric score" can be several judge requests. **Input size.** RAG evaluation prompts carry the retrieved contexts. Contexts are the biggest part of the payload, so a pipeline that retrieves ten chunks costs roughly twice as much to evaluate as one retrieving five, for identical sample counts. Input tokens usually dominate output tokens in this workload by a wide margin. **Retries.** Every retried call is billed. A run pushed hard against a throttling provider can bill a multiple of its nominal cost for the same set of scores. ## The estimation ritual Before any large run: take a random fifty samples from the dataset, run the exact metric set with the parser attached, read `total_tokens()` and `total_cost(...)`, divide by fifty, and multiply by the real dataset size. Do this again whenever the metric set, the judge model, the retrieval depth, or the prompt changes — any of the four moves the per-sample cost. The estimate also tells you which lever to pull if the number is unacceptable. If input tokens dominate, the lever is context size or dataset size, not judge choice. If one metric accounts for most of the calls, the lever is dropping or downgrading that metric. If retries are inflating the total, the lever is concurrency. ## The levers, in rough order of leverage - **Fewer samples.** A well-chosen, stratified few hundred usually tells you what a few thousand does, at a tenth the cost. - **Fewer metrics per run.** Score the two metrics that gate your decision on every run; keep the diagnostic long tail for occasional deep dives. - **A cheaper judge** — but verify it first, because a weaker judge produces more unparseable outputs and therefore more unscored rows, and its scores may not be comparable to your existing history. - **Caching**, so an iteration that re-scores an unchanged dataset does not re-pay for identical judge calls. - **Lower concurrency**, to stop paying for retried requests. ## Reporting it A mature setup records tokens and cost as run metadata beside the scores. That single number turns "we should evaluate more often" from an opinion into a budget line, and it is what lets you argue for a nightly full run versus a per-change subset with actual figures rather than intuition.
- Why does Ragas make you pass the per-token prices instead of knowing them?Provider pricing changes far faster than a pinned library version, so an embedded price table would silently go stale and produce confidently wrong numbers. Supplying the rates also lets you use negotiated pricing, an internal chargeback rate, or a self-hosted cost model rather than list price.
- Your evaluation cost is dominated by input tokens. Which levers actually help?Ones that shrink the prompt or the row count: fewer samples, fewer metrics per run, and smaller retrieved contexts, since RAG evaluation prompts carry the contexts and those dominate the payload. Switching to a judge with cheaper output tokens barely moves an input-dominated bill, and caching helps only for repeated identical calls.
- How do retries show up in the cost figures?As additional billed calls for the same score. A run pushed hard against a throttling provider retries with backoff, and each attempt consumes input tokens for the full prompt. That is why concurrency tuning is a cost lever: sizing workers under the provider's limit can cut the bill without changing a single score.
saying these in an interview costs you the question
- Assuming Ragas tracks token cost automatically
- Estimating cost as samples times metrics, ignoring calls per metric
- Forgetting that retrieved contexts dominate the judge prompt size
- Treating retried calls as free
- Swapping to a cheaper judge without checking its unscored-row rate