skip to content

How do you point Ragas at a judge model different from the app's model?

level: seniorimportance: should knowfreq 52%

answer

  1. ragas cannot see your application's model
  2. judge is just another wrapped object
  3. metric-level llm beats the run default
  4. pin the dated snapshot, not the alias
  5. separate key so eval cannot starve prod

basics

~20 s

Build a second wrapper instance around a different chat model and hand it to ragas — the evaluator LLM is a wholly separate object from anything your pipeline uses. Pin a dated model snapshot, give it its own API key, and record it with the scores.

solid answer

~50 s

Ragas has no connection to your application's model: it receives samples and an evaluator LLM, and those are independent. So "use a different judge" is purely a wiring decision — construct a second `LangchainLLMWrapper` around whichever chat model you want as judge and pass it as `evaluate(llm=...)`. You can go finer: because a metric's own `llm=` beats the run-level default, you can pin a stronger, pricier judge to the metric you trust least (say `Faithfulness`) while the cheaper metrics run on a smaller model in the same run. Three operational habits go with it. Pin a dated model snapshot rather than a floating alias, so a provider-side model update does not silently re-baseline your scores. Give the judge its own API key or project, so an evaluation run cannot burn the rate limit production depends on. And store the judge model identifier alongside every result — a score is meaningless without knowing what produced it.

code

python · 16 lines
python
from langchain_openai import ChatOpenAI
from ragas.llms import LangchainLLMWrapper
from ragas.metrics import AspectCritic, Faithfulness

cheap_judge = LangchainLLMWrapper(ChatOpenAI(model="gpt-4o-mini", temperature=0))
strong_judge = LangchainLLMWrapper(ChatOpenAI(model="gpt-4o", temperature=0))

metrics = [
    Faithfulness(llm=strong_judge),
    AspectCritic(
        name="polite",
        definition="Is the response polite and free of blame?",
        llm=cheap_judge,
    ),
]
print([m.name for m in metrics])

go deeper

for a junior

Know that the judge model is something you choose and pass in, and that it has nothing to do with the model your app uses. Be able to say you would wrap a second chat model and hand it to evaluate().

for a middle

Explain the override mechanics — run-level default versus per-metric llm= — and why you would deliberately mix a cheap and an expensive judge in one run. Mention pinning a dated snapshot and temperature zero.

for a senior

Show the operating discipline: separate key and quota so evals cannot degrade production, judge identifier persisted with every result, and a re-baseline procedure whenever the judge changes. Diagnose a score shift by suspecting the instrument first.

for a principal

Own the standard across teams: which judges are approved, how a judge upgrade is rolled out and re-baselined, who funds evaluation spend, and how score history stays interpretable across a year of model churn.

## The judge is not the application Ragas never sees your pipeline. You hand it samples — user input, retrieved contexts, response, sometimes a reference — plus an evaluator LLM, and it scores what it was given. There is no implicit reuse of the model your RAG app calls, no shared client, no shared key unless you make it so. That means using a different judge is not a feature you enable; it is the default state of the world, and the *only* way the two could end up identical is if you deliberately pass the same model object twice. This is worth stating explicitly in an interview, because a common misconception is that ragas somehow instruments the app. It does not. Which model is under test is a fact about your data-generation step; which model judges is a fact about your `LangchainLLMWrapper` line. ## Mechanically, it is two objects ``` cheap_judge = LangchainLLMWrapper(<small model, temperature 0>) strong_judge = LangchainLLMWrapper(<large model, temperature 0>) ``` Pass one as `evaluate(llm=...)` to set the run default; pass the other into a specific metric's constructor to override it there. Since metric-level wiring wins, a single run can mix judges deliberately: the metric whose failure mode is subtle gets the expensive model, and the metrics that are little more than a yes/no classification get the cheap one. That is the main cost lever available to you inside a run, and it is a specific, checkable answer to "how would you halve your eval bill without dropping metrics". ## Pin the snapshot A floating model alias points at whatever the provider ships this month. Under an alias, an evaluation suite can move without a single commit to your repo — the judge changed underneath you, and your history now contains a discontinuity nobody recorded. Pin the dated snapshot the provider offers, set temperature to zero, and change it only as a deliberate act. When you do change it, re-run the baseline set on both judges before and after so you know the size of the shift, and annotate the point in your score history. ## Isolate the quota An evaluation run is a burst of concurrent model calls: metrics × samples × calls-per-metric, issued in parallel. Pointed at the same API key as production, it competes with live user traffic for the same rate limit, and the failure mode is that a routine eval run degrades the product. Give the judge its own key, project, or provider account. This also makes the eval bill legible as a line item instead of a smear across production spend, which matters the first time someone asks what quality assurance costs. ## Record what judged Ragas returns numbers; it does not stamp the judge onto them. If you keep results — and a CI gate or a quality dashboard means you do — persist the judge model identifier, its temperature and the ragas version next to every run. Two runs judged by different models are not comparable, and without that metadata you cannot tell after the fact which comparison you are looking at. Teams that skip this eventually argue about a regression that was a judge upgrade. ## Should the judge be stronger, or just different? The wiring supports either; what makes one choice better than another is evaluation methodology, not a ragas API. What ragas gives you is the freedom and the obligation: it will faithfully use whatever you hand it, including the exact model that produced the answers, and it will not warn you. So the tool-level discipline is (a) make the choice explicit in code rather than inherited from a default, (b) keep it pinned, and (c) keep it recorded. The interesting engineering conversation — whether a given judge's scores can be trusted — sits on top of that wiring, but it cannot even start until the wiring is deliberate. ## A practical split A pattern that survives contact with production: cheap pinned judge for the fast suite that runs on every pull request, expensive pinned judge for the nightly or pre-release suite over the larger dataset. Same metrics, two judges, two baselines, and each one's numbers only ever compared against its own history.

  • How would you cut a ragas run's judge cost without removing metrics?
    Split judges inside the run. Keep the expensive pinned model on the one or two metrics whose subtlety justifies it, and construct the rest with a small model via their own llm=. Because metric-level wiring overrides the run-level default, that is a per-metric decision, and it typically removes most of the spend while leaving the metric you actually gate on untouched.
  • Your faithfulness scores dropped five points overnight with no code change. What do you check first?
    Whether the judge changed. A floating model alias can be repointed provider-side, so the same data judged by a newer model gives different numbers. Compare the recorded judge identifier for the two runs, then re-run yesterday's data against today's judge to separate a real regression from an instrument shift.
  • Why give the evaluator LLM its own API key rather than reusing the application's?
    Isolation of quota and of accounting. An eval run issues a burst of concurrent judge calls; on a shared key it competes with live traffic for the same rate limit, so a routine evaluation can degrade the product. A separate key also makes evaluation spend a visible line item instead of noise inside production cost.

saying these in an interview costs you the question

  • Believes ragas automatically instruments the app's model
  • Uses a floating model alias as the pinned judge
  • Runs the judge on the production API key
  • Never records which judge produced a score
  • Thinks one judge must serve every metric in a run

context