skip to content

In Ragas, how do you give metrics an evaluator LLM and embedding model?

level: middleimportance: must knowfreq 78%

answer

  1. ragas ships no model of its own
  2. two model kinds, two adapters
  3. wrap the LangChain object, not raw
  4. llm= and embeddings= on evaluate()
  5. a metric's own llm beats the run default

basics

~20 s

Wrap a LangChain chat model in LangchainLLMWrapper and a LangChain embeddings object in LangchainEmbeddingsWrapper, then pass them to evaluate() as llm= and embeddings=, or set them on individual metric objects. Ragas ships no model of its own.

solid answer

~40 s

Ragas metrics are scoring logic, not models: nearly every metric calls a judge LLM, and a few also call an embedding model. You supply both through adapter wrappers — `LangchainLLMWrapper` from `ragas.llms` around any LangChain chat model, and `LangchainEmbeddingsWrapper` from `ragas.embeddings` around a LangChain embeddings object. There are two places to attach them. Run-level: `evaluate(dataset=..., metrics=[...], llm=evaluator_llm, embeddings=evaluator_embeddings)` becomes the default for every metric in that run. Metric-level: `Faithfulness(llm=strong_judge)` pins that one metric, and a metric's own model wins over the run-level default. If you supply neither, ragas falls back to its own default factory, which builds an OpenAI-backed client — so the run either dies on a missing OPENAI_API_KEY or quietly bills an account with a model you never chose. Be explicit, and record which judge produced the scores.

code

python · 24 lines
python
from langchain_openai import ChatOpenAI, OpenAIEmbeddings
from ragas import EvaluationDataset, SingleTurnSample, evaluate
from ragas.embeddings import LangchainEmbeddingsWrapper
from ragas.llms import LangchainLLMWrapper
from ragas.metrics import Faithfulness, ResponseRelevancy

evaluator_llm = LangchainLLMWrapper(ChatOpenAI(model="gpt-4o-mini", temperature=0))
evaluator_embeddings = LangchainEmbeddingsWrapper(OpenAIEmbeddings())

dataset = EvaluationDataset(samples=[
    SingleTurnSample(
        user_input="Who wrote Hamlet?",
        retrieved_contexts=["Hamlet is a tragedy written by William Shakespeare."],
        response="William Shakespeare wrote Hamlet.",
    )
])

result = evaluate(
    dataset=dataset,
    metrics=[Faithfulness(), ResponseRelevancy()],
    llm=evaluator_llm,
    embeddings=evaluator_embeddings,
)
print(result)

go deeper

for a junior

Remember the shape: ragas needs its own judge model, and you give it one by wrapping a LangChain chat model in LangchainLLMWrapper and passing it to evaluate() as llm=. Say plainly that ragas does not bring a model with it.

for a middle

Explain both attachment points — run-level llm=/embeddings= on evaluate() versus llm= on a metric constructor — and state the resolution order, with the metric's own model winning. Mention that a missing wiring falls through to an OpenAI-backed default.

for a senior

Show that you treat the judge as production infrastructure: a pinned model, its own API key and quota so evaluation cannot starve live traffic, and the judge model recorded next to every score so results stay comparable across runs.

for a principal

Own the standard. Decide organisation-wide which judge model evaluations are allowed to use, how that choice is versioned and rolled forward, and who pays for it — because a judge swap silently re-baselines every score anyone has ever quoted.

## Ragas ships metrics, not models Ragas is a library of RAG evaluation metrics — Faithfulness, ResponseRelevancy, LLMContextRecall, LLMContextPrecisionWithReference, FactualCorrectness, NoiseSensitivity, AspectCritic, RubricsScore. Almost all of them are LLM-as-judge metrics: to score one sample the metric builds a prompt, sends it to a language model, and parses the verdict. Ragas does not bundle a model and cannot guess which provider account you are entitled to spend. Supplying an evaluator model is therefore not an optional optimisation — it is step one of every ragas setup, and it is the first thing an interviewer checks you have actually done rather than copied from a quickstart. ## The two wrappers Ragas defines its own narrow interfaces for "a thing that generates text" and "a thing that embeds text", and ships adapters that make third-party objects fit: - `from ragas.llms import LangchainLLMWrapper` — takes a LangChain chat model (`ChatOpenAI`, `AzureChatOpenAI`, `ChatAnthropic`, anything with a LangChain chat class) and exposes it as a ragas evaluator LLM. - `from ragas.embeddings import LangchainEmbeddingsWrapper` — takes a LangChain embeddings object (`OpenAIEmbeddings`, `AzureOpenAIEmbeddings`, a local embedding class) and exposes it as a ragas embedding model. The wrappers are pure adapters. They do not train, cache, or re-configure anything; they translate ragas's calls into the framework's calls, including the async paths the metrics use when they score concurrently. Because the wrapper is the seam, ragas gets provider breadth for free: anything LangChain can talk to, ragas can judge with. Ragas 0.4.x also exposes `llm_factory`, a shorthand that constructs an evaluator LLM directly from a model name and a client — useful when you are not otherwise a LangChain user, and the form the newer `ragas.metrics.collections` classes are usually shown with. ## Two places the wiring can live **Run level.** `evaluate()` accepts `llm=` and `embeddings=`. Whatever you pass becomes the default for every metric in the list that has not been given its own. This is the common case: one judge, one embedding model, one line each. **Metric level.** Metric constructors accept `llm=` (and, where relevant, `embeddings=`). A metric constructed with its own model uses that model, regardless of what `evaluate()` was given. This is how you run a cheap judge for most metrics and an expensive one for the metric you trust least, in a single run. The resolution order is simply: the metric's own model if it has one, otherwise the run-level model, otherwise the library default. Knowing that order is what stops the classic confusion of "I passed `llm=` to evaluate() and my bill still shows the other model" — a metric you constructed earlier with its own judge was ignoring you. ## What happens if you wire nothing Ragas will not simply refuse. It falls back to its default factory, which builds an OpenAI-backed client. In practice that means one of two outcomes, both bad: the run raises because `OPENAI_API_KEY` is not set, or it succeeds and silently charges an OpenAI account with a default model you did not select and cannot cite when someone asks what produced the score. Neither is acceptable in a repeatable evaluation, so treat the wrapper lines as mandatory boilerplate. ## Judge and application are separate objects A frequent misconception is that ragas reuses the model your RAG app already runs. It does not — it has no visibility into your pipeline at all. You hand it samples (question, retrieved contexts, response, sometimes a reference) and a judge, and those two things are independent. That separation is deliberate: it is what lets you evaluate an application built on one model using a different, or stronger, model as the judge, and what lets the judge live on a separate API key with its own quota so an evaluation run cannot exhaust the rate limit your production traffic depends on. ## Cost follows directly from this wiring The object you wrap is the object that gets billed, once per judge call, and metrics make more than one call per sample. Multiply metrics × samples × calls-per-metric and the choice of model inside `LangchainLLMWrapper` is the single biggest lever on what an evaluation run costs. That is why the wrapper line, boring as it looks, is worth arguing about in review.

  • A metric was constructed with its own llm and you also pass llm= to evaluate(). Which one scores it?
    The metric's own model. Run-level `llm=` is a default applied to metrics that have none; it never overrides a model already attached to a metric object. This is the usual explanation for "my evaluate() judge setting had no effect" — some metric in the list was constructed earlier with a judge baked in.
  • Does the evaluator model have to come from LangChain?
    No. The LangChain wrappers are the widest path because any provider with a LangChain chat or embeddings class works through them, but ragas 0.4 also exposes `llm_factory` to build an evaluator LLM from a model name and a client directly, and the newer `ragas.metrics.collections` classes are typically constructed that way. The metric only needs something that satisfies the ragas LLM interface.
  • What extra setup does Azure OpenAI need compared with OpenAI?
    Two deployments, not one. You wrap the Azure chat model in `LangchainLLMWrapper` and a separate Azure embeddings model in `LangchainEmbeddingsWrapper`, each pointing at its own deployment name and endpoint. Teams routinely deploy only the chat model, wire it, and then hit a failure the first time an embedding-based metric runs.

saying these in an interview costs you the question

  • Assumes ragas reuses the application's own LLM automatically
  • Passes a raw LangChain model without the ragas wrapper
  • Wires a judge LLM but no embedding model at all
  • Thinks the wrapper fine-tunes or configures the model
  • Believes evaluate(llm=...) overrides a metric's own judge

context