skip to content

LLM Wiring & Integrations

Ragas needs an evaluator LLM and embedding model of its own, plus a way to plug into whatever pipeline you already run. Expect questions about choosing a judge model that is stronger, or at least different, from the one under test.

on this pageshow

questions

6

In Ragas, how do you give metrics an evaluator LLM and embedding model?

level: middleimportance: must knowfreq 78%

answer

  1. ragas ships no model of its own
  2. two model kinds, two adapters
  3. wrap the LangChain object, not raw
  4. llm= and embeddings= on evaluate()
  5. a metric's own llm beats the run default

basics

~20 s

Wrap a LangChain chat model in LangchainLLMWrapper and a LangChain embeddings object in LangchainEmbeddingsWrapper, then pass them to evaluate() as llm= and embeddings=, or set them on individual metric objects. Ragas ships no model of its own.

solid answer

~40 s

Ragas metrics are scoring logic, not models: nearly every metric calls a judge LLM, and a few also call an embedding model. You supply both through adapter wrappers — `LangchainLLMWrapper` from `ragas.llms` around any LangChain chat model, and `LangchainEmbeddingsWrapper` from `ragas.embeddings` around a LangChain embeddings object. There are two places to attach them. Run-level: `evaluate(dataset=..., metrics=[...], llm=evaluator_llm, embeddings=evaluator_embeddings)` becomes the default for every metric in that run. Metric-level: `Faithfulness(llm=strong_judge)` pins that one metric, and a metric's own model wins over the run-level default. If you supply neither, ragas falls back to its own default factory, which builds an OpenAI-backed client — so the run either dies on a missing OPENAI_API_KEY or quietly bills an account with a model you never chose. Be explicit, and record which judge produced the scores.

code

python · 24 lines
python
from langchain_openai import ChatOpenAI, OpenAIEmbeddings
from ragas import EvaluationDataset, SingleTurnSample, evaluate
from ragas.embeddings import LangchainEmbeddingsWrapper
from ragas.llms import LangchainLLMWrapper
from ragas.metrics import Faithfulness, ResponseRelevancy

evaluator_llm = LangchainLLMWrapper(ChatOpenAI(model="gpt-4o-mini", temperature=0))
evaluator_embeddings = LangchainEmbeddingsWrapper(OpenAIEmbeddings())

dataset = EvaluationDataset(samples=[
    SingleTurnSample(
        user_input="Who wrote Hamlet?",
        retrieved_contexts=["Hamlet is a tragedy written by William Shakespeare."],
        response="William Shakespeare wrote Hamlet.",
    )
])

result = evaluate(
    dataset=dataset,
    metrics=[Faithfulness(), ResponseRelevancy()],
    llm=evaluator_llm,
    embeddings=evaluator_embeddings,
)
print(result)

go deeper

for a junior

Remember the shape: ragas needs its own judge model, and you give it one by wrapping a LangChain chat model in LangchainLLMWrapper and passing it to evaluate() as llm=. Say plainly that ragas does not bring a model with it.

for a middle

Explain both attachment points — run-level llm=/embeddings= on evaluate() versus llm= on a metric constructor — and state the resolution order, with the metric's own model winning. Mention that a missing wiring falls through to an OpenAI-backed default.

for a senior

Show that you treat the judge as production infrastructure: a pinned model, its own API key and quota so evaluation cannot starve live traffic, and the judge model recorded next to every score so results stay comparable across runs.

for a principal

Own the standard. Decide organisation-wide which judge model evaluations are allowed to use, how that choice is versioned and rolled forward, and who pays for it — because a judge swap silently re-baselines every score anyone has ever quoted.

## Ragas ships metrics, not models Ragas is a library of RAG evaluation metrics — Faithfulness, ResponseRelevancy, LLMContextRecall, LLMContextPrecisionWithReference, FactualCorrectness, NoiseSensitivity, AspectCritic, RubricsScore. Almost all of them are LLM-as-judge metrics: to score one sample the metric builds a prompt, sends it to a language model, and parses the verdict. Ragas does not bundle a model and cannot guess which provider account you are entitled to spend. Supplying an evaluator model is therefore not an optional optimisation — it is step one of every ragas setup, and it is the first thing an interviewer checks you have actually done rather than copied from a quickstart. ## The two wrappers Ragas defines its own narrow interfaces for "a thing that generates text" and "a thing that embeds text", and ships adapters that make third-party objects fit: - `from ragas.llms import LangchainLLMWrapper` — takes a LangChain chat model (`ChatOpenAI`, `AzureChatOpenAI`, `ChatAnthropic`, anything with a LangChain chat class) and exposes it as a ragas evaluator LLM. - `from ragas.embeddings import LangchainEmbeddingsWrapper` — takes a LangChain embeddings object (`OpenAIEmbeddings`, `AzureOpenAIEmbeddings`, a local embedding class) and exposes it as a ragas embedding model. The wrappers are pure adapters. They do not train, cache, or re-configure anything; they translate ragas's calls into the framework's calls, including the async paths the metrics use when they score concurrently. Because the wrapper is the seam, ragas gets provider breadth for free: anything LangChain can talk to, ragas can judge with. Ragas 0.4.x also exposes `llm_factory`, a shorthand that constructs an evaluator LLM directly from a model name and a client — useful when you are not otherwise a LangChain user, and the form the newer `ragas.metrics.collections` classes are usually shown with. ## Two places the wiring can live **Run level.** `evaluate()` accepts `llm=` and `embeddings=`. Whatever you pass becomes the default for every metric in the list that has not been given its own. This is the common case: one judge, one embedding model, one line each. **Metric level.** Metric constructors accept `llm=` (and, where relevant, `embeddings=`). A metric constructed with its own model uses that model, regardless of what `evaluate()` was given. This is how you run a cheap judge for most metrics and an expensive one for the metric you trust least, in a single run. The resolution order is simply: the metric's own model if it has one, otherwise the run-level model, otherwise the library default. Knowing that order is what stops the classic confusion of "I passed `llm=` to evaluate() and my bill still shows the other model" — a metric you constructed earlier with its own judge was ignoring you. ## What happens if you wire nothing Ragas will not simply refuse. It falls back to its default factory, which builds an OpenAI-backed client. In practice that means one of two outcomes, both bad: the run raises because `OPENAI_API_KEY` is not set, or it succeeds and silently charges an OpenAI account with a default model you did not select and cannot cite when someone asks what produced the score. Neither is acceptable in a repeatable evaluation, so treat the wrapper lines as mandatory boilerplate. ## Judge and application are separate objects A frequent misconception is that ragas reuses the model your RAG app already runs. It does not — it has no visibility into your pipeline at all. You hand it samples (question, retrieved contexts, response, sometimes a reference) and a judge, and those two things are independent. That separation is deliberate: it is what lets you evaluate an application built on one model using a different, or stronger, model as the judge, and what lets the judge live on a separate API key with its own quota so an evaluation run cannot exhaust the rate limit your production traffic depends on. ## Cost follows directly from this wiring The object you wrap is the object that gets billed, once per judge call, and metrics make more than one call per sample. Multiply metrics × samples × calls-per-metric and the choice of model inside `LangchainLLMWrapper` is the single biggest lever on what an evaluation run costs. That is why the wrapper line, boring as it looks, is worth arguing about in review.

  • A metric was constructed with its own llm and you also pass llm= to evaluate(). Which one scores it?
    The metric's own model. Run-level `llm=` is a default applied to metrics that have none; it never overrides a model already attached to a metric object. This is the usual explanation for "my evaluate() judge setting had no effect" — some metric in the list was constructed earlier with a judge baked in.
  • Does the evaluator model have to come from LangChain?
    No. The LangChain wrappers are the widest path because any provider with a LangChain chat or embeddings class works through them, but ragas 0.4 also exposes `llm_factory` to build an evaluator LLM from a model name and a client directly, and the newer `ragas.metrics.collections` classes are typically constructed that way. The metric only needs something that satisfies the ragas LLM interface.
  • What extra setup does Azure OpenAI need compared with OpenAI?
    Two deployments, not one. You wrap the Azure chat model in `LangchainLLMWrapper` and a separate Azure embeddings model in `LangchainEmbeddingsWrapper`, each pointing at its own deployment name and endpoint. Teams routinely deploy only the chat model, wire it, and then hit a failure the first time an embedding-based metric runs.

saying these in an interview costs you the question

  • Assumes ragas reuses the application's own LLM automatically
  • Passes a raw LangChain model without the ragas wrapper
  • Wires a judge LLM but no embedding model at all
  • Thinks the wrapper fine-tunes or configures the model
  • Believes evaluate(llm=...) overrides a metric's own judge

context

open as a page

Which Ragas metrics need an embedding model, not just a judge LLM?

level: middleimportance: should knowfreq 40%

basics

~20 s

Most ragas metrics are pure LLM-as-judge and need only an evaluator LLM. Similarity-based ones — ResponseRelevancy above all — additionally embed text and compare vectors, so a run containing them fails or misbehaves if you wired an LLM but no embeddings.

open as a page

How do you point Ragas at a judge model different from the app's model?

level: seniorimportance: should knowfreq 52%

basics

~20 s

Build a second wrapper instance around a different chat model and hand it to ragas — the evaluator LLM is a wholly separate object from anything your pipeline uses. Pin a dated model snapshot, give it its own API key, and record it with the scores.

open as a page

In Ragas, how do you export judge LLM calls to Langfuse or LangSmith?

level: seniorimportance: should knowfreq 36%

basics

~20 s

Pass a callback handler to evaluate()'s callbacks argument: because the evaluator LLM is a wrapped LangChain model, a LangChain-compatible handler sees every judge call. Ragas also ships langsmith and tracing.langfuse integration modules for the same purpose.

open as a page

How do you turn a Ragas evaluation into a CI gate on a pull request?

level: principalimportance: should knowfreq 45%

basics

~20 s

Run evaluate() over a small pinned dataset, read aggregate scores, and fail the job when a metric falls below its floor. The hard part is not the assert — it is choosing thresholds that catch regressions without going red at random, and keeping the run cheap enough to sit on every pull request.

open as a page

How do you run Ragas against a LlamaIndex or Haystack pipeline?

level: middleimportance: nice to knowfreq 30%

basics

~20 s

Ragas ships a first-party llama_index integration module, but no Haystack one — Haystack support is the separate deepset-published ragas-haystack distribution providing a RagasEvaluator component. Either way, the universal fallback is to build SingleTurnSample objects yourself from whatever your pipeline returns.

open as a page