skip to content

How do you run Ragas against a LlamaIndex or Haystack pipeline?

level: middleimportance: nice to knowfreq 30%

answer

  1. an integration is only glue to samples
  2. first-party list is shorter than you think
  3. LlamaIndex in the box, Haystack not
  4. the Haystack component ships elsewhere
  5. you can always build samples by hand

basics

~20 s

Ragas ships a first-party llama_index integration module, but no Haystack one — Haystack support is the separate deepset-published ragas-haystack distribution providing a RagasEvaluator component. Either way, the universal fallback is to build SingleTurnSample objects yourself from whatever your pipeline returns.

solid answer

~40 s

Ragas 0.4.3 ships an `ragas.integrations` package containing modules for ag_ui, amazon_bedrock, griptape, helicone, langchain, langgraph, langsmith, llama_index, opik, r2r, swarm and tracing (langfuse, mlflow). LlamaIndex is in that list; Haystack is not. Inside ragas itself the only Haystack surface is `HaystackEmbeddingsWrapper` in `ragas.embeddings.haystack_wrapper`. The Haystack pipeline component people mean when they say "ragas in Haystack" is `RagasEvaluator`, and it ships in a third-party distribution — `ragas-haystack` by deepset — imported as `from haystack_integrations.components.evaluators.ragas import RagasEvaluator`. Its constructor is `RagasEvaluator(ragas_metrics, concurrency_limit=4)`; there is no evaluator-LLM parameter, because each metric arrives already carrying its own model. You call `evaluator.run(query=..., documents=[...], reference=...)` and read `output["result"]`. The point worth making: an integration is convenience, not a requirement. Ragas scores plain samples, so any framework can be evaluated by mapping its output into `SingleTurnSample`.

code

python · 8 lines
python
from haystack_integrations.components.evaluators.ragas import RagasEvaluator
from openai import AsyncOpenAI
from ragas.llms import llm_factory
from ragas.metrics.collections import Faithfulness

judge = llm_factory("gpt-4o-mini", client=AsyncOpenAI())
evaluator = RagasEvaluator(ragas_metrics=[Faithfulness(llm=judge)], concurrency_limit=4)
print(type(evaluator).__name__)

go deeper

for a junior

Know that ragas scores plain sample objects, so any pipeline can be evaluated by collecting its question, retrieved contexts and answer. Recognise that some frameworks have ready-made glue and some do not.

for a middle

Name the packaging split: LlamaIndex has a first-party module inside ragas.integrations, while the Haystack RagasEvaluator comes from the separate ragas-haystack distribution under the haystack_integrations namespace. Describe the component's constructor and run arguments.

for a senior

Show judgement about the seam. Be explicit about which retrieved chunks you count as context, since that choice moves faithfulness and context-precision scores, and prefer a hand-written adapter when you want that control or want to avoid three-way version pinning.

for a principal

Decide organisationally whether evaluation couples to framework integrations at all. A thin internal adapter to SingleTurnSample keeps evaluation portable across the framework churn that a multi-team AI estate will otherwise inflict on every eval suite.

## What "integration" means here Ragas evaluates *data*, not pipelines. A `SingleTurnSample` is a handful of text fields — `user_input`, `retrieved_contexts`, `response`, optionally `reference` — and every metric works off those. An integration, therefore, is nothing more than glue that pulls those fields out of some framework's own objects so you do not have to write the mapping by hand. Nothing about ragas requires an integration to exist for your stack. That framing matters because it answers the interview question underneath the question: "what if we use a framework ragas doesn't support?" The answer is that you write a ten-line adapter from your pipeline's output to `SingleTurnSample` and everything else works unchanged. ## What ragas actually ships (0.4.3) The `ragas.integrations` package contains modules for: ag_ui, amazon_bedrock, griptape, helicone, langchain, langgraph, langsmith, llama_index, opik, r2r, swarm, and a tracing subpackage with langfuse and mlflow. So LlamaIndex has a first-party module inside ragas, alongside LangChain and LangGraph. Haystack does **not** have one, and this trips people up because a Haystack `RagasEvaluator` demonstrably exists. The resolution is a packaging fact rather than an API fact. ## The Haystack surface lives in someone else's distribution Inside ragas the only Haystack-specific object is `HaystackEmbeddingsWrapper`, in `ragas.embeddings.haystack_wrapper` — an adapter so a Haystack embedding model can serve as a ragas evaluator embedding model. That is it. The pipeline component is shipped by deepset, in the distribution `ragas-haystack` (4.1.2), under Haystack's own integrations namespace: `from haystack_integrations.components.evaluators.ragas import RagasEvaluator`. So `pip install ragas` alone will never give it to you; you install `ragas-haystack` and import from `haystack_integrations`, which is the sort of detail that only shows up once you have actually built the thing. Its shape is small and worth knowing precisely: - Constructor: `RagasEvaluator(ragas_metrics, concurrency_limit=4)`. `ragas_metrics` is a list of already-constructed ragas metrics; `concurrency_limit` bounds how many judge calls run at once. - There is **no** evaluator-LLM parameter on the component. This surprises people who expect to hand the component a judge. Instead each metric arrives fully configured with its own model — typically a class from ragas's modern `ragas.metrics.collections` API constructed with `llm=...`. - You run it as `evaluator.run(query=..., documents=[...], reference=...)` and read the score out of `output["result"]`. The `documents` argument taking Haystack `Document` objects is exactly the mapping work the integration exists to do — it turns them into the retrieved-contexts strings ragas metrics expect. ## LlamaIndex On the LlamaIndex side, `ragas.integrations.llama_index` is first-party: you install ragas and it is there. Its job is the same glue — adapting a LlamaIndex query engine and a set of questions into the sample shape ragas metrics score — so that you can evaluate a query engine without hand-assembling contexts and responses from the engine's response objects. Because LlamaIndex framework internals (retrievers, query engines, response modes) belong to LlamaIndex, the thing to be precise about in an interview is the seam: ragas needs the retrieved node texts as `retrieved_contexts` and the synthesised answer as `response`, and the integration's whole value is extracting those two reliably. ## The fallback beats memorising module names If your stack is a bare provider SDK, an internal framework, or a service behind HTTP, do the mapping yourself: - Run your pipeline over each question. - Capture the retrieved context strings and the final answer. - Build a `SingleTurnSample` per question, collect them into an `EvaluationDataset`. - Call `evaluate()` with your metrics and wired judge. This is also often the *better* choice even when an integration exists, because you control exactly which retrieved chunks count as context — a decision that materially changes context-precision and faithfulness scores, and one you would rather make explicitly than inherit from an adapter. ## Watch the version coupling A third-party integration distribution is pinned against a range of ragas versions and a range of the host framework's versions. `ragas-haystack` upgrading, ragas upgrading, and Haystack upgrading are three independent release trains, and the integration sits at the intersection. Budget for the occasional pin conflict, and prefer the hand-rolled adapter if you cannot tolerate that coupling in CI.

  • Why does the Haystack RagasEvaluator take no evaluator-LLM argument?
    Because the judge is attached to the metrics, not to the component. You construct each metric with its own llm — typically from the ragas.metrics.collections API — and pass the finished metric objects in as ragas_metrics. The component's only knobs are that list and concurrency_limit.
  • Your team uses a framework with no ragas integration at all. What changes?
    Almost nothing. Run your pipeline yourself, capture the question, the retrieved context strings and the final answer, build a SingleTurnSample per case and collect them into an EvaluationDataset. Metrics and judge wiring are identical from that point on — the integration was only saving you the mapping code.
  • What is the downside of relying on a third-party integration distribution in CI?
    Version coupling across three independent release trains — ragas, the host framework, and the integration package that pins both. An upgrade to any one can produce a resolver conflict that blocks your build for reasons unrelated to your code. A hand-written sample adapter has only one dependency to track.

saying these in an interview costs you the question

  • Assumes pip install ragas provides the Haystack component
  • Expects a first-party integration for every framework
  • Passes an evaluator LLM to the Haystack component itself
  • Thinks ragas instruments a pipeline rather than scoring samples
  • Believes no integration means the framework cannot be evaluated

context